korentmaj/bipropagation-study

An honest, independent reproduction & component-level decomposition of Bojan Ploj's bipropagation (greedy layer-wise supervised training), benchmarked on MNIST/MLP and CIFAR-10/CNN against tuned modern backprop.

Python

0

7 commits

updated Jun 23, 2026

See the code

See what people are saying

README

Bipropagation: A Decomposition Study Building on Dr. Bojan Ploj's Idea

A component-level decomposition study of greedy, layer-wise, supervised neural-network training, built on Dr. Bojan Ploj's bipropagation idea. Bipropagation is a greedy, layer-wise, supervised training method that trains a deep network one layer at a time using intermediate targets.

Building on Dr. Ploj's insight, this repository decomposes the approach into its parts to understand which component makes layer-wise supervised training effective, and it benchmarks the approach against carefully tuned modern backpropagation baselines on MNIST and CIFAR-10.


TL;DR / Key findings

The thesis. Building on Dr. Bojan Ploj's bipropagation idea, we decompose what makes greedy layer-wise supervised training effective. The key ingredient is per-layer supervision, and the approach is competitive and depth-robust on MNIST and CIFAR-10.

To isolate the operative mechanism, we run a deeply-supervised control: per-layer auxiliary heads driven by a single global gradient. This control matches the local-loss method at every depth (depth-16: 0.9684 vs 0.9685), which clarifies the mechanism constructively. The benefit flows from per-layer supervision itself, with a small, secondary locality effect appearing on CIFAR-10/CNN.

In short: per-layer supervision, the idea at the heart of Ploj's method, is what makes layer-wise training work, and it carries over robustly as networks get deeper.

MNIST / MLP (test accuracy vs. depth)

Full-scale run, 30k train / 10k test, seed 0:

DepthVanilla BPModern BPAnchors (Ploj-style)Local-loss
20.95870.96900.87740.9708
40.96300.97350.87430.9714
80.95860.96270.86530.9701
160.1135 (collapse)0.95130.84150.9685

The locality-isolation control (30k MNIST, seed 0, 30 epochs) shows that deeply-supervised ≈ local-loss at every depth:

DepthResidual BPPlain BP (ReLU+BN, 30ep)Deeply-supervised (global grad + aux heads)Local-loss
80.96740.96910.97250.9701
160.9351*0.96610.96840.9685

*The residual MLP at depth 16 was under-tuned within the epoch budget and is not a load-bearing baseline.

At depth 16, deeply-supervised (0.9684) ≈ local-loss (0.9685): keeping per-layer supervision while restoring the global gradient reproduces the result. The benefit comes from per-layer supervision, which is the core ingredient.

CIFAR-10 / CNN (mean test accuracy ± std over 3 seeds)

15k train / 10k test, 3 seeds, 12 epochs, no augmentation (held identical across methods):

Depth (blocks)E2E backpropGreedy local-lossDeeply-supervised
30.528 ±.0160.567 ±.0030.520 ±.022
60.643 ±.0070.649 ±.0040.577 ±.018
90.557 ±.022 ↓0.626 ±.0050.609 ±.013

On CIFAR the plain (non-residual) end-to-end CNN degrades at depth 9 (0.643 to 0.557). Both per-layer-supervised methods are more depth-robust, and here local (0.626) edges out deepsup (0.609) at depth 9, so locality contributes a small, secondary robustness on CNNs that pure deep supervision does not fully capture. The primary mechanism is still per-layer supervision.


Methods compared

All methods share one framework, architecture, and data pipeline to avoid infrastructure confounds.

MethodDescription
End-to-end (vanilla) backpropNaive init, saturating (tanh) activation, plain SGD. This is the regime where vanishing gradients bite.
Modern backpropAdam + He init + BatchNorm. The strong baseline.
Greedy local-loss (layer-wise)The greedy bipropagation scaffold, but each layer is trained with a temporary softmax head and cross-entropy (à la Belilovsky 2019 / Nøkland 2019); the head is discarded before the next layer.
Deeply-supervised controlPer-layer auxiliary classifier heads with a single global gradient (Lee 2015). The locality-isolation control: same per-layer supervision, but global backprop is retained.
Anchors (Ploj-style)Best-effort reconstruction of Ploj's hand-designed intermediate-target rule: each layer shifts its input toward per-class anchors/prototypes, weights initialized near identity, final softmax readout.
Deterministic centroid-initOne hidden layer constructed analytically from class-centroid geometry (sparse ±1 units over the 3 most-discriminative features), targets = 0.99·layer_output + 0.01·two_hot(class), refined with Adam.

Repository structure

.
├── README.md                       # this file
├── PAPER.md                        # the English paper (authoritative findings & numbers)
├── LICENSE                         # MIT
├── requirements.txt
├── .gitignore
├── experiments/
│   ├── cifar_experiment.py         # CIFAR-10 / CNN decomposition (e2e, local, deepsup)
│   └── mnist_mlp_experiment.py     # MNIST / MLP, all methods (self-contained)
└── archive/
    ├── README.md
    └── ...                         # raw development fragments, kept for provenance

How to run

The experiments are self-contained TensorFlow 2 / Keras scripts that each run in a single Colab cell or locally.

Local

pip install -r requirements.txt
python experiments/cifar_experiment.py        # CIFAR-10 / CNN
python experiments/mnist_mlp_experiment.py    # MNIST / MLP

Colab

Upload a script (or paste it into a cell) and run. A GPU runtime (e.g. T4) is recommended for the full configs.

Notes

  • FAST_MODE flag. Each script has a FAST_MODE toggle near the top. True gives a small, fast indicative smoke run (subset of data, few epochs, 2 to 3 seeds); set it to False for the full benchmark reported in the paper.
  • CIFAR-10 download. cifar_experiment.py downloads CIFAR-10 from cs.toronto.edu via tf.keras.datasets on first run and caches it to disk; subsequent runs reuse the cache.
  • All reported numbers come from actual evaluation runs.

Limitations

  • Seeds. Most decisive numbers (full-scale and control runs) are single-seed (seed 0); the multi-seed evidence is currently FAST_MODE / CIFAR only. A fuller protocol (≥10 seeds, 95% CIs, paired Holm-Bonferroni tests) is left for follow-up.
  • Plain, non-residual baselines. Both testbeds compare against plain baselines that degrade with depth for known optimization reasons. A residual/normalized end-to-end baseline would likely close the depth gap, so the depth-robustness claims are relative to plain architectures, not modern residual networks.
  • Iso-compute. Local-loss sees the data roughly 3-6x more often than a single end-to-end run; a clean accuracy-vs-wall-clock and iso-gradient-step accounting is still outstanding.
  • Reconstruction of Ploj's rule. The "anchors" method reconstructs an unpublished multi-class target rule (the original MNIST.m is auth-walled on ResearchGate). A more faithful target scheme could raise the anchors numbers, though it would not change the per-layer-supervision-is-the-key-ingredient conclusion.

Credit

The bipropagation method and the underlying intuition, that per-layer supervision can help train deep networks, originate with Dr. Bojan Ploj. This repository builds on his work, decomposing it to understand why per-layer supervision is so effective; the credit for the original idea is his. We thank him for making the method and code public, which is what made this study possible.

Dr. Ploj's repositories:

Related foundational work this study builds on includes Deeply-Supervised Nets (Lee et al. 2015), greedy layer-wise learning at scale (Belilovsky et al. 2019), local error signals (Nøkland & Eidnes 2019), and Difference Target Propagation (Lee et al. 2015). See PAPER.md for the full reference list.


Contributing / further research

Contributions and extensions are warmly welcome. This is intended as an open, constructive starting point. Particularly valuable directions:

  • Residual / normalized end-to-end baselines. How does the depth-robustness gap look against a properly modern baseline?
  • More seeds + confidence intervals. ≥10 seeds, 95% CIs, paired Holm-Bonferroni tests on the full-scale and control tables.
  • Harder data. CIFAR-100, Tiny-ImageNet.
  • Other local-learning methods. Forward-Forward, Difference Target Propagation, synthetic gradients, feedback alignment, as additional points of comparison.
  • A more faithful reconstruction of Ploj's multi-class intermediate-target rule (ideally from the original MNIST.m).
  • Iso-compute accounting. Accuracy vs. wall-clock and vs. gradient steps.

Open an issue or a pull request.


Citation

@misc{korent2026bipropagation,
  author       = {Korent, Maj},
  title        = {Bipropagation: A Decomposition Study Building on Dr. Bojan Ploj's Idea},
  year         = {2026},
  howpublished = {\url{https://github.com/korentmaj/bipropagation-study}},
  note         = {Decomposition study building on Bojan Ploj's bipropagation method.}
}

This study is offered in a spirit of constructive collaboration. The aim is to understand and build on the genuine insight at the core of the method, and to credit it generously.

backpropagation
bipropagation
deep-learning
layer-wise-training
machine-learning
neural-networks
reproducibility
tensorflow

korentmaj/bipropagation-study

An honest, independent reproduction & component-level decomposition of Bojan Ploj's bipropagation (greedy layer-wise supervised training), benchmarked on MNIST/MLP and CIFAR-10/CNN against tuned modern backprop.

Python

0

7 commits

updated Jun 23, 2026

See the code

See what people are saying

README

Bipropagation: A Decomposition Study Building on Dr. Bojan Ploj's Idea

A component-level decomposition study of greedy, layer-wise, supervised neural-network training, built on Dr. Bojan Ploj's bipropagation idea. Bipropagation is a greedy, layer-wise, supervised training method that trains a deep network one layer at a time using intermediate targets.

Building on Dr. Ploj's insight, this repository decomposes the approach into its parts to understand which component makes layer-wise supervised training effective, and it benchmarks the approach against carefully tuned modern backpropagation baselines on MNIST and CIFAR-10.


TL;DR / Key findings

The thesis. Building on Dr. Bojan Ploj's bipropagation idea, we decompose what makes greedy layer-wise supervised training effective. The key ingredient is per-layer supervision, and the approach is competitive and depth-robust on MNIST and CIFAR-10.

To isolate the operative mechanism, we run a deeply-supervised control: per-layer auxiliary heads driven by a single global gradient. This control matches the local-loss method at every depth (depth-16: 0.9684 vs 0.9685), which clarifies the mechanism constructively. The benefit flows from per-layer supervision itself, with a small, secondary locality effect appearing on CIFAR-10/CNN.

In short: per-layer supervision, the idea at the heart of Ploj's method, is what makes layer-wise training work, and it carries over robustly as networks get deeper.

MNIST / MLP (test accuracy vs. depth)

Full-scale run, 30k train / 10k test, seed 0:

DepthVanilla BPModern BPAnchors (Ploj-style)Local-loss
20.95870.96900.87740.9708
40.96300.97350.87430.9714
80.95860.96270.86530.9701
160.1135 (collapse)0.95130.84150.9685

The locality-isolation control (30k MNIST, seed 0, 30 epochs) shows that deeply-supervised ≈ local-loss at every depth:

DepthResidual BPPlain BP (ReLU+BN, 30ep)Deeply-supervised (global grad + aux heads)Local-loss
80.96740.96910.97250.9701
160.9351*0.96610.96840.9685

*The residual MLP at depth 16 was under-tuned within the epoch budget and is not a load-bearing baseline.

At depth 16, deeply-supervised (0.9684) ≈ local-loss (0.9685): keeping per-layer supervision while restoring the global gradient reproduces the result. The benefit comes from per-layer supervision, which is the core ingredient.

CIFAR-10 / CNN (mean test accuracy ± std over 3 seeds)

15k train / 10k test, 3 seeds, 12 epochs, no augmentation (held identical across methods):

Depth (blocks)E2E backpropGreedy local-lossDeeply-supervised
30.528 ±.0160.567 ±.0030.520 ±.022
60.643 ±.0070.649 ±.0040.577 ±.018
90.557 ±.022 ↓0.626 ±.0050.609 ±.013

On CIFAR the plain (non-residual) end-to-end CNN degrades at depth 9 (0.643 to 0.557). Both per-layer-supervised methods are more depth-robust, and here local (0.626) edges out deepsup (0.609) at depth 9, so locality contributes a small, secondary robustness on CNNs that pure deep supervision does not fully capture. The primary mechanism is still per-layer supervision.


Methods compared

All methods share one framework, architecture, and data pipeline to avoid infrastructure confounds.

MethodDescription
End-to-end (vanilla) backpropNaive init, saturating (tanh) activation, plain SGD. This is the regime where vanishing gradients bite.
Modern backpropAdam + He init + BatchNorm. The strong baseline.
Greedy local-loss (layer-wise)The greedy bipropagation scaffold, but each layer is trained with a temporary softmax head and cross-entropy (à la Belilovsky 2019 / Nøkland 2019); the head is discarded before the next layer.
Deeply-supervised controlPer-layer auxiliary classifier heads with a single global gradient (Lee 2015). The locality-isolation control: same per-layer supervision, but global backprop is retained.
Anchors (Ploj-style)Best-effort reconstruction of Ploj's hand-designed intermediate-target rule: each layer shifts its input toward per-class anchors/prototypes, weights initialized near identity, final softmax readout.
Deterministic centroid-initOne hidden layer constructed analytically from class-centroid geometry (sparse ±1 units over the 3 most-discriminative features), targets = 0.99·layer_output + 0.01·two_hot(class), refined with Adam.

Repository structure

.
├── README.md                       # this file
├── PAPER.md                        # the English paper (authoritative findings & numbers)
├── LICENSE                         # MIT
├── requirements.txt
├── .gitignore
├── experiments/
│   ├── cifar_experiment.py         # CIFAR-10 / CNN decomposition (e2e, local, deepsup)
│   └── mnist_mlp_experiment.py     # MNIST / MLP, all methods (self-contained)
└── archive/
    ├── README.md
    └── ...                         # raw development fragments, kept for provenance

How to run

The experiments are self-contained TensorFlow 2 / Keras scripts that each run in a single Colab cell or locally.

Local

pip install -r requirements.txt
python experiments/cifar_experiment.py        # CIFAR-10 / CNN
python experiments/mnist_mlp_experiment.py    # MNIST / MLP

Colab

Upload a script (or paste it into a cell) and run. A GPU runtime (e.g. T4) is recommended for the full configs.

Notes

  • FAST_MODE flag. Each script has a FAST_MODE toggle near the top. True gives a small, fast indicative smoke run (subset of data, few epochs, 2 to 3 seeds); set it to False for the full benchmark reported in the paper.
  • CIFAR-10 download. cifar_experiment.py downloads CIFAR-10 from cs.toronto.edu via tf.keras.datasets on first run and caches it to disk; subsequent runs reuse the cache.
  • All reported numbers come from actual evaluation runs.

Limitations

  • Seeds. Most decisive numbers (full-scale and control runs) are single-seed (seed 0); the multi-seed evidence is currently FAST_MODE / CIFAR only. A fuller protocol (≥10 seeds, 95% CIs, paired Holm-Bonferroni tests) is left for follow-up.
  • Plain, non-residual baselines. Both testbeds compare against plain baselines that degrade with depth for known optimization reasons. A residual/normalized end-to-end baseline would likely close the depth gap, so the depth-robustness claims are relative to plain architectures, not modern residual networks.
  • Iso-compute. Local-loss sees the data roughly 3-6x more often than a single end-to-end run; a clean accuracy-vs-wall-clock and iso-gradient-step accounting is still outstanding.
  • Reconstruction of Ploj's rule. The "anchors" method reconstructs an unpublished multi-class target rule (the original MNIST.m is auth-walled on ResearchGate). A more faithful target scheme could raise the anchors numbers, though it would not change the per-layer-supervision-is-the-key-ingredient conclusion.

Credit

The bipropagation method and the underlying intuition, that per-layer supervision can help train deep networks, originate with Dr. Bojan Ploj. This repository builds on his work, decomposing it to understand why per-layer supervision is so effective; the credit for the original idea is his. We thank him for making the method and code public, which is what made this study possible.

Dr. Ploj's repositories:

Related foundational work this study builds on includes Deeply-Supervised Nets (Lee et al. 2015), greedy layer-wise learning at scale (Belilovsky et al. 2019), local error signals (Nøkland & Eidnes 2019), and Difference Target Propagation (Lee et al. 2015). See PAPER.md for the full reference list.


Contributing / further research

Contributions and extensions are warmly welcome. This is intended as an open, constructive starting point. Particularly valuable directions:

  • Residual / normalized end-to-end baselines. How does the depth-robustness gap look against a properly modern baseline?
  • More seeds + confidence intervals. ≥10 seeds, 95% CIs, paired Holm-Bonferroni tests on the full-scale and control tables.
  • Harder data. CIFAR-100, Tiny-ImageNet.
  • Other local-learning methods. Forward-Forward, Difference Target Propagation, synthetic gradients, feedback alignment, as additional points of comparison.
  • A more faithful reconstruction of Ploj's multi-class intermediate-target rule (ideally from the original MNIST.m).
  • Iso-compute accounting. Accuracy vs. wall-clock and vs. gradient steps.

Open an issue or a pull request.


Citation

@misc{korent2026bipropagation,
  author       = {Korent, Maj},
  title        = {Bipropagation: A Decomposition Study Building on Dr. Bojan Ploj's Idea},
  year         = {2026},
  howpublished = {\url{https://github.com/korentmaj/bipropagation-study}},
  note         = {Decomposition study building on Bojan Ploj's bipropagation method.}
}

This study is offered in a spirit of constructive collaboration. The aim is to understand and build on the genuine insight at the core of the method, and to credit it generously.

backpropagation
bipropagation
deep-learning
layer-wise-training
machine-learning
neural-networks
reproducibility
tensorflow