An honest, independent reproduction & component-level decomposition of Bojan Ploj's bipropagation (greedy layer-wise supervised training), benchmarked on MNIST/MLP and CIFAR-10/CNN against tuned modern backprop.
Python
0
7 commits
updated Jun 23, 2026
A component-level decomposition study of greedy, layer-wise, supervised neural-network training, built on Dr. Bojan Ploj's bipropagation idea. Bipropagation is a greedy, layer-wise, supervised training method that trains a deep network one layer at a time using intermediate targets.
Building on Dr. Ploj's insight, this repository decomposes the approach into its parts to understand which component makes layer-wise supervised training effective, and it benchmarks the approach against carefully tuned modern backpropagation baselines on MNIST and CIFAR-10.
The thesis. Building on Dr. Bojan Ploj's bipropagation idea, we decompose what makes greedy layer-wise supervised training effective. The key ingredient is per-layer supervision, and the approach is competitive and depth-robust on MNIST and CIFAR-10.
To isolate the operative mechanism, we run a deeply-supervised control: per-layer auxiliary heads driven by a single global gradient. This control matches the local-loss method at every depth (depth-16: 0.9684 vs 0.9685), which clarifies the mechanism constructively. The benefit flows from per-layer supervision itself, with a small, secondary locality effect appearing on CIFAR-10/CNN.
In short: per-layer supervision, the idea at the heart of Ploj's method, is what makes layer-wise training work, and it carries over robustly as networks get deeper.
Full-scale run, 30k train / 10k test, seed 0:
| Depth | Vanilla BP | Modern BP | Anchors (Ploj-style) | Local-loss |
|---|---|---|---|---|
| 2 | 0.9587 | 0.9690 | 0.8774 | 0.9708 |
| 4 | 0.9630 | 0.9735 | 0.8743 | 0.9714 |
| 8 | 0.9586 | 0.9627 | 0.8653 | 0.9701 |
| 16 | 0.1135 (collapse) | 0.9513 | 0.8415 | 0.9685 |
The locality-isolation control (30k MNIST, seed 0, 30 epochs) shows that deeply-supervised ≈ local-loss at every depth:
| Depth | Residual BP | Plain BP (ReLU+BN, 30ep) | Deeply-supervised (global grad + aux heads) | Local-loss |
|---|---|---|---|---|
| 8 | 0.9674 | 0.9691 | 0.9725 | 0.9701 |
| 16 | 0.9351* | 0.9661 | 0.9684 | 0.9685 |
*The residual MLP at depth 16 was under-tuned within the epoch budget and is not a load-bearing baseline.
At depth 16, deeply-supervised (0.9684) ≈ local-loss (0.9685): keeping per-layer supervision while restoring the global gradient reproduces the result. The benefit comes from per-layer supervision, which is the core ingredient.
15k train / 10k test, 3 seeds, 12 epochs, no augmentation (held identical across methods):
| Depth (blocks) | E2E backprop | Greedy local-loss | Deeply-supervised |
|---|---|---|---|
| 3 | 0.528 ±.016 | 0.567 ±.003 | 0.520 ±.022 |
| 6 | 0.643 ±.007 | 0.649 ±.004 | 0.577 ±.018 |
| 9 | 0.557 ±.022 ↓ | 0.626 ±.005 | 0.609 ±.013 |
On CIFAR the plain (non-residual) end-to-end CNN degrades at depth 9 (0.643 to 0.557). Both per-layer-supervised methods are more depth-robust, and here local (0.626) edges out deepsup (0.609) at depth 9, so locality contributes a small, secondary robustness on CNNs that pure deep supervision does not fully capture. The primary mechanism is still per-layer supervision.
All methods share one framework, architecture, and data pipeline to avoid infrastructure confounds.
| Method | Description |
|---|---|
| End-to-end (vanilla) backprop | Naive init, saturating (tanh) activation, plain SGD. This is the regime where vanishing gradients bite. |
| Modern backprop | Adam + He init + BatchNorm. The strong baseline. |
| Greedy local-loss (layer-wise) | The greedy bipropagation scaffold, but each layer is trained with a temporary softmax head and cross-entropy (à la Belilovsky 2019 / Nøkland 2019); the head is discarded before the next layer. |
| Deeply-supervised control | Per-layer auxiliary classifier heads with a single global gradient (Lee 2015). The locality-isolation control: same per-layer supervision, but global backprop is retained. |
| Anchors (Ploj-style) | Best-effort reconstruction of Ploj's hand-designed intermediate-target rule: each layer shifts its input toward per-class anchors/prototypes, weights initialized near identity, final softmax readout. |
| Deterministic centroid-init | One hidden layer constructed analytically from class-centroid geometry (sparse ±1 units over the 3 most-discriminative features), targets = 0.99·layer_output + 0.01·two_hot(class), refined with Adam. |
.
├── README.md # this file
├── PAPER.md # the English paper (authoritative findings & numbers)
├── LICENSE # MIT
├── requirements.txt
├── .gitignore
├── experiments/
│ ├── cifar_experiment.py # CIFAR-10 / CNN decomposition (e2e, local, deepsup)
│ └── mnist_mlp_experiment.py # MNIST / MLP, all methods (self-contained)
└── archive/
├── README.md
└── ... # raw development fragments, kept for provenance
The experiments are self-contained TensorFlow 2 / Keras scripts that each run in a single Colab cell or locally.
pip install -r requirements.txt
python experiments/cifar_experiment.py # CIFAR-10 / CNN
python experiments/mnist_mlp_experiment.py # MNIST / MLP
Upload a script (or paste it into a cell) and run. A GPU runtime (e.g. T4) is recommended for the full configs.
FAST_MODE flag. Each script has a FAST_MODE toggle near the top. True gives a small, fast indicative smoke run (subset of data, few epochs, 2 to 3 seeds); set it to False for the full benchmark reported in the paper.cifar_experiment.py downloads CIFAR-10 from cs.toronto.edu via tf.keras.datasets on first run and caches it to disk; subsequent runs reuse the cache.MNIST.m is auth-walled on ResearchGate). A more faithful target scheme could raise the anchors numbers, though it would not change the per-layer-supervision-is-the-key-ingredient conclusion.The bipropagation method and the underlying intuition, that per-layer supervision can help train deep networks, originate with Dr. Bojan Ploj. This repository builds on his work, decomposing it to understand why per-layer supervision is so effective; the credit for the original idea is his. We thank him for making the method and code public, which is what made this study possible.
Dr. Ploj's repositories:
Related foundational work this study builds on includes Deeply-Supervised Nets (Lee et al. 2015), greedy layer-wise learning at scale (Belilovsky et al. 2019), local error signals (Nøkland & Eidnes 2019), and Difference Target Propagation (Lee et al. 2015). See PAPER.md for the full reference list.
Contributions and extensions are warmly welcome. This is intended as an open, constructive starting point. Particularly valuable directions:
MNIST.m).Open an issue or a pull request.
@misc{korent2026bipropagation,
author = {Korent, Maj},
title = {Bipropagation: A Decomposition Study Building on Dr. Bojan Ploj's Idea},
year = {2026},
howpublished = {\url{https://github.com/korentmaj/bipropagation-study}},
note = {Decomposition study building on Bojan Ploj's bipropagation method.}
}
This study is offered in a spirit of constructive collaboration. The aim is to understand and build on the genuine insight at the core of the method, and to credit it generously.
An honest, independent reproduction & component-level decomposition of Bojan Ploj's bipropagation (greedy layer-wise supervised training), benchmarked on MNIST/MLP and CIFAR-10/CNN against tuned modern backprop.
Python
0
7 commits
updated Jun 23, 2026
A component-level decomposition study of greedy, layer-wise, supervised neural-network training, built on Dr. Bojan Ploj's bipropagation idea. Bipropagation is a greedy, layer-wise, supervised training method that trains a deep network one layer at a time using intermediate targets.
Building on Dr. Ploj's insight, this repository decomposes the approach into its parts to understand which component makes layer-wise supervised training effective, and it benchmarks the approach against carefully tuned modern backpropagation baselines on MNIST and CIFAR-10.
The thesis. Building on Dr. Bojan Ploj's bipropagation idea, we decompose what makes greedy layer-wise supervised training effective. The key ingredient is per-layer supervision, and the approach is competitive and depth-robust on MNIST and CIFAR-10.
To isolate the operative mechanism, we run a deeply-supervised control: per-layer auxiliary heads driven by a single global gradient. This control matches the local-loss method at every depth (depth-16: 0.9684 vs 0.9685), which clarifies the mechanism constructively. The benefit flows from per-layer supervision itself, with a small, secondary locality effect appearing on CIFAR-10/CNN.
In short: per-layer supervision, the idea at the heart of Ploj's method, is what makes layer-wise training work, and it carries over robustly as networks get deeper.
Full-scale run, 30k train / 10k test, seed 0:
| Depth | Vanilla BP | Modern BP | Anchors (Ploj-style) | Local-loss |
|---|---|---|---|---|
| 2 | 0.9587 | 0.9690 | 0.8774 | 0.9708 |
| 4 | 0.9630 | 0.9735 | 0.8743 | 0.9714 |
| 8 | 0.9586 | 0.9627 | 0.8653 | 0.9701 |
| 16 | 0.1135 (collapse) | 0.9513 | 0.8415 | 0.9685 |
The locality-isolation control (30k MNIST, seed 0, 30 epochs) shows that deeply-supervised ≈ local-loss at every depth:
| Depth | Residual BP | Plain BP (ReLU+BN, 30ep) | Deeply-supervised (global grad + aux heads) | Local-loss |
|---|---|---|---|---|
| 8 | 0.9674 | 0.9691 | 0.9725 | 0.9701 |
| 16 | 0.9351* | 0.9661 | 0.9684 | 0.9685 |
*The residual MLP at depth 16 was under-tuned within the epoch budget and is not a load-bearing baseline.
At depth 16, deeply-supervised (0.9684) ≈ local-loss (0.9685): keeping per-layer supervision while restoring the global gradient reproduces the result. The benefit comes from per-layer supervision, which is the core ingredient.
15k train / 10k test, 3 seeds, 12 epochs, no augmentation (held identical across methods):
| Depth (blocks) | E2E backprop | Greedy local-loss | Deeply-supervised |
|---|---|---|---|
| 3 | 0.528 ±.016 | 0.567 ±.003 | 0.520 ±.022 |
| 6 | 0.643 ±.007 | 0.649 ±.004 | 0.577 ±.018 |
| 9 | 0.557 ±.022 ↓ | 0.626 ±.005 | 0.609 ±.013 |
On CIFAR the plain (non-residual) end-to-end CNN degrades at depth 9 (0.643 to 0.557). Both per-layer-supervised methods are more depth-robust, and here local (0.626) edges out deepsup (0.609) at depth 9, so locality contributes a small, secondary robustness on CNNs that pure deep supervision does not fully capture. The primary mechanism is still per-layer supervision.
All methods share one framework, architecture, and data pipeline to avoid infrastructure confounds.
| Method | Description |
|---|---|
| End-to-end (vanilla) backprop | Naive init, saturating (tanh) activation, plain SGD. This is the regime where vanishing gradients bite. |
| Modern backprop | Adam + He init + BatchNorm. The strong baseline. |
| Greedy local-loss (layer-wise) | The greedy bipropagation scaffold, but each layer is trained with a temporary softmax head and cross-entropy (à la Belilovsky 2019 / Nøkland 2019); the head is discarded before the next layer. |
| Deeply-supervised control | Per-layer auxiliary classifier heads with a single global gradient (Lee 2015). The locality-isolation control: same per-layer supervision, but global backprop is retained. |
| Anchors (Ploj-style) | Best-effort reconstruction of Ploj's hand-designed intermediate-target rule: each layer shifts its input toward per-class anchors/prototypes, weights initialized near identity, final softmax readout. |
| Deterministic centroid-init | One hidden layer constructed analytically from class-centroid geometry (sparse ±1 units over the 3 most-discriminative features), targets = 0.99·layer_output + 0.01·two_hot(class), refined with Adam. |
.
├── README.md # this file
├── PAPER.md # the English paper (authoritative findings & numbers)
├── LICENSE # MIT
├── requirements.txt
├── .gitignore
├── experiments/
│ ├── cifar_experiment.py # CIFAR-10 / CNN decomposition (e2e, local, deepsup)
│ └── mnist_mlp_experiment.py # MNIST / MLP, all methods (self-contained)
└── archive/
├── README.md
└── ... # raw development fragments, kept for provenance
The experiments are self-contained TensorFlow 2 / Keras scripts that each run in a single Colab cell or locally.
pip install -r requirements.txt
python experiments/cifar_experiment.py # CIFAR-10 / CNN
python experiments/mnist_mlp_experiment.py # MNIST / MLP
Upload a script (or paste it into a cell) and run. A GPU runtime (e.g. T4) is recommended for the full configs.
FAST_MODE flag. Each script has a FAST_MODE toggle near the top. True gives a small, fast indicative smoke run (subset of data, few epochs, 2 to 3 seeds); set it to False for the full benchmark reported in the paper.cifar_experiment.py downloads CIFAR-10 from cs.toronto.edu via tf.keras.datasets on first run and caches it to disk; subsequent runs reuse the cache.MNIST.m is auth-walled on ResearchGate). A more faithful target scheme could raise the anchors numbers, though it would not change the per-layer-supervision-is-the-key-ingredient conclusion.The bipropagation method and the underlying intuition, that per-layer supervision can help train deep networks, originate with Dr. Bojan Ploj. This repository builds on his work, decomposing it to understand why per-layer supervision is so effective; the credit for the original idea is his. We thank him for making the method and code public, which is what made this study possible.
Dr. Ploj's repositories:
Related foundational work this study builds on includes Deeply-Supervised Nets (Lee et al. 2015), greedy layer-wise learning at scale (Belilovsky et al. 2019), local error signals (Nøkland & Eidnes 2019), and Difference Target Propagation (Lee et al. 2015). See PAPER.md for the full reference list.
Contributions and extensions are warmly welcome. This is intended as an open, constructive starting point. Particularly valuable directions:
MNIST.m).Open an issue or a pull request.
@misc{korent2026bipropagation,
author = {Korent, Maj},
title = {Bipropagation: A Decomposition Study Building on Dr. Bojan Ploj's Idea},
year = {2026},
howpublished = {\url{https://github.com/korentmaj/bipropagation-study}},
note = {Decomposition study building on Bojan Ploj's bipropagation method.}
}
This study is offered in a spirit of constructive collaboration. The aim is to understand and build on the genuine insight at the core of the method, and to credit it generously.