Biologically-constrained JEPA for disease progression - predicts who will decline, not just labels. Real ADNI data (2,347 patients), patient-level inference, honest evaluation incl. LLM head-to-head.
Python
1
3 commits
updated Sep 28, 2026
Disease progression modeling with a biologically-constrained JEPA — from recognition to reasoning.
Medical AI today labels a scan; it does not reason about how the disease will move. This project implements and tests the idea that a JEPA (Joint-Embedding Predictive Architecture) trained to predict the latent future state of a patient — and penalized whenever its predicted trajectory violates known biology — produces disease trajectories that are more biologically plausible than standard pattern-matching models.
Evaluated on real longitudinal data: an ADNI sample (2,347 participants, ~9,300 training trajectories, baseline amyloid/tau/hippocampal markers, longitudinal MMSE/ADAS13/CDR-SB, conversion labels) and OASIS-2 (150 subjects, 373 visits), with a synthetic Alzheimer cascade used as ground truth for rule verification.

Why latent prediction: raw-value prediction forces the model to also model measurement noise, scanner variation, and missing-entry artifacts — the "messy and often missing" reality of medical data. Predicting in representation space lets unpredictable detail dissolve and keeps the disease trajectory as the learning target. The biological penalty then shapes what trajectories the model is allowed to imagine: an effect requires a cause.
| JEPA-v2 | JEPA-v2+Bio | GRU | HGB | Ridge | Carry-fwd | |
|---|---|---|---|---|---|---|
| MMSE MAE @12 mo ↓ | 1.530 | 1.517 (λ2: 1.518) | 1.577 | 1.545 | 1.599 | 1.696 |
| CDR-SB MAE @36 mo ↓ | 1.222 | 1.196 (λ2: 1.206) | 1.201 | 1.202 | 1.298 | 1.408 |
| bio-penalty ablation (patient-level permutation) | — | ADAS13 p=0.0025 ✓ (survives Bonferroni) · CDRSB p=0.0001 ✓ · MMSE p=0.34 | ||||
| MAE at 75% input degradation ↓ | 3.99 | 4.08 | 3.98 | 4.43 | 4.12 | — |
Conversion AUC — two different measurements, never mixed:
Three honest statements:
Full tables, significance details, and limitations: docs/RESULTS.md.
The rule engine is verified for implementation consistency against a synthetic Alzheimer cascade with known ground truth (11/11 tests, including two regression tests added after external review: rules constrain the prediction — not gated by future ground truth — and R3 operates on unique patients with observed cells only). Compliant trajectories incur ≈ 0 penalty; unsupported decline, pathology reversal, and anti-causal ordering are flagged with >10× margin. These tests validate the code against its own pre-specified rules; they do not validate the rules as clinical causal truth. See docs/BIOLOGICAL_RULES.md.
git clone https://github.com/0ans/biological-jepa.git && cd biological-jepa
make setup # venv + torch/pandas/sklearn
make data # downloads & verifies both real datasets (checksums printed)
make test # 11 unit tests incl. rule-engine vs ground truth
make experiments # 8 models × 15 runs × 2 studies + hard tests (~45 min on a laptop)
Try it on a patient (trains in ~40 s, then predicts the 3-year trajectory):
python scripts/predict_patient.py # 3 held-out patients
python scripts/predict_patient.py --sid ADNI_77 # a specific held-out patient
Results land in experiments/{adni,oasis}/: results.json (15-run CV +
significance), hard_tests.json (missingness stress + rollout), and figures.
See experiments/README.md for a map of every artifact.
src/biojepa/
├── data/ adni.py · oasis.py · synthetic.py · dataset.py (splits, K-fold, pairs, masks)
├── model/
│ ├── jepa.py v1 snapshot encoder · v2 history encoder · EMA target · VICReg · latent rollout
│ ├── bio_rules.py differentiable rule engine (R1 capacity, R2 monotonicity, R3 ordering)
│ └── baselines.py carry-forward · ridge · HGB · supervised GRU
├── pipeline.py train-fit standardizers, per-pair tensors, capacity calibration
├── train.py · evaluate.py (bootstrap · stress · rollout) · run_experiments.py
docs/ RESULTS · RESEARCH_LOG (incl. failures) · BIOLOGICAL_RULES · ADNI_ACCESS
experiments/ committed results.json + hard_tests.json + figures (evidence)
tests/ rule-engine ground-truth tests + end-to-end smoke
Participant-level data are not committed. scripts/download_data.py
fetches the ADNI sample (via the abaR R package redistribution) and the
OASIS-2 longitudinal CSV from public research mirrors, verifies them, and
prints checksums. For production-grade runs, register at
adni.loni.usc.edu (free, DUA) — the loader
targets the ADNIMERGE schema; see docs/ADNI_ACCESS.md.
See docs/RESEARCH_LOG.md for the full honest account, including everything that failed along the way.
@software{alharbi2026biologicaljepa,
author = {Al-Harbi, Anas},
title = {biological-jepa: biologically-constrained JEPA for disease progression},
year = {2026},
url = {https://github.com/0ans/biological-jepa}
}
Foundations: JEPA/LeCun et al.; V-JEPA 2 (Assran et al., 2025); LeJEPA (Balestriero & LeCun, 2025); AD biomarker cascade (Jack et al., 2010, 2013, 2016); ADNI and OASIS-2 datasets.
MIT — see LICENSE.
Python
99.4%
Biologically-constrained JEPA for disease progression - predicts who will decline, not just labels. Real ADNI data (2,347 patients), patient-level inference, honest evaluation incl. LLM head-to-head.
Python
1
3 commits
updated Sep 28, 2026
Disease progression modeling with a biologically-constrained JEPA — from recognition to reasoning.
Medical AI today labels a scan; it does not reason about how the disease will move. This project implements and tests the idea that a JEPA (Joint-Embedding Predictive Architecture) trained to predict the latent future state of a patient — and penalized whenever its predicted trajectory violates known biology — produces disease trajectories that are more biologically plausible than standard pattern-matching models.
Evaluated on real longitudinal data: an ADNI sample (2,347 participants, ~9,300 training trajectories, baseline amyloid/tau/hippocampal markers, longitudinal MMSE/ADAS13/CDR-SB, conversion labels) and OASIS-2 (150 subjects, 373 visits), with a synthetic Alzheimer cascade used as ground truth for rule verification.

Why latent prediction: raw-value prediction forces the model to also model measurement noise, scanner variation, and missing-entry artifacts — the "messy and often missing" reality of medical data. Predicting in representation space lets unpredictable detail dissolve and keeps the disease trajectory as the learning target. The biological penalty then shapes what trajectories the model is allowed to imagine: an effect requires a cause.
| JEPA-v2 | JEPA-v2+Bio | GRU | HGB | Ridge | Carry-fwd | |
|---|---|---|---|---|---|---|
| MMSE MAE @12 mo ↓ | 1.530 | 1.517 (λ2: 1.518) | 1.577 | 1.545 | 1.599 | 1.696 |
| CDR-SB MAE @36 mo ↓ | 1.222 | 1.196 (λ2: 1.206) | 1.201 | 1.202 | 1.298 | 1.408 |
| bio-penalty ablation (patient-level permutation) | — | ADAS13 p=0.0025 ✓ (survives Bonferroni) · CDRSB p=0.0001 ✓ · MMSE p=0.34 | ||||
| MAE at 75% input degradation ↓ | 3.99 | 4.08 | 3.98 | 4.43 | 4.12 | — |
Conversion AUC — two different measurements, never mixed:
Three honest statements:
Full tables, significance details, and limitations: docs/RESULTS.md.
The rule engine is verified for implementation consistency against a synthetic Alzheimer cascade with known ground truth (11/11 tests, including two regression tests added after external review: rules constrain the prediction — not gated by future ground truth — and R3 operates on unique patients with observed cells only). Compliant trajectories incur ≈ 0 penalty; unsupported decline, pathology reversal, and anti-causal ordering are flagged with >10× margin. These tests validate the code against its own pre-specified rules; they do not validate the rules as clinical causal truth. See docs/BIOLOGICAL_RULES.md.
git clone https://github.com/0ans/biological-jepa.git && cd biological-jepa
make setup # venv + torch/pandas/sklearn
make data # downloads & verifies both real datasets (checksums printed)
make test # 11 unit tests incl. rule-engine vs ground truth
make experiments # 8 models × 15 runs × 2 studies + hard tests (~45 min on a laptop)
Try it on a patient (trains in ~40 s, then predicts the 3-year trajectory):
python scripts/predict_patient.py # 3 held-out patients
python scripts/predict_patient.py --sid ADNI_77 # a specific held-out patient
Results land in experiments/{adni,oasis}/: results.json (15-run CV +
significance), hard_tests.json (missingness stress + rollout), and figures.
See experiments/README.md for a map of every artifact.
src/biojepa/
├── data/ adni.py · oasis.py · synthetic.py · dataset.py (splits, K-fold, pairs, masks)
├── model/
│ ├── jepa.py v1 snapshot encoder · v2 history encoder · EMA target · VICReg · latent rollout
│ ├── bio_rules.py differentiable rule engine (R1 capacity, R2 monotonicity, R3 ordering)
│ └── baselines.py carry-forward · ridge · HGB · supervised GRU
├── pipeline.py train-fit standardizers, per-pair tensors, capacity calibration
├── train.py · evaluate.py (bootstrap · stress · rollout) · run_experiments.py
docs/ RESULTS · RESEARCH_LOG (incl. failures) · BIOLOGICAL_RULES · ADNI_ACCESS
experiments/ committed results.json + hard_tests.json + figures (evidence)
tests/ rule-engine ground-truth tests + end-to-end smoke
Participant-level data are not committed. scripts/download_data.py
fetches the ADNI sample (via the abaR R package redistribution) and the
OASIS-2 longitudinal CSV from public research mirrors, verifies them, and
prints checksums. For production-grade runs, register at
adni.loni.usc.edu (free, DUA) — the loader
targets the ADNIMERGE schema; see docs/ADNI_ACCESS.md.
See docs/RESEARCH_LOG.md for the full honest account, including everything that failed along the way.
@software{alharbi2026biologicaljepa,
author = {Al-Harbi, Anas},
title = {biological-jepa: biologically-constrained JEPA for disease progression},
year = {2026},
url = {https://github.com/0ans/biological-jepa}
}
Foundations: JEPA/LeCun et al.; V-JEPA 2 (Assran et al., 2025); LeJEPA (Balestriero & LeCun, 2025); AD biomarker cascade (Jack et al., 2010, 2013, 2016); ADNI and OASIS-2 datasets.
MIT — see LICENSE.
Python
99.4%