dsadawq3/DynamicTune

DynamicTune: Cross-model hidden trajectory transport and direct MLP weight surgery across dimensions (Qwen3.5 4B -> 0.8B on 8GB RX 580)

Python

0

6 commits

updated Oct 2, 2026

See the code

See what people are saying

README

DynamicTune

Challenging the trillion-token orthodoxy: cross-model hidden trajectory transport and closed-form weight surgery across architectures and model widths.

Tested across radically different model families: modern hybrid Qwen3.5 (4B with d=2560 -> 0.8B with d=1024) and notoriously fragile GPT-2 (XL with d=1600 -> small with d=768).


Breaking the Trillion-Token Orthodoxy

The prevailing consensus in deep learning is that transferring capabilities from a larger teacher model to a smaller student demands billions or trillions of tokens, massive synthetic dataset pipelines, and weeks of GPU cluster compute running token-level cross-entropy or KL divergence minimization.

DynamicTune challenges this dogma:

  1. A transformer stack is fundamentally a discrete dynamical system over depth: $h_{l+1} = h_l + f_l(h_l)$.
  2. A capable teacher traces an informational velocity field through representation space.
  3. By aligning these trajectories through a local orthogonal Procrustes atlas and solving for closed-form weight updates in the student's MLP blocks, we can physically transfer teacher trajectory dynamics into the student without backpropagation or training runs.
  4. Cross-architecture stability: We tested this on both modern Qwen3.5 (SwiGLU, hybrid linear/sliding attention) and the notoriously fragile GPT-2 small (where Conv1D layers famously collapse or degrade into gibberish at the slightest weight disturbance). In both architectures, baseline language modeling integrity is preserved with well-behaved, bounded degradation margins, while target domain accuracy improves.
  5. The spectral entropy discovery: Editing all 24 student layers destroys the model (+64.78% NLL) because intermediate layers (1-22) are high-entropy polysemantic knots (>0.90 spectral entropy). Restricting the surgery to 4 anchor blocks (layers 0, 7, 15, and 23) avoids destructive interference, cuts multi-domain held-out NLL by -10.8%, and boosts HellaSwag (+0.50% across 400 tasks) on native Vulkan llama.cpp.

Raw reproducible benchmark logs: benchmarks/.


Theoretical Foundations: Where This Fits in the Literature

DynamicTune bridges established machine learning theory and mechanistic interpretability into an empirical weight surgery engine:

1. Residual Networks as Dynamical Systems and Neural ODEs

Residual connections allow layers to be viewed as Euler discretization steps of an underlying continuous ordinary differential equation $\frac{dh}{dt} = f(h(t), t)$.

In DynamicTune, we do not view weights as static feature matrices. We treat the step $\Delta h_l = h_{l+1} - h_l$ as the velocity of a dynamical trajectory through hidden space. If a 4B teacher takes a more direct path toward the solution manifold than a 0.8B student, the teacher's velocity field carries transferable corrective force.

2. The Linear Representation Hypothesis and Procrustes Alignment

Concepts in large language models are represented as linear directions in representation space, and different models often learn linearly or orthogonally equivalent geometries up to rotation and scaling.

Because teacher and student models have different hidden dimensions (e.g. 2560 vs 1024), a single global orthogonal matrix cannot capture non-linear curvature across different semantic clusters. DynamicTune builds a piecewise local Procrustes atlas (ManifoldChartAtlas in faytuna_flow/manifold_charts.py): we cluster hidden states with K-Means into $K$ local charts and fit temperature-weighted local rotations $P_k \in \mathbb{R}^{d_T \times d_S}$.

3. Superposition, Polysemanticity, and the Spectral Entropy Barrier

Why did past attempts at layer-wise weight transfer fail? Anthropic's research into mechanistic interpretability showed that neural networks pack more features than they have dimensions via superposition, creating polysemantic neurons that activate on multiple unrelated concepts.

When a student model has only 1024 dimensions, intermediate layers (layers 1 to 22) are forced to operate in dense superposition. In DynamicTune, we compute the singular value distribution of the flow residual and calculate its normalized Shannon spectral entropy $H \in [0, 1]$ (faytuna_flow/knots.py).

  • Layer 0 exhibits low entropy ($H = 0.7138$): clean, coherent semantic grounding.
  • Layers 1-22 exhibit high entropy ($H \in [0.8957, 0.9669]$): chaotic superposition knots. Forcing a linear weight update here creates catastrophic interference and ruins the model.
  • Layer 23 ($H = 0.9310$): pre-unembed boundary where features unpack toward vocabulary logits.

By discovering this entropy barrier, we learned that weight surgery must respect superposition boundaries: edit sparse anchor points, bypass the knots.

4. Direct Closed-Form Model Editing

Instead of gradient descent, direct weight updates can be computed as closed-form linear projections that satisfy key-value associations.

DynamicTune extends this concept from individual fact-editing to depth-wise dynamical flow transport: we pull the projected trajectory deltas back through the SwiGLU MLP blocks via a damped Tikhonov pseudoinverse and rank-constrained SVD projections with explicit spectral trust-region bounds.


24-Layer Spectral Entropy Scan

Running scripts/scan_24_layers_autogate.py across all layers of Qwen3.5-0.8B mapped to Qwen3.5-4B reveals why full-model transfer fails:

Student LayerMapped Teacher LayerMulti-Chart Atlas Spectral Entropy $H$Diagnosis
Layer 0Layer 00.7138Coherent semantic anchor (safe for surgery)
Layer 1Layer 10.9669Polysemantic knot (skip)
Layer 2Layer 30.9419Polysemantic knot (skip)
Layer 3Layer 40.9032Polysemantic knot (skip)
Layer 4Layer 50.9572Polysemantic knot (skip)
Layer 5Layer 70.9250Polysemantic knot (skip)
Layer 6Layer 80.9388Polysemantic knot (skip)
Layer 7Layer 90.9384Intermediate bridge anchor (damped)
Layer 8Layer 110.9387Polysemantic knot (skip)
Layer 9Layer 120.9531Polysemantic knot (skip)
Layer 10Layer 130.9182Polysemantic knot (skip)
Layer 11Layer 150.9482Polysemantic knot (skip)
Layer 12Layer 160.9504Polysemantic knot (skip)
Layer 13Layer 170.9183Polysemantic knot (skip)
Layer 14Layer 190.9169Polysemantic knot (skip)
Layer 15Layer 200.9158Intermediate bridge anchor (damped)
Layer 16Layer 210.8957Polysemantic knot (skip)
Layer 17Layer 230.9458Polysemantic knot (skip)
Layer 18Layer 240.9033Polysemantic knot (skip)
Layer 19Layer 250.9298Polysemantic knot (skip)
Layer 20Layer 270.9372Polysemantic knot (skip)
Layer 21Layer 280.9375Polysemantic knot (skip)
Layer 22Layer 300.9410Polysemantic knot (skip)
Layer 23Layer 310.9310Pre-head output boundary (low-alpha anchor)
  • All 24 layers edited: Perplexity explodes from 17.34 to 76.59 (+64.78% NLL).
  • 4-block anchor surgery ([0, 7, 15, 23]): Model remains stable, baseline integrity is preserved, and held-out benchmarks improve.

Empirical Benchmark Results

Evaluated on exported GGUF models (qwen35_0.8b_base_f16.gguf vs qwen35_0.8b_transferred_f16.gguf) using stock Vulkan llama.cpp tools.

1. HellaSwag: 400 Tasks (Vulkan llama-perplexity --hellaswag, Seed 42)

CheckpointBase 0.8B (acc_norm)Transferred 0.8B (acc_norm)Delta
50 tasks54.00%56.00%+2.00%
100 tasks50.00%51.00%+1.00%
150 tasks54.00%54.67%+0.67%
200 tasks53.50%54.50%+1.00%
250 tasks53.20%54.00%+0.80%
300 tasks54.67%55.67%+1.00%
350 tasks53.43%54.29%+0.86%
400 tasks (Final)54.75%55.25%+0.50%

See benchmarks/hellaswag_benchmark_report.json.

2. Multi-Domain Perplexity Audit (30 Held-Out Tasks via llama.cpp)

Domain (6 tasks each)Base NLLTransferred NLLNLL Delta (%)
Biomedicine & Nature0.6390.487-23.8%
Mathematics & Logic0.6560.560-14.6%
Python Algorithms0.3400.314-7.6%
Deep Learning Architecture0.7230.698-3.5%
Russian Reasoning & Nuance0.5500.538-2.2%
Overall Average (30 tasks)0.5820.519-10.8%

See benchmarks/llama_cpp_hardcore_benchmark_report.json.

3. Strictly Masked Target-Only QA Cross-Entropy (25 Pairs)

Evaluated with prompt tokens masked to -100, measuring loss strictly on target answer tokens:

  • Science & Medicine: -16.76% NLL (2.4730 -> 2.0585)
  • Russian QA: -6.99% NLL (2.5701 -> 2.3905)
  • Logic & Math: -6.18% NLL (2.3736 -> 2.2268)
  • History & Geography: -5.27% NLL (2.0630 -> 1.9543)

See benchmarks/strict_qa_report.json.


Qualitative Generation Differences (Greedy, Seed 42)

None of these tasks appeared in the 8 calibration prompts.

1. Binary Tree Inversion (algo_invert_binary_tree)

Prompt: Question: Write a clean Python function invert_tree(root) that recursively inverts a binary tree node with .left and .right pointers and returns the root.\nAnswer:

  • Base 0.8B (NLL: 0.2347): Outputs commented-out dead code:
# def invert_tree(root):
#     if root is None:
#         return None
#     root.left, root.right = root.right, root.left
  • Transferred 0.8B (NLL: 0.1057, -55.0%): Outputs valid, executable Python with recursive traversal:
def invert_tree(root):
    if root is None:
        return None
    root.left, root.right = root.right, root.left
    invert_tree(root.left)
    invert_tree(root.right)
    return root
if __name__ == "__main__":

2. Russian Knights & Knaves Paradox (ru_knights_knaves_paradox)

Prompt: Question: На острове живут рыцари (всегда говорят правду) и лжецы (всегда лгут). Житель А говорит: «Я лжец». Кто житель А?\nAnswer:

  • Base 0.8B (NLL: 0.5535): Immediately hallucinates a single wrong sentence without reasoning: Житель А - лжец.
  • Transferred 0.8B (NLL: 0.2862, -48.3%): Spontaneously enters a structured <think> reasoning chain:
<think>
Мы рассматриваем ситуацию с двумя типами людей: рыцари (которые всегда правдивы) и лжецы (которые всегда лгут).
Житель А говорит: "Я лжец". Нужно определить, кем является житель А. Рассмотрим возможные варианты:
1. Если А - рыцарь...

3. CPython Global Interpreter Lock (algo_gil_python)

Prompt: Question: Why does Python's standard CPython runtime use a Global Interpreter Lock (GIL)?\nAnswer:

  • Base 0.8B: Gives generic filler ("prevents multiple threads from executing Python bytecodes simultaneously, which can lead to inefficiencies").
  • Transferred 0.8B: Identifies the exact low-level systems reason ("prevents multiple threads from accessing the same memory locations simultaneously, which could lead to race conditions and data corruption. The GIL is essential for maintaining thread safety").

Math & Solver Highlights

  1. Damped Tikhonov Inversion (faytuna_flow/nonlinear_transfer.py): Standard pseudoinverse np.linalg.pinv(w_down.T) blows up on near-zero singular values, forcing trust-region gates to crush updates down to 0.0009. We use Levenberg-Marquardt damping: $$\sigma_i^+ = \frac{\sigma_i}{\sigma_i^2 + \lambda}, \quad \lambda = \text{ridge} \cdot \frac{|A|_F^2}{\min(M, N)}$$

  2. Adaptive Spectral Rank Selection: Dynamically determines SVD truncation rank $r \in [16, 128]$ such that $\sum_{i=1}^r \sigma_i^2 / \sum \sigma_i^2 \ge 0.85$, retaining 85% of variance instead of fixed rank-16 truncation.

  3. Spectral Directional Rescale: Enforces $|\Delta W|2 \le 0.05 \cdot |W{\text{orig}}|_2$, ensuring that updates cannot alter the principal spectral direction of the original weight matrix.

  4. Null-Space Projector for Tied-Embedding Instruct Models (scripts/run_instruct_flow_transfer.py): For Instruct models where Layer 23 maps directly into tied token embeddings, we construct a semantic null-space projector: $$P_\perp = I - V_{16} V_{16}^T$$ derived from the top eigenvectors of the token embedding Gram matrix $W_{\text{embed}}^T W_{\text{embed}}$, protecting RLHF and vocabulary margins while computing the closed-form scale $\alpha^* = \frac{\langle \hat{\Delta}, \Delta Y \rangle_F}{|\hat{\Delta}|_F^2}$ strictly on assistant response tokens.


Hardware Accessibility: Consumer GPU Layer-Streaming

You do not need a cluster of H100s to perform trajectory transfer. To prove that direct weight surgery is accessible on commodity local hardware, all experiments were run on a single consumer 8GB AMD Radeon RX 580 using Layer-Outer VRAM Streaming (faytuna_flow/layer_streaming.py):

  1. Since we only need forward trajectories over a small calibration batch (8 to 32 prompts), we load embed_tokens and Layer 0 (~250 MB for 4B) into GPU VRAM via DirectML (torch-directml).
  2. All calibration prompts pass through Layer 0 in one batch.
  3. The resulting hidden states $H_1$ are saved to system RAM, Layer 0 is deleted from VRAM, and Layer 1 is loaded.
  4. We repeat this across all 32 layers.

Zero-Logit speedup: Calling model.model(...) directly instead of model(...) skips the final lm_head projection onto Qwen's 248,320 vocabulary tokens during trace collection, cutting extraction time by 38%.


Quickstart

1. Install dependencies

git clone https://github.com/dsadawq3/DynamicTune.git
cd DynamicTune
pip install -e .

2. Run unit tests

python -m pytest

3. Scan 24 layers for spectral entropy and knot detection

python scripts/scan_24_layers_autogate.py

4. Run 4-Block Anchor Surgery (4B -> 0.8B on GPU)

python scripts/run_qwen35_transfer.py \
  --student-dir /path/to/Qwen3.5-0.8B-Base \
  --teacher-dir /path/to/Qwen3.5-4B-Base \
  --output-dir runs/qwen35_anchor_surgery \
  --device dml \
  --use-layer-streaming \
  --export-gguf

5. Reproduce benchmarks via Vulkan llama.cpp

python scripts/run_hellaswag_audit.py
python scripts/bench_llama_cpp_real.py

Community, High-Priority Model Pairs, and Open Issues

We invite community members and researchers with modern 24GB-32GB+ GPUs (RTX 5090, RTX 4090, dual GPUs, or cloud nodes) to run trajectory surgery on frontier model pairs:

  • Next-Gen Qwen Reasoning & Flow Surgery:
    • Projecting trajectories from Qwen/Qwen3.8-27B (or Qwen3.8-Flash-Next) into Qwen/Qwen3.5-9B, 4B, or 2B.
    • Testing transfer between Mixture-of-Experts and dense backbones (Qwen3.6-35B-A3B -> Qwen3.5-4B).
  • Google Gemma-4 Cross-Scale Transfer:
    • Compressing the representation manifold of google/gemma-4-31B (or gemma-4-26B-A4B) into mobile-class edge models like google/gemma-4-E4B or gemma-4-E2B.
  • Abliteration & Agentic Trait Transplants:
    • Transferring refusal-ablation and guardrail-relaxation vectors from uncensored agentic models (e.g. Huihui-NeoHorse-1-4B-abliterated or Hermes) into restricted compact base models without destructive retraining.
  • Unknotting Superposition with Pre-Trained SAEs:
    • Can official sparse autoencoders like Qwen/SAE-Res-Qwen3.5-27B-W80K-L0_100 and Qwen/SAE-Res-Qwen3.5-9B-Base-W64K isolate monosemantic feature directions in layers 1-22, allowing trajectory transport across all layers rather than just 4 anchor blocks?

Open an Issue or Share Your Results

If you run DynamicTune on Qwen3.8, Gemma-4, or custom fine-tunes on an RTX 5090 or cloud cluster, open a GitHub Issue with your scan_24_layers_autogate.py entropy table, NLL logs, and GGUF outputs. We actively review PRs and discussion threads.


License

Apache 2.0. See LICENSE for details.

dsadawq3/DynamicTune

DynamicTune: Cross-model hidden trajectory transport and direct MLP weight surgery across dimensions (Qwen3.5 4B -> 0.8B on 8GB RX 580)

Python

0

6 commits

updated Oct 2, 2026

See the code

See what people are saying

README

DynamicTune

Challenging the trillion-token orthodoxy: cross-model hidden trajectory transport and closed-form weight surgery across architectures and model widths.

Tested across radically different model families: modern hybrid Qwen3.5 (4B with d=2560 -> 0.8B with d=1024) and notoriously fragile GPT-2 (XL with d=1600 -> small with d=768).


Breaking the Trillion-Token Orthodoxy

The prevailing consensus in deep learning is that transferring capabilities from a larger teacher model to a smaller student demands billions or trillions of tokens, massive synthetic dataset pipelines, and weeks of GPU cluster compute running token-level cross-entropy or KL divergence minimization.

DynamicTune challenges this dogma:

  1. A transformer stack is fundamentally a discrete dynamical system over depth: $h_{l+1} = h_l + f_l(h_l)$.
  2. A capable teacher traces an informational velocity field through representation space.
  3. By aligning these trajectories through a local orthogonal Procrustes atlas and solving for closed-form weight updates in the student's MLP blocks, we can physically transfer teacher trajectory dynamics into the student without backpropagation or training runs.
  4. Cross-architecture stability: We tested this on both modern Qwen3.5 (SwiGLU, hybrid linear/sliding attention) and the notoriously fragile GPT-2 small (where Conv1D layers famously collapse or degrade into gibberish at the slightest weight disturbance). In both architectures, baseline language modeling integrity is preserved with well-behaved, bounded degradation margins, while target domain accuracy improves.
  5. The spectral entropy discovery: Editing all 24 student layers destroys the model (+64.78% NLL) because intermediate layers (1-22) are high-entropy polysemantic knots (>0.90 spectral entropy). Restricting the surgery to 4 anchor blocks (layers 0, 7, 15, and 23) avoids destructive interference, cuts multi-domain held-out NLL by -10.8%, and boosts HellaSwag (+0.50% across 400 tasks) on native Vulkan llama.cpp.

Raw reproducible benchmark logs: benchmarks/.


Theoretical Foundations: Where This Fits in the Literature

DynamicTune bridges established machine learning theory and mechanistic interpretability into an empirical weight surgery engine:

1. Residual Networks as Dynamical Systems and Neural ODEs

Residual connections allow layers to be viewed as Euler discretization steps of an underlying continuous ordinary differential equation $\frac{dh}{dt} = f(h(t), t)$.

In DynamicTune, we do not view weights as static feature matrices. We treat the step $\Delta h_l = h_{l+1} - h_l$ as the velocity of a dynamical trajectory through hidden space. If a 4B teacher takes a more direct path toward the solution manifold than a 0.8B student, the teacher's velocity field carries transferable corrective force.

2. The Linear Representation Hypothesis and Procrustes Alignment

Concepts in large language models are represented as linear directions in representation space, and different models often learn linearly or orthogonally equivalent geometries up to rotation and scaling.

Because teacher and student models have different hidden dimensions (e.g. 2560 vs 1024), a single global orthogonal matrix cannot capture non-linear curvature across different semantic clusters. DynamicTune builds a piecewise local Procrustes atlas (ManifoldChartAtlas in faytuna_flow/manifold_charts.py): we cluster hidden states with K-Means into $K$ local charts and fit temperature-weighted local rotations $P_k \in \mathbb{R}^{d_T \times d_S}$.

3. Superposition, Polysemanticity, and the Spectral Entropy Barrier

Why did past attempts at layer-wise weight transfer fail? Anthropic's research into mechanistic interpretability showed that neural networks pack more features than they have dimensions via superposition, creating polysemantic neurons that activate on multiple unrelated concepts.

When a student model has only 1024 dimensions, intermediate layers (layers 1 to 22) are forced to operate in dense superposition. In DynamicTune, we compute the singular value distribution of the flow residual and calculate its normalized Shannon spectral entropy $H \in [0, 1]$ (faytuna_flow/knots.py).

  • Layer 0 exhibits low entropy ($H = 0.7138$): clean, coherent semantic grounding.
  • Layers 1-22 exhibit high entropy ($H \in [0.8957, 0.9669]$): chaotic superposition knots. Forcing a linear weight update here creates catastrophic interference and ruins the model.
  • Layer 23 ($H = 0.9310$): pre-unembed boundary where features unpack toward vocabulary logits.

By discovering this entropy barrier, we learned that weight surgery must respect superposition boundaries: edit sparse anchor points, bypass the knots.

4. Direct Closed-Form Model Editing

Instead of gradient descent, direct weight updates can be computed as closed-form linear projections that satisfy key-value associations.

DynamicTune extends this concept from individual fact-editing to depth-wise dynamical flow transport: we pull the projected trajectory deltas back through the SwiGLU MLP blocks via a damped Tikhonov pseudoinverse and rank-constrained SVD projections with explicit spectral trust-region bounds.


24-Layer Spectral Entropy Scan

Running scripts/scan_24_layers_autogate.py across all layers of Qwen3.5-0.8B mapped to Qwen3.5-4B reveals why full-model transfer fails:

Student LayerMapped Teacher LayerMulti-Chart Atlas Spectral Entropy $H$Diagnosis
Layer 0Layer 00.7138Coherent semantic anchor (safe for surgery)
Layer 1Layer 10.9669Polysemantic knot (skip)
Layer 2Layer 30.9419Polysemantic knot (skip)
Layer 3Layer 40.9032Polysemantic knot (skip)
Layer 4Layer 50.9572Polysemantic knot (skip)
Layer 5Layer 70.9250Polysemantic knot (skip)
Layer 6Layer 80.9388Polysemantic knot (skip)
Layer 7Layer 90.9384Intermediate bridge anchor (damped)
Layer 8Layer 110.9387Polysemantic knot (skip)
Layer 9Layer 120.9531Polysemantic knot (skip)
Layer 10Layer 130.9182Polysemantic knot (skip)
Layer 11Layer 150.9482Polysemantic knot (skip)
Layer 12Layer 160.9504Polysemantic knot (skip)
Layer 13Layer 170.9183Polysemantic knot (skip)
Layer 14Layer 190.9169Polysemantic knot (skip)
Layer 15Layer 200.9158Intermediate bridge anchor (damped)
Layer 16Layer 210.8957Polysemantic knot (skip)
Layer 17Layer 230.9458Polysemantic knot (skip)
Layer 18Layer 240.9033Polysemantic knot (skip)
Layer 19Layer 250.9298Polysemantic knot (skip)
Layer 20Layer 270.9372Polysemantic knot (skip)
Layer 21Layer 280.9375Polysemantic knot (skip)
Layer 22Layer 300.9410Polysemantic knot (skip)
Layer 23Layer 310.9310Pre-head output boundary (low-alpha anchor)
  • All 24 layers edited: Perplexity explodes from 17.34 to 76.59 (+64.78% NLL).
  • 4-block anchor surgery ([0, 7, 15, 23]): Model remains stable, baseline integrity is preserved, and held-out benchmarks improve.

Empirical Benchmark Results

Evaluated on exported GGUF models (qwen35_0.8b_base_f16.gguf vs qwen35_0.8b_transferred_f16.gguf) using stock Vulkan llama.cpp tools.

1. HellaSwag: 400 Tasks (Vulkan llama-perplexity --hellaswag, Seed 42)

CheckpointBase 0.8B (acc_norm)Transferred 0.8B (acc_norm)Delta
50 tasks54.00%56.00%+2.00%
100 tasks50.00%51.00%+1.00%
150 tasks54.00%54.67%+0.67%
200 tasks53.50%54.50%+1.00%
250 tasks53.20%54.00%+0.80%
300 tasks54.67%55.67%+1.00%
350 tasks53.43%54.29%+0.86%
400 tasks (Final)54.75%55.25%+0.50%

See benchmarks/hellaswag_benchmark_report.json.

2. Multi-Domain Perplexity Audit (30 Held-Out Tasks via llama.cpp)

Domain (6 tasks each)Base NLLTransferred NLLNLL Delta (%)
Biomedicine & Nature0.6390.487-23.8%
Mathematics & Logic0.6560.560-14.6%
Python Algorithms0.3400.314-7.6%
Deep Learning Architecture0.7230.698-3.5%
Russian Reasoning & Nuance0.5500.538-2.2%
Overall Average (30 tasks)0.5820.519-10.8%

See benchmarks/llama_cpp_hardcore_benchmark_report.json.

3. Strictly Masked Target-Only QA Cross-Entropy (25 Pairs)

Evaluated with prompt tokens masked to -100, measuring loss strictly on target answer tokens:

  • Science & Medicine: -16.76% NLL (2.4730 -> 2.0585)
  • Russian QA: -6.99% NLL (2.5701 -> 2.3905)
  • Logic & Math: -6.18% NLL (2.3736 -> 2.2268)
  • History & Geography: -5.27% NLL (2.0630 -> 1.9543)

See benchmarks/strict_qa_report.json.


Qualitative Generation Differences (Greedy, Seed 42)

None of these tasks appeared in the 8 calibration prompts.

1. Binary Tree Inversion (algo_invert_binary_tree)

Prompt: Question: Write a clean Python function invert_tree(root) that recursively inverts a binary tree node with .left and .right pointers and returns the root.\nAnswer:

  • Base 0.8B (NLL: 0.2347): Outputs commented-out dead code:
# def invert_tree(root):
#     if root is None:
#         return None
#     root.left, root.right = root.right, root.left
  • Transferred 0.8B (NLL: 0.1057, -55.0%): Outputs valid, executable Python with recursive traversal:
def invert_tree(root):
    if root is None:
        return None
    root.left, root.right = root.right, root.left
    invert_tree(root.left)
    invert_tree(root.right)
    return root
if __name__ == "__main__":

2. Russian Knights & Knaves Paradox (ru_knights_knaves_paradox)

Prompt: Question: На острове живут рыцари (всегда говорят правду) и лжецы (всегда лгут). Житель А говорит: «Я лжец». Кто житель А?\nAnswer:

  • Base 0.8B (NLL: 0.5535): Immediately hallucinates a single wrong sentence without reasoning: Житель А - лжец.
  • Transferred 0.8B (NLL: 0.2862, -48.3%): Spontaneously enters a structured <think> reasoning chain:
<think>
Мы рассматриваем ситуацию с двумя типами людей: рыцари (которые всегда правдивы) и лжецы (которые всегда лгут).
Житель А говорит: "Я лжец". Нужно определить, кем является житель А. Рассмотрим возможные варианты:
1. Если А - рыцарь...

3. CPython Global Interpreter Lock (algo_gil_python)

Prompt: Question: Why does Python's standard CPython runtime use a Global Interpreter Lock (GIL)?\nAnswer:

  • Base 0.8B: Gives generic filler ("prevents multiple threads from executing Python bytecodes simultaneously, which can lead to inefficiencies").
  • Transferred 0.8B: Identifies the exact low-level systems reason ("prevents multiple threads from accessing the same memory locations simultaneously, which could lead to race conditions and data corruption. The GIL is essential for maintaining thread safety").

Math & Solver Highlights

  1. Damped Tikhonov Inversion (faytuna_flow/nonlinear_transfer.py): Standard pseudoinverse np.linalg.pinv(w_down.T) blows up on near-zero singular values, forcing trust-region gates to crush updates down to 0.0009. We use Levenberg-Marquardt damping: $$\sigma_i^+ = \frac{\sigma_i}{\sigma_i^2 + \lambda}, \quad \lambda = \text{ridge} \cdot \frac{|A|_F^2}{\min(M, N)}$$

  2. Adaptive Spectral Rank Selection: Dynamically determines SVD truncation rank $r \in [16, 128]$ such that $\sum_{i=1}^r \sigma_i^2 / \sum \sigma_i^2 \ge 0.85$, retaining 85% of variance instead of fixed rank-16 truncation.

  3. Spectral Directional Rescale: Enforces $|\Delta W|2 \le 0.05 \cdot |W{\text{orig}}|_2$, ensuring that updates cannot alter the principal spectral direction of the original weight matrix.

  4. Null-Space Projector for Tied-Embedding Instruct Models (scripts/run_instruct_flow_transfer.py): For Instruct models where Layer 23 maps directly into tied token embeddings, we construct a semantic null-space projector: $$P_\perp = I - V_{16} V_{16}^T$$ derived from the top eigenvectors of the token embedding Gram matrix $W_{\text{embed}}^T W_{\text{embed}}$, protecting RLHF and vocabulary margins while computing the closed-form scale $\alpha^* = \frac{\langle \hat{\Delta}, \Delta Y \rangle_F}{|\hat{\Delta}|_F^2}$ strictly on assistant response tokens.


Hardware Accessibility: Consumer GPU Layer-Streaming

You do not need a cluster of H100s to perform trajectory transfer. To prove that direct weight surgery is accessible on commodity local hardware, all experiments were run on a single consumer 8GB AMD Radeon RX 580 using Layer-Outer VRAM Streaming (faytuna_flow/layer_streaming.py):

  1. Since we only need forward trajectories over a small calibration batch (8 to 32 prompts), we load embed_tokens and Layer 0 (~250 MB for 4B) into GPU VRAM via DirectML (torch-directml).
  2. All calibration prompts pass through Layer 0 in one batch.
  3. The resulting hidden states $H_1$ are saved to system RAM, Layer 0 is deleted from VRAM, and Layer 1 is loaded.
  4. We repeat this across all 32 layers.

Zero-Logit speedup: Calling model.model(...) directly instead of model(...) skips the final lm_head projection onto Qwen's 248,320 vocabulary tokens during trace collection, cutting extraction time by 38%.


Quickstart

1. Install dependencies

git clone https://github.com/dsadawq3/DynamicTune.git
cd DynamicTune
pip install -e .

2. Run unit tests

python -m pytest

3. Scan 24 layers for spectral entropy and knot detection

python scripts/scan_24_layers_autogate.py

4. Run 4-Block Anchor Surgery (4B -> 0.8B on GPU)

python scripts/run_qwen35_transfer.py \
  --student-dir /path/to/Qwen3.5-0.8B-Base \
  --teacher-dir /path/to/Qwen3.5-4B-Base \
  --output-dir runs/qwen35_anchor_surgery \
  --device dml \
  --use-layer-streaming \
  --export-gguf

5. Reproduce benchmarks via Vulkan llama.cpp

python scripts/run_hellaswag_audit.py
python scripts/bench_llama_cpp_real.py

Community, High-Priority Model Pairs, and Open Issues

We invite community members and researchers with modern 24GB-32GB+ GPUs (RTX 5090, RTX 4090, dual GPUs, or cloud nodes) to run trajectory surgery on frontier model pairs:

  • Next-Gen Qwen Reasoning & Flow Surgery:
    • Projecting trajectories from Qwen/Qwen3.8-27B (or Qwen3.8-Flash-Next) into Qwen/Qwen3.5-9B, 4B, or 2B.
    • Testing transfer between Mixture-of-Experts and dense backbones (Qwen3.6-35B-A3B -> Qwen3.5-4B).
  • Google Gemma-4 Cross-Scale Transfer:
    • Compressing the representation manifold of google/gemma-4-31B (or gemma-4-26B-A4B) into mobile-class edge models like google/gemma-4-E4B or gemma-4-E2B.
  • Abliteration & Agentic Trait Transplants:
    • Transferring refusal-ablation and guardrail-relaxation vectors from uncensored agentic models (e.g. Huihui-NeoHorse-1-4B-abliterated or Hermes) into restricted compact base models without destructive retraining.
  • Unknotting Superposition with Pre-Trained SAEs:
    • Can official sparse autoencoders like Qwen/SAE-Res-Qwen3.5-27B-W80K-L0_100 and Qwen/SAE-Res-Qwen3.5-9B-Base-W64K isolate monosemantic feature directions in layers 1-22, allowing trajectory transport across all layers rather than just 4 anchor blocks?

Open an Issue or Share Your Results

If you run DynamicTune on Qwen3.8, Gemma-4, or custom fine-tunes on an RTX 5090 or cloud cluster, open a GitHub Issue with your scan_24_layers_autogate.py entropy table, NLL logs, and GGUF outputs. We actively review PRs and discussion threads.


License

Apache 2.0. See LICENSE for details.

Languages

Python

100.0%