
I have many little research ideas I've been curious to try (especially after being out on paternity leave!). Some are more research-oriented, some are more abstract/artistic/weird.
I'd love to give more time to each one, but as a forcing function I wanted to do the advent of ML ideas, for 25 days until Christmas I want to put out a half-baked idea with some initial experimentation. (And 1-2 posts will just be catching up on my reading list.)
Each day will get a folder under the src/ directory. Each day should have a README that describes the experiment, code, and results - but will also link to the Twitter post that gives more detail.
Each day/folder is totally self-contained. Even if some days build off previous days, each folder will be a self-contained project, even if it means repeated code. While all experiments share the same uv environment (defined by the root uv.lock), each day's code is independent and can be run on its own without dependencies on other days.
Yes, this means there will be repeated code across folders. But the goal here is simplicity and ease of experimentation - you can test things out, copy a folder, modify it, and not worry about breaking other experiments.
This project uses uv for package management and uses a shared environment across all experiments. The repo includes a uv.lock file, which allows you to recreate the environment identically.
To get started:
Install uv (if you haven't already):
curl -LsSf https://astral.sh/uv/install.sh | sh
Sync the environment (creates virtual environment and installs exact pinned versions from uv.lock):
uv sync
Each day's experiment lives in src/day_XX/ with its own README documenting the approach, code, and results.
| Day | Summary |
|---|---|
| Day 1 | Reading List — Curated ~80 papers, repos, and startup ideas from my bookmarks backlog (July → early November). |
| Day 2 | Unsupervised VLM Training — CycleGAN-style training for vision-language models: describe a chart, regenerate it with Flux, use cosine similarity as reward. ~8% improvement on CharXiv reasoning without any labeled data. |
| Day 3 | Adversarial VLM Training — Extension of Day 2 with an adversary model that learns to generate prompts for challenging images, creating a competitive training dynamic. |
| Day 4 | Steering Vectors for Reasoning — Extract activation differences between base and GRPO-finetuned models to create lightweight steering vectors that boost math reasoning performance. |
| Day 5 | Bayesian Optimization for Steering — Use Gaussian Processes to find optimal per-layer-group steering weights, matching full finetuning performance with just a few vectors. |
| Day 6 | Steering Toward Novelty — Can we steer LLMs toward more creative outputs? Using award-winning paper abstracts to build "creativity vectors" that guide models away from generic text. |
| Day 7 | Entropy-Based Rewards — Reward models for maintaining higher entropy in middle layers during reasoning, inspired by findings that reasoning models have higher internal entropy. |
| Day 8 | LLM Personality Training — Measure and modify LLM personality using the Big Five (OCEAN) framework. Train models toward target personalities like "the jerk" who actually pushes back on bad ideas. |
| Day 9 | GEPA vs GRPO — Can prompt optimization match finetuning? Compare evolving prompts with LLM-powered reflection (GEPA) against weight updates (GRPO) on math reasoning. |
| Day 10 | GEPA + GRPO Composition — Do prompt optimization and finetuning stack? Yes! Best results come from evolving a prompt with GEPA, then finetuning with GRPO using that prompt. |
| Day 11 | Human Preference GRPO — Train a model to write image prompts that make you happy, using round-robin tournaments with real human feedback. Impractical but surprisingly effective. |
| Day 12 | Emotion-Based GRPO — Replace human clicking with facial expression analysis. Hume AI reads your face as you view generated images, turning emotional reactions into reward signals. |
| Day 13 | Cartridges — Compress long documents into tiny learnable KV caches. A 6000+ token document distilled to 1024 vectors achieves ~72% of full-context accuracy with 16% of the tokens. |
| Day 14 | ENGRAM Continual Learning — Mimic human learning: develop skills via prompt optimization, then distill into cartridge KV vectors. Iteratively build frozen cartridges while resetting skills, enabling continual learning without weight updates. |
| Day 15 | Teaching a Model to Daydream — Apply ENGRAM to creative thinking itself. Train a model to find novel connections between random concepts (entropy ↔ democracy), distilling the skill of "seeing connections" into a cartridge. |
| Day 16 | ENGRAM for Wiki Search — ENGRAM meets tool use. A multi-turn agent learns Wikipedia search strategies through skill refinement and cartridge distillation, improving from 0% to 40% accuracy on trivia questions. |
| Day 17 | Synthetic Persona Simulation — Simulate how 1M AI personas (NVIDIA's Nemotron-Personas-USA) react to any content. Interactive NYT-style dashboard visualizes results across geography, age, education, and more. |
| Day 18 | GRPO with Persona Judges — Train models to write content that resonates with specific demographics. Use 1M synthetic personas as judges in GRPO training, watch your model learn to beat GPT-4.1 with your target audience. |
| Day 19 | Evolution Strategies for LLMs — Gradient-free fine-tuning using Evolution Strategies. Perturb weights, evaluate with greedy decoding, update based on reward-weighted noise. Testing claims from "ES at Scale" paper on MATH dataset. |
| Day 20 | Training Small Reasoning Models on SYNTH — Train a 56M parameter reasoning model from scratch on PleIAs's SYNTH dataset. Full pipeline: download 68M samples, pre-tokenize, distributed training on 4x H100 in ~1 hour. |
| Day 21 | NEFTune for Format Learning — Add noisy embeddings (NEFTune) to small model training. Surprising result: +19% improvement in format compliance (<think>...</think> structure) even though training loss didn't improve. |
| Day 22 | Weighted Loss for Reasoning — Weight <think> tokens at 0.5× compared to answer tokens during training. Best format compliance yet: 82.3% valid rate (+31% over baseline), though accuracy slightly lower. |
| Day 23 | Looped Reasoning Layers — Custom GPT-style transformer where the last 10% of layers loop multiple times with LayerNorm between iterations. Testing if iterative computation on reasoning layers improves performance. |
| Day 24 | GRPO with LLM-as-Judge — Train reasoning without ground truth labels. GPT-4.1 judges which of 4 reasoning chains is "better" in round-robin comparisons; win rate becomes the reward signal. Moved the needle (+0.2%) purely from preference over reasoning quality. |
| Day 25 | Reading List — Final day! Curated reading list from mid-November to December 20th - papers, repos, and ideas found on Twitter. |
Some (or maybe all) of these ideas might be worth fleshing out more - if you're interested in that or just want to talk more, please message me on Twitter @brendanh0gan
1,021 followers · starred Dec 2025
514 followers · starred Sep 2026
Python
91.3%
JavaScript
5.2%
CSS
2.5%

I have many little research ideas I've been curious to try (especially after being out on paternity leave!). Some are more research-oriented, some are more abstract/artistic/weird.
I'd love to give more time to each one, but as a forcing function I wanted to do the advent of ML ideas, for 25 days until Christmas I want to put out a half-baked idea with some initial experimentation. (And 1-2 posts will just be catching up on my reading list.)
Each day will get a folder under the src/ directory. Each day should have a README that describes the experiment, code, and results - but will also link to the Twitter post that gives more detail.
Each day/folder is totally self-contained. Even if some days build off previous days, each folder will be a self-contained project, even if it means repeated code. While all experiments share the same uv environment (defined by the root uv.lock), each day's code is independent and can be run on its own without dependencies on other days.
Yes, this means there will be repeated code across folders. But the goal here is simplicity and ease of experimentation - you can test things out, copy a folder, modify it, and not worry about breaking other experiments.
This project uses uv for package management and uses a shared environment across all experiments. The repo includes a uv.lock file, which allows you to recreate the environment identically.
To get started:
Install uv (if you haven't already):
curl -LsSf https://astral.sh/uv/install.sh | sh
Sync the environment (creates virtual environment and installs exact pinned versions from uv.lock):
uv sync
Each day's experiment lives in src/day_XX/ with its own README documenting the approach, code, and results.
| Day | Summary |
|---|---|
| Day 1 | Reading List — Curated ~80 papers, repos, and startup ideas from my bookmarks backlog (July → early November). |
| Day 2 | Unsupervised VLM Training — CycleGAN-style training for vision-language models: describe a chart, regenerate it with Flux, use cosine similarity as reward. ~8% improvement on CharXiv reasoning without any labeled data. |
| Day 3 | Adversarial VLM Training — Extension of Day 2 with an adversary model that learns to generate prompts for challenging images, creating a competitive training dynamic. |
| Day 4 | Steering Vectors for Reasoning — Extract activation differences between base and GRPO-finetuned models to create lightweight steering vectors that boost math reasoning performance. |
| Day 5 | Bayesian Optimization for Steering — Use Gaussian Processes to find optimal per-layer-group steering weights, matching full finetuning performance with just a few vectors. |
| Day 6 | Steering Toward Novelty — Can we steer LLMs toward more creative outputs? Using award-winning paper abstracts to build "creativity vectors" that guide models away from generic text. |
| Day 7 | Entropy-Based Rewards — Reward models for maintaining higher entropy in middle layers during reasoning, inspired by findings that reasoning models have higher internal entropy. |
| Day 8 | LLM Personality Training — Measure and modify LLM personality using the Big Five (OCEAN) framework. Train models toward target personalities like "the jerk" who actually pushes back on bad ideas. |
| Day 9 | GEPA vs GRPO — Can prompt optimization match finetuning? Compare evolving prompts with LLM-powered reflection (GEPA) against weight updates (GRPO) on math reasoning. |
| Day 10 | GEPA + GRPO Composition — Do prompt optimization and finetuning stack? Yes! Best results come from evolving a prompt with GEPA, then finetuning with GRPO using that prompt. |
| Day 11 | Human Preference GRPO — Train a model to write image prompts that make you happy, using round-robin tournaments with real human feedback. Impractical but surprisingly effective. |
| Day 12 | Emotion-Based GRPO — Replace human clicking with facial expression analysis. Hume AI reads your face as you view generated images, turning emotional reactions into reward signals. |
| Day 13 | Cartridges — Compress long documents into tiny learnable KV caches. A 6000+ token document distilled to 1024 vectors achieves ~72% of full-context accuracy with 16% of the tokens. |
| Day 14 | ENGRAM Continual Learning — Mimic human learning: develop skills via prompt optimization, then distill into cartridge KV vectors. Iteratively build frozen cartridges while resetting skills, enabling continual learning without weight updates. |
| Day 15 | Teaching a Model to Daydream — Apply ENGRAM to creative thinking itself. Train a model to find novel connections between random concepts (entropy ↔ democracy), distilling the skill of "seeing connections" into a cartridge. |
| Day 16 | ENGRAM for Wiki Search — ENGRAM meets tool use. A multi-turn agent learns Wikipedia search strategies through skill refinement and cartridge distillation, improving from 0% to 40% accuracy on trivia questions. |
| Day 17 | Synthetic Persona Simulation — Simulate how 1M AI personas (NVIDIA's Nemotron-Personas-USA) react to any content. Interactive NYT-style dashboard visualizes results across geography, age, education, and more. |
| Day 18 | GRPO with Persona Judges — Train models to write content that resonates with specific demographics. Use 1M synthetic personas as judges in GRPO training, watch your model learn to beat GPT-4.1 with your target audience. |
| Day 19 | Evolution Strategies for LLMs — Gradient-free fine-tuning using Evolution Strategies. Perturb weights, evaluate with greedy decoding, update based on reward-weighted noise. Testing claims from "ES at Scale" paper on MATH dataset. |
| Day 20 | Training Small Reasoning Models on SYNTH — Train a 56M parameter reasoning model from scratch on PleIAs's SYNTH dataset. Full pipeline: download 68M samples, pre-tokenize, distributed training on 4x H100 in ~1 hour. |
| Day 21 | NEFTune for Format Learning — Add noisy embeddings (NEFTune) to small model training. Surprising result: +19% improvement in format compliance (<think>...</think> structure) even though training loss didn't improve. |
| Day 22 | Weighted Loss for Reasoning — Weight <think> tokens at 0.5× compared to answer tokens during training. Best format compliance yet: 82.3% valid rate (+31% over baseline), though accuracy slightly lower. |
| Day 23 | Looped Reasoning Layers — Custom GPT-style transformer where the last 10% of layers loop multiple times with LayerNorm between iterations. Testing if iterative computation on reasoning layers improves performance. |
| Day 24 | GRPO with LLM-as-Judge — Train reasoning without ground truth labels. GPT-4.1 judges which of 4 reasoning chains is "better" in round-robin comparisons; win rate becomes the reward signal. Moved the needle (+0.2%) purely from preference over reasoning quality. |
| Day 25 | Reading List — Final day! Curated reading list from mid-November to December 20th - papers, repos, and ideas found on Twitter. |
Some (or maybe all) of these ideas might be worth fleshing out more - if you're interested in that or just want to talk more, please message me on Twitter @brendanh0gan
1,021 followers · starred Dec 2025
514 followers · starred Sep 2026
Python
91.3%
JavaScript
5.2%
CSS
2.5%