Collection of important articles to be treated as a textbook
Jupyter Notebook
880
284 commits
updated Aug 31, 2025
2025-05-28 - I haven't been updating this as diligently as I used to. Here's what I'm currently reading -- sort by the first column to see papers that I've revisited, and the second to see papers I've invested time on.
A curated collection of significant/impactful articles to be treated as a textbook, because sometimes it's just best to go straight to the source. My hope is to provide a reference for understanding important developments in the historical context that motivated them, e.g. the problems the authors were attempting to solve, what particular features of the discovery were considered especially novel or impressive when it was first published, what the competing theories or techniques at the time were, etc.
Someday this will be organized better.
Signal Processing Interpretation of Harmonic Analysis
Lasso/elasticnet
Boosting
Bagging
random forest
Adaboost
gradient boosting
bias-variance tradeoff
non-parametric bootstrap
permutation testing (target shuffle)
PCA
nonlinear PCA variants
ICA
LSI/LSA
LDA
SVM
NMF
random projections
MCMC - metropolis-hastings, HMC, reversible jump, NUTS
SMOTE
tSNE
UMAP
LSH
Feature hashing / hashing trick
the kernel trick
naive bayes
HMM
CRF
RBM - Restricted Boltzman Machine
GAM - General Additive Models
MARS - Multivariate adaptive regression splines
Decision Trees
KNN
Benford's Law
Guassian KDE
Boruta
Step-wise regression / forward selection / backwards elimination / recursive feature elimination
kalman filter
restricted boltzman machine
Deep belief networks
Scree plot
Collaborative Filtering (SVD and otherwise)
Market basket analysis
Process mining
self-organizing maps
Good overview of modeling process
poisson bootstrap
constrained optimization / Linear Programming
compressed sensing
perceptron algorithm
SGD / backprop
Adagrad / RMSProp
Adam
reverse-mode autodiff
gradient clipping
learning rate scheduling
distributed training
ZeRo Offload, Zero Redundandancy Optimizers
federated learning
K-FAC for approximating Fisher Information, Hessian
Good list here: https://spinningup.openai.com/en/latest/spinningup/rl_intro2.html#citations-below
Occupancy Networks
SIREN
Neural Radiance Fields (NeRF)
NeRF + Triplanes
TensoRT
Video Segmentation
LeNet
alexnet - Demonstrated importance of network depth (specifically stacking convolutions), and ReLU capability over the then conventional sigmoid and tanh activations
GAN, DCGAN, WGAN
StyleGAN -> StyleGANv2 -> StyleGAN2-ADA
U-net
VGG
inception/DeepDream
style transfer, content-texture decomposition, weight covariance transfer
cyclegan/discogan
YOLO
EfficientNet - Scaling laws for conv-resnets
FPN - Feature Pyramid Networks
Mask R-CNN
Mobilenet
Generative Diffusion models
VQVAE/VQGAN
Stable Diffusion
ViT
text-to-3d, score distillation sampling, dreamfusion
Gaussian Splatting
SAM
ConvNeXt
autoencoders
VAE (w inference amortization)
siamese network
student-teacher transfer learning, catastrophic forgetting
InfoNCE, contrastive learning
DINO
CLIP
Johnson-Lindenstrauss lemma
Neural ODE
Neural PDE
seq2seq
pix2pix
BNN
The Netflix Prize
Kaggle Galaxy Zoo
Capsule Networks
BiDirectional RNN
WaveNet
Normalizing flows
AlphaFold, EvoFormer
AlphaGo
IBM Watson on Jeopardy
Understanding the GPU compute paradigm
VC Dimension
gradient double descent
neural tangent kernel
lottery ticket hypothesis
manifold hypothesis
information bottleneck
generalized degrees of freedom
AIC / BIC
dropout as ensemblification
knowledge distillation
model quantization
SGD = MAP inference
shapley scoring
LIME
adversarial examples
(AU)ROC Curve / PR curve
Cohen's Kappa / meta-analysis
PAC learning
L1/L2 regularization expressable as bayesian priors
Hilbert spaces
No Free Lunch
Significance test for the LASSO
RNN's are near-sighted
Relationship between logistic regression and naive bayes
understanding softmax and its relation to log-sum-exp
Word2vec approximates a matrix factorization
The distributional hypothesis (computational linguistics)
Johnson–Lindenstrauss lemma (high dim manifolds can be accurately projected onto lower dim embeddings)
Empirical Risk Minimization
Loss geometry
generalization, overparameterization, effective model capacity, grokking
Strong inductive bias in CNN structure
prediction calibration
batch-norm and dropout in tug-of-war
deep learning model fitting process
Universal approximation theorem
Dropout as approximate bayesian inference
CMA-ES learns an approximation to the hessian of the loss
Inductive Biases
Neural Networks are essentially high dimensional decision trees. The latent space can be datum-specific. The learned manifold is not smooth, but heavily faceted. each neuron (relu) adds a hyperplane.
Extrapolation
Formalizing "intelligence"
PEFT: LoRA
Fourier features
NoPe - positional encodings not needed, learned implicitly
Intrinsic Dimension, PEFT
GAN training dynamics, TTUR
The Bitter Lesson
sigma reparameterization to stabilize transformer training by mitigating "entropy collapse" (concentration of density in attention)
buffer tokens, register tokens, attention sinks
Classifier-free Guidance (CFG)
SDEdit
Denoising diffusion as generic de-corruption
k-samplers, variance-preserving, variance exploding
Cross Attention guidance
Controlnet/T2I adaptors
Text inversion
null text inversion
Chain of thought
LLMs as role-play
learning to use tools
Chinchilla
Instruct tuning, InstructGPT
Speculative decoding
largely via https://twitter.com/karpathy/status/1668302116576976906
284 commits
Jupyter Notebook
98.4%
Python
1.0%
Collection of important articles to be treated as a textbook
Jupyter Notebook
880
284 commits
updated Aug 31, 2025
2025-05-28 - I haven't been updating this as diligently as I used to. Here's what I'm currently reading -- sort by the first column to see papers that I've revisited, and the second to see papers I've invested time on.
A curated collection of significant/impactful articles to be treated as a textbook, because sometimes it's just best to go straight to the source. My hope is to provide a reference for understanding important developments in the historical context that motivated them, e.g. the problems the authors were attempting to solve, what particular features of the discovery were considered especially novel or impressive when it was first published, what the competing theories or techniques at the time were, etc.
Someday this will be organized better.
Signal Processing Interpretation of Harmonic Analysis
Lasso/elasticnet
Boosting
Bagging
random forest
Adaboost
gradient boosting
bias-variance tradeoff
non-parametric bootstrap
permutation testing (target shuffle)
PCA
nonlinear PCA variants
ICA
LSI/LSA
LDA
SVM
NMF
random projections
MCMC - metropolis-hastings, HMC, reversible jump, NUTS
SMOTE
tSNE
UMAP
LSH
Feature hashing / hashing trick
the kernel trick
naive bayes
HMM
CRF
RBM - Restricted Boltzman Machine
GAM - General Additive Models
MARS - Multivariate adaptive regression splines
Decision Trees
KNN
Benford's Law
Guassian KDE
Boruta
Step-wise regression / forward selection / backwards elimination / recursive feature elimination
kalman filter
restricted boltzman machine
Deep belief networks
Scree plot
Collaborative Filtering (SVD and otherwise)
Market basket analysis
Process mining
self-organizing maps
Good overview of modeling process
poisson bootstrap
constrained optimization / Linear Programming
compressed sensing
perceptron algorithm
SGD / backprop
Adagrad / RMSProp
Adam
reverse-mode autodiff
gradient clipping
learning rate scheduling
distributed training
ZeRo Offload, Zero Redundandancy Optimizers
federated learning
K-FAC for approximating Fisher Information, Hessian
Good list here: https://spinningup.openai.com/en/latest/spinningup/rl_intro2.html#citations-below
Occupancy Networks
SIREN
Neural Radiance Fields (NeRF)
NeRF + Triplanes
TensoRT
Video Segmentation
LeNet
alexnet - Demonstrated importance of network depth (specifically stacking convolutions), and ReLU capability over the then conventional sigmoid and tanh activations
GAN, DCGAN, WGAN
StyleGAN -> StyleGANv2 -> StyleGAN2-ADA
U-net
VGG
inception/DeepDream
style transfer, content-texture decomposition, weight covariance transfer
cyclegan/discogan
YOLO
EfficientNet - Scaling laws for conv-resnets
FPN - Feature Pyramid Networks
Mask R-CNN
Mobilenet
Generative Diffusion models
VQVAE/VQGAN
Stable Diffusion
ViT
text-to-3d, score distillation sampling, dreamfusion
Gaussian Splatting
SAM
ConvNeXt
autoencoders
VAE (w inference amortization)
siamese network
student-teacher transfer learning, catastrophic forgetting
InfoNCE, contrastive learning
DINO
CLIP
Johnson-Lindenstrauss lemma
Neural ODE
Neural PDE
seq2seq
pix2pix
BNN
The Netflix Prize
Kaggle Galaxy Zoo
Capsule Networks
BiDirectional RNN
WaveNet
Normalizing flows
AlphaFold, EvoFormer
AlphaGo
IBM Watson on Jeopardy
Understanding the GPU compute paradigm
VC Dimension
gradient double descent
neural tangent kernel
lottery ticket hypothesis
manifold hypothesis
information bottleneck
generalized degrees of freedom
AIC / BIC
dropout as ensemblification
knowledge distillation
model quantization
SGD = MAP inference
shapley scoring
LIME
adversarial examples
(AU)ROC Curve / PR curve
Cohen's Kappa / meta-analysis
PAC learning
L1/L2 regularization expressable as bayesian priors
Hilbert spaces
No Free Lunch
Significance test for the LASSO
RNN's are near-sighted
Relationship between logistic regression and naive bayes
understanding softmax and its relation to log-sum-exp
Word2vec approximates a matrix factorization
The distributional hypothesis (computational linguistics)
Johnson–Lindenstrauss lemma (high dim manifolds can be accurately projected onto lower dim embeddings)
Empirical Risk Minimization
Loss geometry
generalization, overparameterization, effective model capacity, grokking
Strong inductive bias in CNN structure
prediction calibration
batch-norm and dropout in tug-of-war
deep learning model fitting process
Universal approximation theorem
Dropout as approximate bayesian inference
CMA-ES learns an approximation to the hessian of the loss
Inductive Biases
Neural Networks are essentially high dimensional decision trees. The latent space can be datum-specific. The learned manifold is not smooth, but heavily faceted. each neuron (relu) adds a hyperplane.
Extrapolation
Formalizing "intelligence"
PEFT: LoRA
Fourier features
NoPe - positional encodings not needed, learned implicitly
Intrinsic Dimension, PEFT
GAN training dynamics, TTUR
The Bitter Lesson
sigma reparameterization to stabilize transformer training by mitigating "entropy collapse" (concentration of density in attention)
buffer tokens, register tokens, attention sinks
Classifier-free Guidance (CFG)
SDEdit
Denoising diffusion as generic de-corruption
k-samplers, variance-preserving, variance exploding
Cross Attention guidance
Controlnet/T2I adaptors
Text inversion
null text inversion
Chain of thought
LLMs as role-play
learning to use tools
Chinchilla
Instruct tuning, InstructGPT
Speculative decoding
largely via https://twitter.com/karpathy/status/1668302116576976906
284 commits
Jupyter Notebook
98.4%
Python
1.0%