
Here are my personal deep learning notes. I've written this cheatsheet for keep track my knowledge but you can use it as a guide for learning deep learning aswell.
| 🗂 Data | 🧠 Layers | 📉 Loss | 📈 Metrics | 🔥 Training | ✅ Production |
|---|---|---|---|---|---|
| Pytorch dataset | Weight init | Cross entropy | Optimizers | Ensemble | |
| Pytorch dataloader | Activations | Weight Decay | Transfer learning | TTA | |
| Split | Self Attention | Label Smoothing | Clean mem | Pseudolabeling | |
| Normalization | Trained CNN | Mixup | Half precision | Webserver (Flask) | |
| Data augmentation | CoordConv | SoftF1 | Multiple GPUs | Distillation | |
| Deal imbalance | Precomputation | Pruning | |||
| Set seed | Quantization (int8) | ||||
| TorchScript | |||||
| ONNX |
https://github.com/lucidrains?tab=repositories https://walkwithfastai.com/
If you can not get more data of the underrepresented classes, you can fix the imbalance with code:
sampler:
torch.utils.data.WeightedRandomSampler(weights=[…])catalyst.data.sampler.BalanceClassSampler(labels=ds.targets, mode="downsampling")catalyst.data.sampler.BalanceClassSampler(labels=ds.targets, mode="upsampling")CrossEntropyLoss(weight=[…])
class BalanceClassSampler(torch.utils.data.Sampler):
"""
Allows you to create stratified sample on unbalanced classes.
Inspired from Catalyst's BalanceClassSampler:
https://catalyst-team.github.io/catalyst/_modules/catalyst/data/sampler.html#BalanceClassSampler
Args:
labels: list of class label for each elem in the dataset
mode: Strategy to balance classes. Must be one of [downsampling, upsampling]
"""
def __init__(self, labels:list[int], mode:str = "upsampling"):
labels = np.array(labels)
self.unique_labels = set(labels)
########## STEP 1:
# Compute the final_num_samples_per_label
# An Integer
num_samples_per_label = {label: (labels == label).sum() for label in self.unique_labels}
if mode == "upsampling": self.final_num_samples_per_label = max(num_samples_per_label.values())
elif mode == "downsampling": self.final_num_samples_per_label = min(num_samples_per_label.values())
else: raise Exception("mode should be: \"downsampling\" or \"upsampling\"")
########## STEP 2:
# Compute actual indices of every label.
# A Diccionary of lists
self.indices_per_label = {label: np.arange(len(labels))[labels==label].tolist() for label in self.unique_labels}
def __iter__(self): #-> Iterator[int]:
indices = []
for label in self.unique_labels:
label_indices = self.indices_per_label[label]
repeat_all_elementes = self.final_num_samples_per_label // len(label_indices)
pick_random_elementes = self.final_num_samples_per_label % len(label_indices)
indices += label_indices * repeat_all_elementes # repeat the list several times
indices += random.sample(label_indices, k=pick_random_elementes) # pick random idxs without repetition
assert len(indices) == self.__len__()
np.random.shuffle(indices) # Inplace shuffle the list
return iter(indices)
def __len__(self) -> int:
return self.final_num_samples_per_label * len(self.unique_labels)
10% or 20% of your train set.
10Scale the inputs to have mean 0 and a variance of 1. Also linear decorrelation/whitening/pca helps a lot. Normalization parameters are obtained only from train set, and then applied to both train and valid sets.
x = x-x.mean() / x.std() Most used
x = x - x.mean() fights vanishing and exploding gradientsx = x / x.std() improves convergence speed and accuracyx = x - x.mean()whitened = decorrelated / np.sqrt(eigVals + 1e-5)(x-x.min()) / (x.max()-x.min()): Values from 0 to 12*(x-x.min()) / (x.max()-x.min()) - 1: Values from -1 to 1
- In case of images, the scale is from 0 to 255, so it is not strictly necessary normalize.
- neural networks data preparation
x = λxᵢ + (1−λ)xⱼ & y = λyᵢ + (1−λ)yⱼ. Fast.ai doc
λ sampleando la distribución beta α=β=0.4 ó 0.2 (Así pocas veces la imgs se mezclarán)
| Augmentation | Description | Pillow |
|---|---|---|
| Rotate | Rotate some degrees | pil_img.rotate() |
| Translate | pil_img.transform() | |
| Shear | Affine transform | pil_img.transform() |
| Autocontrast | Equalize the histogram (linear) | PIL.ImageOps.autocontrast() |
| Equalize | Equalize the histogram (non-linear) | PIL.ImageOps.equalize() |
| Posterize | Reducing pixel bits | PIL.ImageOps.posterize() |
| Solarize | Inverting colors above a threshold | PIL.ImageOps.solarize() |
| Color | PIL.ImageEnhance.Color() | |
| Contrast | PIL.ImageEnhance.Contrast() | |
| Brightness | PIL.ImageEnhance.Brightness() | |
| Sharpness | Sharpen or blurs the image | PIL.ImageEnhance.Sharpness() |
Interpolations when rotate, translate or affine:
Depends on the models architecture. Try to avoid vanishing or exploding outputs. blog1, blog2.
ReLU(x) equals to min(x,0) - 0.5 for a correct mean (0)def weight_init(m):
# LINEAR
if type(m) == nn.Linear:
torch.nn.init.xavier_uniform(m.weight)
m.bias.data.fill_(0.01)
# CONVS
classname = m.__class__.__name__
if classname.find('Conv') != -1:
nn.init.xavier_uniform_(m.weight, gain=nn.init.calculate_gain('relu'))
nn.init.zeros_(m.bias)
model.apply(weight_init)
linear1(x) * sigmoid(linear2(x))x * sigmoid(x) paper (2017)xxxx paper (2018)x * tanh( ln(1 + e^x) ) paper (2019)0.5 * x * ( tanh(x) + 1 )0.5 * x * ( tanh (x+1) + 1)x * ((x+x+1)/(abs(x+1) + abs(x)) * 0.5 + 0.5)class AddCoord2D(torch.nn.Module):
def __init__(self, len):
super(AddCoord2D, self).__init__()
i_coord = torch.linspace(start=1/len, end=1, steps=len).view(len, -1).expand(-1, len)
j_coord = torch.linspace(start=1/len, end=1, steps=len).view(-1, len).expand(len, -1)
self.coords = torch.stack([i_coord, j_coord])
print(self.coords.shape)
def forward(self, x): # X shape: [BS, C, X, Y]
BS = x.shape[0]
return torch.cat((x, self.coords.expand(BS,-1,-1,-1)), dim=1)
During training, some neurons will be deactivated randomly. Hinton, 2012, Srivasta, 2014

Weight penalty: Regularization in loss function (penalice high weights). Weight decay hyper-parameter usually 0.0005.
Visually, the weights only can take a value inside the blue region, and the red circles represent the minimum. Here, there are 2 weight variables.
| L1 (LASSO) | L2 (Ridge) | Elastic Net |
|---|---|---|
![]() | ![]() | ![]() |
| Shrinks coefficients to 0. Good for variable selection | Most used. Makes coefficients smaller | Tradeoff between variable selection and small coefficients |
| Penalizes the sum of absolute weights | Penalizes the sum of squared weights | Combination of 2 before |
loss + wd * weights.abs().sum() | loss + wd * weights.pow(2).sum() |
At training and inference, some connections (weights) will be deactivated permanently. LeCun, 2013. This is very useful at the firsts layers.

Knowledge Distillation (teacher-student) A teacher model teach a student model.
mean(GT - pred) It could determine if the model has positive bias or negative bias.mean(|GT - pred|) The most simple.mean((GT-pred)²) Penalice large errors more than MAE. Most usedsqrt(MSE) Proportional to MSE. Value closer to MAE.nn.CrossEntropyLoss.
nn.NLLLoss()nn.BCELossnn.HingeEmbeddingLoss()-(1-p)^gamma * log(p) paper(Pred ∩ GT)/(Pred ∪ GT) = TP / TP + FP * FN2 * (Pred ∩ GT)/(Pred + GT) = 2·TP / 2·TP + FP * FN
0 (worst) to 1 (best)1 − DiceSmooth the one-hot target label.
LabelSmoothingCrossEntropy(eps:float=0.1, reduction='mean')
Referennce
Dataset with 5 disease images and 20 normal images. If the model predicts all images to be normal, its accuracy is 80%, and F1-score of such a model is 0.88
TP + TN / TP + TN + FP + FN2 * (Prec*Rec)/(Prec+Rec)
TP / TP + FP = TP / predicted possitivesTP / TP + FN = TP / actual possitives2 * (Pred ∩ GT)/(Pred + GT)How big the steps are during training.
lr_find())Number of samples to learn simultaneously.
Batch size = 1: Train each sample individually. (Online gradient descent) ❌Batch size = length(dataset): Train the whole dataset at once, as a batch. (Batch gradient descent) ❌Batch size = number: Train disjoint groups of samples (Mini-batch gradient descent). ✅
32 or 64 are good values.4: Lot of updates. Very noisy random updates in the net (bad).512 Few updates. Very general common updates (bad).
Some people are tring to make a batch size finder according to this paper.
Times to learn the whole dataset.
def seed_everything(seed):
os.environ['PYTHONHASHSEED'] = str(seed)
random.seed(seed) # Random
np.random.seed(seed) # Numpy
torch.manual_seed(seed) # Pytorch
torch.cuda.manual_seed(seed)
torch.backends.cudnn.deterministic = True
torch.backends.cudnn.benchmark = False
#tf.random.set_seed(seed) # Tensorflow
Read this
def clean_mem():
gc.collect()
torch.cuda.empty_cache()
learn.to_parallel()
Reference
learn.to_fp16()
learn.to_fp32()
Reference
import numpy as np
import torch
from torchvision import models
import torchvision.transforms as transforms
from PIL import Image
from flask import Flask, jsonify, request
import json
app = Flask(__name__)
app.config['JSON_SORT_KEYS'] = False
classes = json.load(open('imagenet_classes.json'))
model = models.densenet121(pretrained=True)
model.eval()
def pre_process(image_file):
my_transforms = transforms.Compose([transforms.Resize(255),
transforms.CenterCrop(224),
transforms.ToTensor(),
transforms.Normalize([0.485, 0.456, 0.406], [0.229, 0.224, 0.225])])
image = Image.open(image_file)
return my_transforms(image).unsqueeze(0) # unsqueeze is for the BS dim
def post_process(logits):
vals, idxs = logits.softmax(1).topk(5)
vals = vals[0].numpy()
idxs = idxs[0].numpy()
result = {}
for idx, val in zip(idxs, vals):
result[classes[idx]] = round(float(val), 4)
return result
def get_prediction(image_file):
with torch.no_grad():
image_tensor = pre_process(image_file)
output = model.forward(image_tensor)
return post_process(output)
@app.route('/predict', methods=['POST'])
def predict():
if request.method == 'POST':
image_file = request.files['my_img_file']
result_dict = get_prediction(image_file)
#return jsonify(result_dict)
#return json.dumps(result_dict)
return result_dict
if __name__ == '__main__':
app.run()
FLASK_ENV=development FLASK_APP=app.py flask run
curl -X POST -F my_img_file=@cardigan.jpg http://localhost:5000/predict
import requests
resp = requests.post("http://localhost:5000/predict",
files={"my_img_file": open('cardigan.jpg','rb')})
print(resp.json())
{
"cardigan": 0.7083,
"wool": 0.0837,
"suit": 0.0431,
"Windsor_tie": 0.031,
"trench_coat": 0.0307
}
| What | Accuracy | Pytorch API | |
|---|---|---|---|
| Dynamic Quantization | Weights only | Good | qmodel = torch.quantization.quantize_dynamic(model, dtype=torch.qint8) |
| Post Training Quantization | Weights and activations | Good | model.qconfig = torch.quantization.default_qconfig torch.quantization.prepare(model, inplace=True) torch.quantization.convert(model, inplace=True) |
| Quantization-Aware Training | Weights and activations | Best | torch.quantization.prepare_qat -> torch.quantization.convert |
Reference
import torch.nn.utils.prune as prune
parameters_to_prune = (
(model.conv1, 'weight'),
(model.conv2, 'weight'),
(model.fc1, 'weight'),
(model.fc2, 'weight'),
(model.fc3, 'weight'),
)
prune.global_unstructured(
parameters_to_prune,
pruning_method=prune.L1Unstructured,
amount=0.2,
)
class ThresholdPruning(prune.BasePruningMethod):
PRUNING_TYPE = "unstructured"
def __init__(self, threshold): self.threshold = threshold
def compute_mask(self, tensor, default_mask): return torch.abs(tensor) > self.threshold
prune.global_unstructured(
parameters_to_prune,
pruning_method=ThresholdPruning,
threshold=0.01
)
def pruned_info(model):
print("Weights pruned:")
print("==============")
total_pruned, total_weights = 0,0
for name, chil in model.named_children():
layer_pruned = torch.sum(chil.weight == 0)
layer_weights = chil.weight.nelement()
total_pruned += layer_pruned
total_weights += layer_weights
print(name, "\t{:.2f}%".format(100 * float(layer_pruned)/ float(layer_weights)))
print("==============")
print("Total\t{:.2f}%".format(100 * float(total_pruned)/ float(total_weights)))
# Weights pruned:
# ==============
# conv1 1.85%
# conv2 8.10%
# fc1 19.76%
# fc2 10.66%
# fc3 9.40%
# ==============
# Total 17.90%
Iterative magnitude pruning is iterative process of removing connections (Prune/Train/Repeat):

At the end you can have pruned the 15%, 30%, 45%, 60%, 75%, and 90% of your original model.
Reference
- Code:
- Papers:
- Deep Compression (2015)
- Train Large, Then Compress (2020)
- Neural Networks are Surprisingly Modular (2020)
torch_script = torch.jit.script(MyModel())
torch_script.save("my_model_script.pt")
Reference
torch.onnx.export(model, img, f, verbose=False, opset_version=11) # Export to onnx
# Check onnx model
import onnx
model = onnx.load(f) # load onnx model
onnx.checker.check_model(model) # check onnx model
print(onnx.helper.printable_graph(model.graph)) # print a human readable representation of the graph
print('Export complete. ONNX model saved to %s\nView with https://github.com/lutzroeder/netron' % f)
Reference
0.50.0005
0.01 or 0.1.wd * w. Sometimes mathematically identical to L2 reg.0.632 Prob of sample in Out Of Bag 0.368Other tricks:
- Label Smoothing: Smooth the one-hot target label
- Knowledge Distillation: A bigger trained net (teacher) helps the network paper
loss = recontruction loss + latent loss
- 2D: [x,y]->[R,G,B]
- 3D: [x,y,z]->[R,G,B,alpha]
| Description | Website | Video | Paper |
|---|---|---|---|
| NeRF in the Wild | web | 3:41 | Aug 2020 |
| NeRF++ | Oct 2020 | ||
| Deformable NeRF (nerfies) | web | 7:26 | Nov 2020 |
| NeRF with time dimension | web | 2:21 | Nov 2020 |
| NeRF with better weight init | web | 3:54 | Dec 2020 |
Check this kaggle discussion
Reinforcement learning reference
Projections (BAD REPRESENTATION) (complicated things with voxels) Dense matrix (antor) - Its a depth map i think - Not projections - NAtive output of the sensor but condensed in a dense matrix
TODO
- Multi-Task Learning: Train a model on a variety of learning tasks
- Meta-learning: Learn new tasks with minimal data using prior knowledge.
- N-Shot Learning
- Zero-shot: 0 trainning examples of that class.
- One-shot: 1 trainning example of that class.
- Few-shot: 2...5 trainning examples of that class.
- Models
- Naive approach: re-training the model on the new data, would severely overfit.
- Siamese Networks (2015) Knows if to inputs are the same or not. (2 Feature extraction shares wights)
- Matching Networks (2016) Weighted nearest-neighbor classifier applied within an embedding space.
- Model-Agnostic Meta-Learning (MAML) (2017)
- Prototypical Networks (2017): Better nearest-neighbor classifier of embeddings.
- Meta-Learning for Semi-Supervised classification (2018) Extensions of Prototypical Networks. SotA.
- Meta-Transfer Learning (MTL) (2018)
- Online Meta-Learning (2019)
- Neural Turing machine. paper, code
- Neural Arithmetic Logic Units (NALU) paper
- Remember the math:
- Matrix calculus
- Einsum: link 1, link 2
nvidia-smi daemon: Check that sm% is near to 100% for a good GPU usage.
Jupyter Notebook
99.9%

Here are my personal deep learning notes. I've written this cheatsheet for keep track my knowledge but you can use it as a guide for learning deep learning aswell.
| 🗂 Data | 🧠 Layers | 📉 Loss | 📈 Metrics | 🔥 Training | ✅ Production |
|---|---|---|---|---|---|
| Pytorch dataset | Weight init | Cross entropy | Optimizers | Ensemble | |
| Pytorch dataloader | Activations | Weight Decay | Transfer learning | TTA | |
| Split | Self Attention | Label Smoothing | Clean mem | Pseudolabeling | |
| Normalization | Trained CNN | Mixup | Half precision | Webserver (Flask) | |
| Data augmentation | CoordConv | SoftF1 | Multiple GPUs | Distillation | |
| Deal imbalance | Precomputation | Pruning | |||
| Set seed | Quantization (int8) | ||||
| TorchScript | |||||
| ONNX |
https://github.com/lucidrains?tab=repositories https://walkwithfastai.com/
If you can not get more data of the underrepresented classes, you can fix the imbalance with code:
sampler:
torch.utils.data.WeightedRandomSampler(weights=[…])catalyst.data.sampler.BalanceClassSampler(labels=ds.targets, mode="downsampling")catalyst.data.sampler.BalanceClassSampler(labels=ds.targets, mode="upsampling")CrossEntropyLoss(weight=[…])
class BalanceClassSampler(torch.utils.data.Sampler):
"""
Allows you to create stratified sample on unbalanced classes.
Inspired from Catalyst's BalanceClassSampler:
https://catalyst-team.github.io/catalyst/_modules/catalyst/data/sampler.html#BalanceClassSampler
Args:
labels: list of class label for each elem in the dataset
mode: Strategy to balance classes. Must be one of [downsampling, upsampling]
"""
def __init__(self, labels:list[int], mode:str = "upsampling"):
labels = np.array(labels)
self.unique_labels = set(labels)
########## STEP 1:
# Compute the final_num_samples_per_label
# An Integer
num_samples_per_label = {label: (labels == label).sum() for label in self.unique_labels}
if mode == "upsampling": self.final_num_samples_per_label = max(num_samples_per_label.values())
elif mode == "downsampling": self.final_num_samples_per_label = min(num_samples_per_label.values())
else: raise Exception("mode should be: \"downsampling\" or \"upsampling\"")
########## STEP 2:
# Compute actual indices of every label.
# A Diccionary of lists
self.indices_per_label = {label: np.arange(len(labels))[labels==label].tolist() for label in self.unique_labels}
def __iter__(self): #-> Iterator[int]:
indices = []
for label in self.unique_labels:
label_indices = self.indices_per_label[label]
repeat_all_elementes = self.final_num_samples_per_label // len(label_indices)
pick_random_elementes = self.final_num_samples_per_label % len(label_indices)
indices += label_indices * repeat_all_elementes # repeat the list several times
indices += random.sample(label_indices, k=pick_random_elementes) # pick random idxs without repetition
assert len(indices) == self.__len__()
np.random.shuffle(indices) # Inplace shuffle the list
return iter(indices)
def __len__(self) -> int:
return self.final_num_samples_per_label * len(self.unique_labels)
10% or 20% of your train set.
10Scale the inputs to have mean 0 and a variance of 1. Also linear decorrelation/whitening/pca helps a lot. Normalization parameters are obtained only from train set, and then applied to both train and valid sets.
x = x-x.mean() / x.std() Most used
x = x - x.mean() fights vanishing and exploding gradientsx = x / x.std() improves convergence speed and accuracyx = x - x.mean()whitened = decorrelated / np.sqrt(eigVals + 1e-5)(x-x.min()) / (x.max()-x.min()): Values from 0 to 12*(x-x.min()) / (x.max()-x.min()) - 1: Values from -1 to 1
- In case of images, the scale is from 0 to 255, so it is not strictly necessary normalize.
- neural networks data preparation
x = λxᵢ + (1−λ)xⱼ & y = λyᵢ + (1−λ)yⱼ. Fast.ai doc
λ sampleando la distribución beta α=β=0.4 ó 0.2 (Así pocas veces la imgs se mezclarán)
| Augmentation | Description | Pillow |
|---|---|---|
| Rotate | Rotate some degrees | pil_img.rotate() |
| Translate | pil_img.transform() | |
| Shear | Affine transform | pil_img.transform() |
| Autocontrast | Equalize the histogram (linear) | PIL.ImageOps.autocontrast() |
| Equalize | Equalize the histogram (non-linear) | PIL.ImageOps.equalize() |
| Posterize | Reducing pixel bits | PIL.ImageOps.posterize() |
| Solarize | Inverting colors above a threshold | PIL.ImageOps.solarize() |
| Color | PIL.ImageEnhance.Color() | |
| Contrast | PIL.ImageEnhance.Contrast() | |
| Brightness | PIL.ImageEnhance.Brightness() | |
| Sharpness | Sharpen or blurs the image | PIL.ImageEnhance.Sharpness() |
Interpolations when rotate, translate or affine:
Depends on the models architecture. Try to avoid vanishing or exploding outputs. blog1, blog2.
ReLU(x) equals to min(x,0) - 0.5 for a correct mean (0)def weight_init(m):
# LINEAR
if type(m) == nn.Linear:
torch.nn.init.xavier_uniform(m.weight)
m.bias.data.fill_(0.01)
# CONVS
classname = m.__class__.__name__
if classname.find('Conv') != -1:
nn.init.xavier_uniform_(m.weight, gain=nn.init.calculate_gain('relu'))
nn.init.zeros_(m.bias)
model.apply(weight_init)
linear1(x) * sigmoid(linear2(x))x * sigmoid(x) paper (2017)xxxx paper (2018)x * tanh( ln(1 + e^x) ) paper (2019)0.5 * x * ( tanh(x) + 1 )0.5 * x * ( tanh (x+1) + 1)x * ((x+x+1)/(abs(x+1) + abs(x)) * 0.5 + 0.5)class AddCoord2D(torch.nn.Module):
def __init__(self, len):
super(AddCoord2D, self).__init__()
i_coord = torch.linspace(start=1/len, end=1, steps=len).view(len, -1).expand(-1, len)
j_coord = torch.linspace(start=1/len, end=1, steps=len).view(-1, len).expand(len, -1)
self.coords = torch.stack([i_coord, j_coord])
print(self.coords.shape)
def forward(self, x): # X shape: [BS, C, X, Y]
BS = x.shape[0]
return torch.cat((x, self.coords.expand(BS,-1,-1,-1)), dim=1)
During training, some neurons will be deactivated randomly. Hinton, 2012, Srivasta, 2014

Weight penalty: Regularization in loss function (penalice high weights). Weight decay hyper-parameter usually 0.0005.
Visually, the weights only can take a value inside the blue region, and the red circles represent the minimum. Here, there are 2 weight variables.
| L1 (LASSO) | L2 (Ridge) | Elastic Net |
|---|---|---|
![]() | ![]() | ![]() |
| Shrinks coefficients to 0. Good for variable selection | Most used. Makes coefficients smaller | Tradeoff between variable selection and small coefficients |
| Penalizes the sum of absolute weights | Penalizes the sum of squared weights | Combination of 2 before |
loss + wd * weights.abs().sum() | loss + wd * weights.pow(2).sum() |
At training and inference, some connections (weights) will be deactivated permanently. LeCun, 2013. This is very useful at the firsts layers.

Knowledge Distillation (teacher-student) A teacher model teach a student model.
mean(GT - pred) It could determine if the model has positive bias or negative bias.mean(|GT - pred|) The most simple.mean((GT-pred)²) Penalice large errors more than MAE. Most usedsqrt(MSE) Proportional to MSE. Value closer to MAE.nn.CrossEntropyLoss.
nn.NLLLoss()nn.BCELossnn.HingeEmbeddingLoss()-(1-p)^gamma * log(p) paper(Pred ∩ GT)/(Pred ∪ GT) = TP / TP + FP * FN2 * (Pred ∩ GT)/(Pred + GT) = 2·TP / 2·TP + FP * FN
0 (worst) to 1 (best)1 − DiceSmooth the one-hot target label.
LabelSmoothingCrossEntropy(eps:float=0.1, reduction='mean')
Referennce
Dataset with 5 disease images and 20 normal images. If the model predicts all images to be normal, its accuracy is 80%, and F1-score of such a model is 0.88
TP + TN / TP + TN + FP + FN2 * (Prec*Rec)/(Prec+Rec)
TP / TP + FP = TP / predicted possitivesTP / TP + FN = TP / actual possitives2 * (Pred ∩ GT)/(Pred + GT)How big the steps are during training.
lr_find())Number of samples to learn simultaneously.
Batch size = 1: Train each sample individually. (Online gradient descent) ❌Batch size = length(dataset): Train the whole dataset at once, as a batch. (Batch gradient descent) ❌Batch size = number: Train disjoint groups of samples (Mini-batch gradient descent). ✅
32 or 64 are good values.4: Lot of updates. Very noisy random updates in the net (bad).512 Few updates. Very general common updates (bad).
Some people are tring to make a batch size finder according to this paper.
Times to learn the whole dataset.
def seed_everything(seed):
os.environ['PYTHONHASHSEED'] = str(seed)
random.seed(seed) # Random
np.random.seed(seed) # Numpy
torch.manual_seed(seed) # Pytorch
torch.cuda.manual_seed(seed)
torch.backends.cudnn.deterministic = True
torch.backends.cudnn.benchmark = False
#tf.random.set_seed(seed) # Tensorflow
Read this
def clean_mem():
gc.collect()
torch.cuda.empty_cache()
learn.to_parallel()
Reference
learn.to_fp16()
learn.to_fp32()
Reference
import numpy as np
import torch
from torchvision import models
import torchvision.transforms as transforms
from PIL import Image
from flask import Flask, jsonify, request
import json
app = Flask(__name__)
app.config['JSON_SORT_KEYS'] = False
classes = json.load(open('imagenet_classes.json'))
model = models.densenet121(pretrained=True)
model.eval()
def pre_process(image_file):
my_transforms = transforms.Compose([transforms.Resize(255),
transforms.CenterCrop(224),
transforms.ToTensor(),
transforms.Normalize([0.485, 0.456, 0.406], [0.229, 0.224, 0.225])])
image = Image.open(image_file)
return my_transforms(image).unsqueeze(0) # unsqueeze is for the BS dim
def post_process(logits):
vals, idxs = logits.softmax(1).topk(5)
vals = vals[0].numpy()
idxs = idxs[0].numpy()
result = {}
for idx, val in zip(idxs, vals):
result[classes[idx]] = round(float(val), 4)
return result
def get_prediction(image_file):
with torch.no_grad():
image_tensor = pre_process(image_file)
output = model.forward(image_tensor)
return post_process(output)
@app.route('/predict', methods=['POST'])
def predict():
if request.method == 'POST':
image_file = request.files['my_img_file']
result_dict = get_prediction(image_file)
#return jsonify(result_dict)
#return json.dumps(result_dict)
return result_dict
if __name__ == '__main__':
app.run()
FLASK_ENV=development FLASK_APP=app.py flask run
curl -X POST -F my_img_file=@cardigan.jpg http://localhost:5000/predict
import requests
resp = requests.post("http://localhost:5000/predict",
files={"my_img_file": open('cardigan.jpg','rb')})
print(resp.json())
{
"cardigan": 0.7083,
"wool": 0.0837,
"suit": 0.0431,
"Windsor_tie": 0.031,
"trench_coat": 0.0307
}
| What | Accuracy | Pytorch API | |
|---|---|---|---|
| Dynamic Quantization | Weights only | Good | qmodel = torch.quantization.quantize_dynamic(model, dtype=torch.qint8) |
| Post Training Quantization | Weights and activations | Good | model.qconfig = torch.quantization.default_qconfig torch.quantization.prepare(model, inplace=True) torch.quantization.convert(model, inplace=True) |
| Quantization-Aware Training | Weights and activations | Best | torch.quantization.prepare_qat -> torch.quantization.convert |
Reference
import torch.nn.utils.prune as prune
parameters_to_prune = (
(model.conv1, 'weight'),
(model.conv2, 'weight'),
(model.fc1, 'weight'),
(model.fc2, 'weight'),
(model.fc3, 'weight'),
)
prune.global_unstructured(
parameters_to_prune,
pruning_method=prune.L1Unstructured,
amount=0.2,
)
class ThresholdPruning(prune.BasePruningMethod):
PRUNING_TYPE = "unstructured"
def __init__(self, threshold): self.threshold = threshold
def compute_mask(self, tensor, default_mask): return torch.abs(tensor) > self.threshold
prune.global_unstructured(
parameters_to_prune,
pruning_method=ThresholdPruning,
threshold=0.01
)
def pruned_info(model):
print("Weights pruned:")
print("==============")
total_pruned, total_weights = 0,0
for name, chil in model.named_children():
layer_pruned = torch.sum(chil.weight == 0)
layer_weights = chil.weight.nelement()
total_pruned += layer_pruned
total_weights += layer_weights
print(name, "\t{:.2f}%".format(100 * float(layer_pruned)/ float(layer_weights)))
print("==============")
print("Total\t{:.2f}%".format(100 * float(total_pruned)/ float(total_weights)))
# Weights pruned:
# ==============
# conv1 1.85%
# conv2 8.10%
# fc1 19.76%
# fc2 10.66%
# fc3 9.40%
# ==============
# Total 17.90%
Iterative magnitude pruning is iterative process of removing connections (Prune/Train/Repeat):

At the end you can have pruned the 15%, 30%, 45%, 60%, 75%, and 90% of your original model.
Reference
- Code:
- Papers:
- Deep Compression (2015)
- Train Large, Then Compress (2020)
- Neural Networks are Surprisingly Modular (2020)
torch_script = torch.jit.script(MyModel())
torch_script.save("my_model_script.pt")
Reference
torch.onnx.export(model, img, f, verbose=False, opset_version=11) # Export to onnx
# Check onnx model
import onnx
model = onnx.load(f) # load onnx model
onnx.checker.check_model(model) # check onnx model
print(onnx.helper.printable_graph(model.graph)) # print a human readable representation of the graph
print('Export complete. ONNX model saved to %s\nView with https://github.com/lutzroeder/netron' % f)
Reference
0.50.0005
0.01 or 0.1.wd * w. Sometimes mathematically identical to L2 reg.0.632 Prob of sample in Out Of Bag 0.368Other tricks:
- Label Smoothing: Smooth the one-hot target label
- Knowledge Distillation: A bigger trained net (teacher) helps the network paper
loss = recontruction loss + latent loss
- 2D: [x,y]->[R,G,B]
- 3D: [x,y,z]->[R,G,B,alpha]
| Description | Website | Video | Paper |
|---|---|---|---|
| NeRF in the Wild | web | 3:41 | Aug 2020 |
| NeRF++ | Oct 2020 | ||
| Deformable NeRF (nerfies) | web | 7:26 | Nov 2020 |
| NeRF with time dimension | web | 2:21 | Nov 2020 |
| NeRF with better weight init | web | 3:54 | Dec 2020 |
Check this kaggle discussion
Reinforcement learning reference
Projections (BAD REPRESENTATION) (complicated things with voxels) Dense matrix (antor) - Its a depth map i think - Not projections - NAtive output of the sensor but condensed in a dense matrix
TODO
- Multi-Task Learning: Train a model on a variety of learning tasks
- Meta-learning: Learn new tasks with minimal data using prior knowledge.
- N-Shot Learning
- Zero-shot: 0 trainning examples of that class.
- One-shot: 1 trainning example of that class.
- Few-shot: 2...5 trainning examples of that class.
- Models
- Naive approach: re-training the model on the new data, would severely overfit.
- Siamese Networks (2015) Knows if to inputs are the same or not. (2 Feature extraction shares wights)
- Matching Networks (2016) Weighted nearest-neighbor classifier applied within an embedding space.
- Model-Agnostic Meta-Learning (MAML) (2017)
- Prototypical Networks (2017): Better nearest-neighbor classifier of embeddings.
- Meta-Learning for Semi-Supervised classification (2018) Extensions of Prototypical Networks. SotA.
- Meta-Transfer Learning (MTL) (2018)
- Online Meta-Learning (2019)
- Neural Turing machine. paper, code
- Neural Arithmetic Logic Units (NALU) paper
- Remember the math:
- Matrix calculus
- Einsum: link 1, link 2
nvidia-smi daemon: Check that sm% is near to 100% for a good GPU usage.
Jupyter Notebook
99.9%