Advanced deep learning learning techniques, layers, activations loss functions, all in keras / tensorflow
Python
26
6,603 commits
updated Oct 3, 2026
In the rapidly evolving landscape of AI research, groundbreaking techniques are often scattered across countless repositories, disparate implementations, and dense academic papers. dl_techniques emerges as a unified, curated, and production-ready arsenal for advanced deep learning. Pioneered and sponsored by Electi Consulting, this library is more than a collection of components—it is the definitive toolkit for researchers and engineers to design, train, and dissect state-of-the-art neural networks.
We bridge the chasm between theoretical innovation and practical application, providing faithful, efficient, and enterprise-validated components. From next-generation attention mechanisms and graph neural networks to information-theoretic loss functions and an unparalleled model analysis suite, dl_techniques is your strategic advantage for pushing the boundaries of deep learning.
dl_techniques?: Your Unfair Advantage in AI R&D.This library is a comprehensive suite of tools organized into five key pillars, developed through rigorous research and validated in real-world enterprise applications:
BERT, Gemma 3, Qwen3 (dense, Qwen3Next and embeddings), Mamba, ModernBERT, DistilBERT and GPT-2, alongside the Byte Latent Transformer (BLT), the Hierarchical Reasoning Model (HRM), the Tiny Recursive Model, Tree Transformer, late-interaction retrieval with ColBERT (v1 and v2), and Fourier token mixing in FNet and FFTNet.CLIP, MobileCLIP v1 and v2 in one package (v2 is the faithful port, on a faithful FastViT image tower; v1 is deliberately non-faithful on the image side and is kept, not deprecated), FastVLM (vision-only despite the name), NanoVLM, DINOv1/v2/v3, ViT-SigLIP (a ViT with a two-stage conv patch stem, not SigLIP's sigmoid contrastive objective), BEiT, Swin Transformer, MAE, Video-JEPA, the latent-energy world model LEWM, object detection with DETR and YOLOv12, keypoints with SuperPoint, and Segment Anything (SAM 1, 2 and 3). Three packages are named for something they are not; each is listed with its measured correction in src/dl_techniques/models/README.md § Names that misattribute.ConvNeXtV1/V2, ConvUNeXt, MobileNetV1-V4, ResNet, the recursively-defined FractalNet, the complex shearlet-based CoShNet, attention-augmented CBAM, and the ultra-efficient SqueezeNet family.TiRex with quantile prediction, an enhanced implementation of N-BEATS, autoregressive DeepAR, PRISM, and the novel xLSTM.SD3 MMDiT, Ideogram4), a complete Variational Autoencoder (VAE) framework with VQ-VAE variants (including rotation-based codebook updates), arbitrary-scale super-resolution (THERA, PFT-SR), and restoration/denoising backbones (DarkIR, SCUNet, ACC-UNet, PW-FNet, and a family of bias-free denoisers).DepthAnything for monocular depth estimation, full Capsule Networks (CapsNet) with dynamic routing, TabM for tabular data, model-agnostic inference-time power_sampling for any causal LM/VLM, and unsupervised aligners like Mini-Vec2Vec.Graph Neural Networks (GNNs) in RELGT and the Simplified Hyperbolic GCN SHGCN, the Energy Transformer family (image and graph domains), geometric-algebra networks in CliffordNet, Kolmogorov-Arnold Networks (KAN) and PowerMLP, external-memory computers (NTM, the Neural Arithmetic Module NAM), Self-Organizing Maps, point-cloud latent_gmm_registration, and bio-mimetic models like MothNet.
DifferentialMultiHeadAttention, modern HopfieldAttention, GroupQueryAttention, CapsuleRoutingAttention, and efficient alternatives like FNetFourierTransform and RingAttention.BandRMS, LogitNorm), and 21 Feed-Forward Network (FFN) types (SwiGLU, GeGLU, OrthoGLU) with a single line of code — plus registries for 22 activations, 13 embeddings, and task heads dispatched by domain via create_head.GCN, GAT, GraphSAGE), Relational Graph Transformer (RELGT) blocks, and Entity-Graph Refinement for learning hierarchical relationships.Mixture Density Networks (MDN), Normalizing Flows, and time series analysis layers for residual autocorrelation (ResidualACFLayer).Rotary Position Embedding (RoPE), and its modern variants like DualRotaryPositionEmbedding.
ModelAnalyzer to benchmark models across five critical dimensions: training dynamics, weight health, prediction calibration, information flow, and advanced spectral analysis. Its modular design includes specialized analyzers like CalibrationAnalyzer, WeightAnalyzer, and InformationFlowAnalyzer.SpectralAnalyzer to assess generalization potential by analyzing the spectral properties (eigenvalues) of weight matrices—often without needing test data.CalibrationAnalyzer), information bottlenecks (InformationFlowAnalyzer), weight decay and similarity (WeightAnalyzer), and learning efficiency (TrainingDynamicsAnalyzer).
AnyLoss: A groundbreaking framework that transforms any confusion-matrix-based metric (e.g., F1-score, Balanced Accuracy, Matthews Correlation Coefficient) into a differentiable loss function for direct optimization on imbalanced data.GoodhartAwareLoss (cross-entropy plus a per-sample confidence penalty, with an optional anti-collapse term), calibration-focused losses like BrierScoreLoss, the uncertainty-aware FocalUncertaintyLoss, and DINO's self-distillation loss.CLIPContrastiveLoss, SigLIPLoss), segmentation (Dice, Focal, Tversky), time series (MASELoss, SMAPELoss), and generative modeling (WassersteinLoss with gradient penalty).WarmupSchedule, utilities for DeepSupervision in multi-scale architectures, and a suite of advanced regularizers (SoftOrthogonal, SRIP).
train_*.py entry points across 47 trainer directories (src/train/, plus the shared src/train/common/ library), establishing standardized and reproducible workflows for training, validation, and testing across domains like NLP, Vision, and Time Series.VisualizationManager), and enhanced model serialization with custom object support.tests/) ensures the correctness and stability of every component, with dedicated fixtures for mixed-precision and TF32-sensitive regressions. It mirrors src/dl_techniques/ directory-for-directory, with the deliberate exception of tests/test_models/, which stays flat rather than following the model family nesting.src/dl_techniques/layers/fastvit/reference.py) rather than asserting parity in prose.dl_techniques?dl_techniques adheres to the highest standards of software engineering, ensuring it's easy to use, extend, and maintain.Note: This library requires Python 3.11+ and Keras 3.8.0 with the TensorFlow 2.18.0 backend.
Clone the repository:
git clone https://github.com/nikolasmarkou/dl_techniques.git
cd dl_techniques
Install dependencies: For standard usage, install the library and its core dependencies:
pip install .
Editable Install (for developers): If you plan to contribute or modify the library's source code, install it in editable mode with development tools:
pip install -e ".[dev]"
This installs additional tools such as pytest, pytest-cov, pylint, and pre-commit.
Optional Extras:
The Streamlit front-end under src/applications/bias_free_denoiser/ needs an extra that is not part of a core install:
pip install ".[apps]"
The tokenizer utilities and the HuggingFace/tensorflow-datasets-backed loaders under dl_techniques.datasets need their own extra:
pip install ".[data]"
Verify Installation:
python -c "import dl_techniques; print('Installation successful!')"
make test # full pytest suite (~1.5 hours; also the pre-push hook)
make clean # remove build artifacts and __pycache__
make structure # print the src/ tree
Scope pytest to what you changed rather than running the whole suite as a regression check:
pytest tests/test_layers/test_norms/ -vvv
Training scripts need a non-interactive matplotlib backend on headless machines:
MPLBACKEND=Agg python -m train.<pipeline>.train_<script> [args]
No trained weights ship with this repository. Training outputs live under
results/, which is gitignored, so a fresh clone contains source, tests and prose only. Anyresults/<run>/final_model.keraspath quoted in a trainer default, a paper or an application is a local artifact — you train it yourself first. SeeREPO_MAP.md§ What a fresh clone actually contains.
Effortlessly construct and experiment with modern transformer components using our unified factory system.
import keras
from dl_techniques.layers.attention.factory import create_attention_layer
from dl_techniques.layers.norms.factory import create_normalization_layer
from dl_techniques.layers.ffn.factory import create_ffn_layer
inputs = keras.Input(shape=(1024, 512))
# Use factories for consistent, validated component creation
attention = create_attention_layer(
'differential',
dim=512,
num_heads=8,
head_dim=64
)
norm = create_normalization_layer('rms_norm', epsilon=1e-6)
ffn = create_ffn_layer('swiglu', output_dim=512)
# Build a modern transformer block
x = attention(inputs)
x = norm(x)
x = ffn(x)
model = keras.Model(inputs, x)
model.summary()
Instantiate a state-of-the-art univariate time series model capable of generating robust, uncertainty-aware forecasts.
from dl_techniques.models.time_series.tirex.model import create_tirex_model
from dl_techniques.losses.quantile_loss import QuantileLoss
# TiRex: probabilistic univariate forecasting with quantile prediction
model = create_tirex_model(
input_length=100, # length of the input context window
prediction_length=24, # forecast horizon
quantile_levels=[0.1, 0.5, 0.9], # 80% prediction interval + median
)
# Train directly for calibrated uncertainty with a quantile loss.
# Predictions are (batch, prediction_length, n_quantiles), so a plain
# point-forecast metric like 'mae' would broadcast across the quantile axis.
model.compile(
optimizer='adamw',
loss=QuantileLoss(quantiles=[0.1, 0.5, 0.9]),
)
Map images and text into a single shared embedding space with a CLIP dual encoder, unlocking zero-shot classification and cross-modal retrieval.
import keras
from dl_techniques.models.vision_language.clip.model import create_clip_variant
# CLIP ViT-B/32: a dual encoder mapping images and text into one space
model = create_clip_variant("ViT-B/32")
model.build({"image": (None, 224, 224, 3), "text": (None, 77)})
images = keras.random.normal((4, 224, 224, 3))
tokens = keras.ops.cast(keras.random.uniform((4, 77), 0, 49408), "int32")
# Full contrastive forward pass
outputs = model({"image": images, "text": tokens}, training=False)
# outputs["image_features"] -> (4, 512), L2-normalised
# outputs["text_features"] -> (4, 512), L2-normalised
# outputs["logits_per_image"] -> (4, 4), similarity matrix
# Or encode each modality independently for retrieval
image_features = model.encode_image(images)
text_features = model.encode_text(tokens)
Go beyond surface-level metrics and gain deep, actionable insights into your models' performance and behavior.
from dl_techniques.analyzer import ModelAnalyzer, AnalysisConfig, DataInput
# Compare multiple models across a suite of deep diagnostics
models = {'TiRex_Model': tirex_model, 'Baseline_LSTM': lstm_model}
histories = {'TiRex_Model': tirex_history, 'Baseline_LSTM': lstm_history}
test_data = DataInput(x_test, y_test)
# Configure a comprehensive analysis run
config = AnalysisConfig(
analyze_training_dynamics=True,
analyze_calibration=True,
analyze_weights=True,
analyze_spectral=True, # Unleash spectral analysis for generalization insights
save_plots=True,
plot_style='publication'
)
analyzer = ModelAnalyzer(models, config=config, training_history=histories)
# Execute the complete analysis and generate a full suite of visualizations
results = analyzer.analyze(test_data)
# Access detailed, structured metrics programmatically for automated reporting
print(f"Best calibrated model (by ECE): {min(results.calibration_metrics.items(), key=lambda x: x[1]['ece'])}")
print(f"Training efficiency ranking (epochs to converge): {results.training_metrics.epochs_to_convergence}")
Stop tuning class weights and start optimizing your target metric directly with the AnyLoss framework.
from dl_techniques.losses.any_loss import F1Loss, BalancedAccuracyLoss
# For imbalanced datasets, optimize F1-score directly
model.compile(
optimizer='adamw',
loss=F1Loss(from_logits=True),
metrics=['accuracy', 'precision', 'recall']
)
# Alternatively, optimize for balanced accuracy
model.compile(
optimizer='adamw',
loss=BalancedAccuracyLoss(from_logits=True),
metrics=['accuracy']
)
Unlock insights from graph-structured data with our powerful and configurable GNN implementations.
import keras
from dl_techniques.layers.graphs.graph_neural_network import GraphNeuralNetworkLayer
# Create a Graph Attention Network (GAT) to process relational data
gnn = GraphNeuralNetworkLayer(
concept_dim=256,
num_layers=3,
message_passing='gat', # Use attention for message passing
aggregation='attention',
dropout_rate=0.1
)
# Apply the layer to graph-structured inputs
node_features = keras.Input(shape=(None, 256)) # Variable number of nodes
adjacency_matrix = keras.Input(shape=(None, None))
node_embeddings = gnn([node_features, adjacency_matrix])
This library is engineered to be a living knowledge base, bridging the gap between academia and industry. Our documentation is split into three primary resources: a navigational map of the repository, in-depth theoretical guides, and a comprehensive API reference.
REPO_MAP.mdBefore anything else, read REPO_MAP.md — a path-verified router for the repository. It answers where code of a given kind lives, how the registry and factory dispatch is wired, which trainer trains which model, and which of the many in-tree docs answers your question.
research/)Our research/ directory contains over 120 articles providing the theoretical foundations, implementation details, and best practices behind key components. Highlights include:
research/papers/)Five LaTeX manuscripts written against this codebase live under research/papers/, several with built PDFs: band_rms (band-constrained RMS normalization), bfunet (bias-free denoisers as image priors), cliffordnet_extensions, correlations, and logical_net.
There is no committed documentation directory and no doc generator — generate_docs.py and the make docs target were deleted as deprecated. For detailed documentation on every module, class, and function, browse the source tree directly: each subpackage ships a focused README.md (e.g. src/dl_techniques/analyzer/README.md) and a per-package AGENTS.md describing its conventions, patterns, and components. Every one of the 84 leaf model packages carries its own README.md. The research/ guides above complement these with the underlying theory.
The repository is organized for clarity, maintainability, and ease of contribution.
src/ is the import root: the library is imported as dl_techniques.*, the training pipelines as train.*, and the applications as applications.*.
dl_techniques/
├── src/dl_techniques/ # THE LIBRARY — 13 subpackages
│ ├── models/ # 80 leaf model packages in 11 FAMILY directories
│ │ │ # full catalogue: src/dl_techniques/models/README.md
│ │ ├── vision/ # 35 — resnet, convnext, vit, dino, swin_transformer, beit,
│ │ │ # yolo12, detr, vae, vq_vae, depth_anything, and the
│ │ │ # nested image_restoration/, super_resolution/, keypoints/
│ │ ├── language/ # 17 — bert, gemma, qwen, mamba, modern_bert, gpt2, colbert,
│ │ │ # byte_latent_transformer, hierarchical_reasoning_model, ...
│ │ ├── vision_language/ # 9 — clip, mobile_clip, fastvlm, nano_vlm, sd3_mmdit,
│ │ │ # ideogram4, and the nested sam/{sam1,sam2,sam3}
│ │ ├── time_series/ # 7 — tirex, nbeats, xlstm, deepar, prism, mdn, adaptive_ema
│ │ ├── general_purpose/ # 3 — kan, mothnet, power_mlp
│ │ ├── graph/ # 3 — relgt, graph_energy_transformer, shgcn
│ │ ├── neural_computer/ # 2 — ntm, nam
│ │ ├── common/ # 1 — power_sampling (model-agnostic inference machinery)
│ │ ├── memory/ # 1 — som
│ │ ├── point_cloud/ # 1 — latent_gmm_registration
│ │ └── tabular/ # 1 — tabm
│ ├── layers/ # 275 modules — 200 in 21 themed subpackages,
│ │ │ # 75 loose at the top level
│ │ ├── attention/ # 33 registered attention mechanisms (factory)
│ │ ├── ffn/ # 21 registered feed-forward networks (factory)
│ │ ├── activations/ # 22 registered activations (factory)
│ │ ├── norms/ # 18 registered normalization layers (factory)
│ │ ├── embedding/ # 13 registered positional/semantic embeddings (factory)
│ │ ├── transformers/ # Assembled transformer/encoder/decoder blocks
│ │ ├── heads/ # Task heads dispatched by domain (nlp/vision/vlm)
│ │ ├── time_series/ # Forecasting-specific layers
│ │ ├── fastvit/ # FastViT primitives (+ a committed reference impl)
│ │ ├── graphs/ # Graph neural network components
│ │ ├── moe/ # Mixture of Experts (MoE) system
│ │ ├── memory/ # External / associative memory layers
│ │ ├── statistics/ # Statistical and probabilistic layers (MDN, Flows)
│ │ └── ... # fusion, geometric, logic, mixtures, physics,
│ │ # reasoning, sequence_pooling, tokenizers
│ ├── losses/ # 44 specialized loss modules (AnyLoss, Goodhart, etc.)
│ ├── metrics/ # Custom Keras metrics (PSNR, SSIM, perplexity, Brier)
│ ├── optimization/ # Optimizers (Muon, VSGD, SGLD, Gefen, WW-PGD), LR schedules
│ ├── analyzer/ # Comprehensive model analysis toolkit and visualizers
│ ├── visualization/ # Plotting helpers for training and evaluation
│ ├── datasets/ # Dataset loaders and synthetic generators
│ ├── callbacks/ # Reusable Keras callbacks
│ ├── regularizers/ # Advanced regularization techniques (SRIP, Orthogonal)
│ ├── initializers/ # Structured initializers (Gabor, Haar, orthonormal, KAN)
│ ├── constraints/ # Weight constraints
│ └── utils/ # Core utilities, loggers, masking, and data handlers
├── src/train/ # 70 train_*.py entry points across 47 trainer directories
├── src/applications/ # Deployable apps (today: bias_free_denoiser, Streamlit)
├── research/ # 120+ in-depth articles, guides, and LaTeX papers
├── tests/ # 895 test modules mirroring src/dl_techniques/
└── REPO_MAP.md # Path-verified router — read this first
We welcome contributions from the research community. Whether you are implementing a new technique, improving documentation, or fixing bugs, your input is valuable.
pip install -e ".[dev]".git checkout -b feature/new-technique.dl_techniques.utils.logger (no print)..venv environment and write comprehensive tests using pytest, scoped to the modules you change (make test runs the full ~1.5h suite). Set MPLBACKEND=Agg when running training scripts on headless machines.research/ directory. Docstring style is not uniform across this repo — layers/ is predominantly Sphinx/reST (:param:), models/ is measurably mixed with the Sphinx exemplar models/language/bert/model.py as the model for new packages, and losses//metrics//utils//optimization//analyzer//visualization/ are Google-Args:-majority. Match the package you are editing and never convert a file wholesale; the measured counts and the greps that re-derive them are in src/dl_techniques/AGENTS.md § Core Conventions → Code Style.ModelAnalyzer toolkit.This project is licensed under the GNU General Public License v3.0.
Important Considerations:
See the LICENSE file for complete details.
This library is proudly sponsored and pioneered by Electi Consulting, a premier AI consultancy specializing in enterprise artificial intelligence, blockchain technology, and cryptographic solutions. The practical validation and enterprise-ready nature of these components has been made possible through Electi's extensive experience deploying state-of-the-art AI solutions across diverse industries:
Special recognition is extended to the open-source community and the many researchers whose groundbreaking work forms the foundation of this library.
This library stands on the shoulders of giants. Our implementations are grounded in a rigorous study of the source papers that have defined the field of modern deep learning.
Complete bibliographic information is available in the documentation for individual modules.
Python
99.1%
Advanced deep learning learning techniques, layers, activations loss functions, all in keras / tensorflow
Python
26
6,603 commits
updated Oct 3, 2026
In the rapidly evolving landscape of AI research, groundbreaking techniques are often scattered across countless repositories, disparate implementations, and dense academic papers. dl_techniques emerges as a unified, curated, and production-ready arsenal for advanced deep learning. Pioneered and sponsored by Electi Consulting, this library is more than a collection of components—it is the definitive toolkit for researchers and engineers to design, train, and dissect state-of-the-art neural networks.
We bridge the chasm between theoretical innovation and practical application, providing faithful, efficient, and enterprise-validated components. From next-generation attention mechanisms and graph neural networks to information-theoretic loss functions and an unparalleled model analysis suite, dl_techniques is your strategic advantage for pushing the boundaries of deep learning.
dl_techniques?: Your Unfair Advantage in AI R&D.This library is a comprehensive suite of tools organized into five key pillars, developed through rigorous research and validated in real-world enterprise applications:
BERT, Gemma 3, Qwen3 (dense, Qwen3Next and embeddings), Mamba, ModernBERT, DistilBERT and GPT-2, alongside the Byte Latent Transformer (BLT), the Hierarchical Reasoning Model (HRM), the Tiny Recursive Model, Tree Transformer, late-interaction retrieval with ColBERT (v1 and v2), and Fourier token mixing in FNet and FFTNet.CLIP, MobileCLIP v1 and v2 in one package (v2 is the faithful port, on a faithful FastViT image tower; v1 is deliberately non-faithful on the image side and is kept, not deprecated), FastVLM (vision-only despite the name), NanoVLM, DINOv1/v2/v3, ViT-SigLIP (a ViT with a two-stage conv patch stem, not SigLIP's sigmoid contrastive objective), BEiT, Swin Transformer, MAE, Video-JEPA, the latent-energy world model LEWM, object detection with DETR and YOLOv12, keypoints with SuperPoint, and Segment Anything (SAM 1, 2 and 3). Three packages are named for something they are not; each is listed with its measured correction in src/dl_techniques/models/README.md § Names that misattribute.ConvNeXtV1/V2, ConvUNeXt, MobileNetV1-V4, ResNet, the recursively-defined FractalNet, the complex shearlet-based CoShNet, attention-augmented CBAM, and the ultra-efficient SqueezeNet family.TiRex with quantile prediction, an enhanced implementation of N-BEATS, autoregressive DeepAR, PRISM, and the novel xLSTM.SD3 MMDiT, Ideogram4), a complete Variational Autoencoder (VAE) framework with VQ-VAE variants (including rotation-based codebook updates), arbitrary-scale super-resolution (THERA, PFT-SR), and restoration/denoising backbones (DarkIR, SCUNet, ACC-UNet, PW-FNet, and a family of bias-free denoisers).DepthAnything for monocular depth estimation, full Capsule Networks (CapsNet) with dynamic routing, TabM for tabular data, model-agnostic inference-time power_sampling for any causal LM/VLM, and unsupervised aligners like Mini-Vec2Vec.Graph Neural Networks (GNNs) in RELGT and the Simplified Hyperbolic GCN SHGCN, the Energy Transformer family (image and graph domains), geometric-algebra networks in CliffordNet, Kolmogorov-Arnold Networks (KAN) and PowerMLP, external-memory computers (NTM, the Neural Arithmetic Module NAM), Self-Organizing Maps, point-cloud latent_gmm_registration, and bio-mimetic models like MothNet.
DifferentialMultiHeadAttention, modern HopfieldAttention, GroupQueryAttention, CapsuleRoutingAttention, and efficient alternatives like FNetFourierTransform and RingAttention.BandRMS, LogitNorm), and 21 Feed-Forward Network (FFN) types (SwiGLU, GeGLU, OrthoGLU) with a single line of code — plus registries for 22 activations, 13 embeddings, and task heads dispatched by domain via create_head.GCN, GAT, GraphSAGE), Relational Graph Transformer (RELGT) blocks, and Entity-Graph Refinement for learning hierarchical relationships.Mixture Density Networks (MDN), Normalizing Flows, and time series analysis layers for residual autocorrelation (ResidualACFLayer).Rotary Position Embedding (RoPE), and its modern variants like DualRotaryPositionEmbedding.
ModelAnalyzer to benchmark models across five critical dimensions: training dynamics, weight health, prediction calibration, information flow, and advanced spectral analysis. Its modular design includes specialized analyzers like CalibrationAnalyzer, WeightAnalyzer, and InformationFlowAnalyzer.SpectralAnalyzer to assess generalization potential by analyzing the spectral properties (eigenvalues) of weight matrices—often without needing test data.CalibrationAnalyzer), information bottlenecks (InformationFlowAnalyzer), weight decay and similarity (WeightAnalyzer), and learning efficiency (TrainingDynamicsAnalyzer).
AnyLoss: A groundbreaking framework that transforms any confusion-matrix-based metric (e.g., F1-score, Balanced Accuracy, Matthews Correlation Coefficient) into a differentiable loss function for direct optimization on imbalanced data.GoodhartAwareLoss (cross-entropy plus a per-sample confidence penalty, with an optional anti-collapse term), calibration-focused losses like BrierScoreLoss, the uncertainty-aware FocalUncertaintyLoss, and DINO's self-distillation loss.CLIPContrastiveLoss, SigLIPLoss), segmentation (Dice, Focal, Tversky), time series (MASELoss, SMAPELoss), and generative modeling (WassersteinLoss with gradient penalty).WarmupSchedule, utilities for DeepSupervision in multi-scale architectures, and a suite of advanced regularizers (SoftOrthogonal, SRIP).
train_*.py entry points across 47 trainer directories (src/train/, plus the shared src/train/common/ library), establishing standardized and reproducible workflows for training, validation, and testing across domains like NLP, Vision, and Time Series.VisualizationManager), and enhanced model serialization with custom object support.tests/) ensures the correctness and stability of every component, with dedicated fixtures for mixed-precision and TF32-sensitive regressions. It mirrors src/dl_techniques/ directory-for-directory, with the deliberate exception of tests/test_models/, which stays flat rather than following the model family nesting.src/dl_techniques/layers/fastvit/reference.py) rather than asserting parity in prose.dl_techniques?dl_techniques adheres to the highest standards of software engineering, ensuring it's easy to use, extend, and maintain.Note: This library requires Python 3.11+ and Keras 3.8.0 with the TensorFlow 2.18.0 backend.
Clone the repository:
git clone https://github.com/nikolasmarkou/dl_techniques.git
cd dl_techniques
Install dependencies: For standard usage, install the library and its core dependencies:
pip install .
Editable Install (for developers): If you plan to contribute or modify the library's source code, install it in editable mode with development tools:
pip install -e ".[dev]"
This installs additional tools such as pytest, pytest-cov, pylint, and pre-commit.
Optional Extras:
The Streamlit front-end under src/applications/bias_free_denoiser/ needs an extra that is not part of a core install:
pip install ".[apps]"
The tokenizer utilities and the HuggingFace/tensorflow-datasets-backed loaders under dl_techniques.datasets need their own extra:
pip install ".[data]"
Verify Installation:
python -c "import dl_techniques; print('Installation successful!')"
make test # full pytest suite (~1.5 hours; also the pre-push hook)
make clean # remove build artifacts and __pycache__
make structure # print the src/ tree
Scope pytest to what you changed rather than running the whole suite as a regression check:
pytest tests/test_layers/test_norms/ -vvv
Training scripts need a non-interactive matplotlib backend on headless machines:
MPLBACKEND=Agg python -m train.<pipeline>.train_<script> [args]
No trained weights ship with this repository. Training outputs live under
results/, which is gitignored, so a fresh clone contains source, tests and prose only. Anyresults/<run>/final_model.keraspath quoted in a trainer default, a paper or an application is a local artifact — you train it yourself first. SeeREPO_MAP.md§ What a fresh clone actually contains.
Effortlessly construct and experiment with modern transformer components using our unified factory system.
import keras
from dl_techniques.layers.attention.factory import create_attention_layer
from dl_techniques.layers.norms.factory import create_normalization_layer
from dl_techniques.layers.ffn.factory import create_ffn_layer
inputs = keras.Input(shape=(1024, 512))
# Use factories for consistent, validated component creation
attention = create_attention_layer(
'differential',
dim=512,
num_heads=8,
head_dim=64
)
norm = create_normalization_layer('rms_norm', epsilon=1e-6)
ffn = create_ffn_layer('swiglu', output_dim=512)
# Build a modern transformer block
x = attention(inputs)
x = norm(x)
x = ffn(x)
model = keras.Model(inputs, x)
model.summary()
Instantiate a state-of-the-art univariate time series model capable of generating robust, uncertainty-aware forecasts.
from dl_techniques.models.time_series.tirex.model import create_tirex_model
from dl_techniques.losses.quantile_loss import QuantileLoss
# TiRex: probabilistic univariate forecasting with quantile prediction
model = create_tirex_model(
input_length=100, # length of the input context window
prediction_length=24, # forecast horizon
quantile_levels=[0.1, 0.5, 0.9], # 80% prediction interval + median
)
# Train directly for calibrated uncertainty with a quantile loss.
# Predictions are (batch, prediction_length, n_quantiles), so a plain
# point-forecast metric like 'mae' would broadcast across the quantile axis.
model.compile(
optimizer='adamw',
loss=QuantileLoss(quantiles=[0.1, 0.5, 0.9]),
)
Map images and text into a single shared embedding space with a CLIP dual encoder, unlocking zero-shot classification and cross-modal retrieval.
import keras
from dl_techniques.models.vision_language.clip.model import create_clip_variant
# CLIP ViT-B/32: a dual encoder mapping images and text into one space
model = create_clip_variant("ViT-B/32")
model.build({"image": (None, 224, 224, 3), "text": (None, 77)})
images = keras.random.normal((4, 224, 224, 3))
tokens = keras.ops.cast(keras.random.uniform((4, 77), 0, 49408), "int32")
# Full contrastive forward pass
outputs = model({"image": images, "text": tokens}, training=False)
# outputs["image_features"] -> (4, 512), L2-normalised
# outputs["text_features"] -> (4, 512), L2-normalised
# outputs["logits_per_image"] -> (4, 4), similarity matrix
# Or encode each modality independently for retrieval
image_features = model.encode_image(images)
text_features = model.encode_text(tokens)
Go beyond surface-level metrics and gain deep, actionable insights into your models' performance and behavior.
from dl_techniques.analyzer import ModelAnalyzer, AnalysisConfig, DataInput
# Compare multiple models across a suite of deep diagnostics
models = {'TiRex_Model': tirex_model, 'Baseline_LSTM': lstm_model}
histories = {'TiRex_Model': tirex_history, 'Baseline_LSTM': lstm_history}
test_data = DataInput(x_test, y_test)
# Configure a comprehensive analysis run
config = AnalysisConfig(
analyze_training_dynamics=True,
analyze_calibration=True,
analyze_weights=True,
analyze_spectral=True, # Unleash spectral analysis for generalization insights
save_plots=True,
plot_style='publication'
)
analyzer = ModelAnalyzer(models, config=config, training_history=histories)
# Execute the complete analysis and generate a full suite of visualizations
results = analyzer.analyze(test_data)
# Access detailed, structured metrics programmatically for automated reporting
print(f"Best calibrated model (by ECE): {min(results.calibration_metrics.items(), key=lambda x: x[1]['ece'])}")
print(f"Training efficiency ranking (epochs to converge): {results.training_metrics.epochs_to_convergence}")
Stop tuning class weights and start optimizing your target metric directly with the AnyLoss framework.
from dl_techniques.losses.any_loss import F1Loss, BalancedAccuracyLoss
# For imbalanced datasets, optimize F1-score directly
model.compile(
optimizer='adamw',
loss=F1Loss(from_logits=True),
metrics=['accuracy', 'precision', 'recall']
)
# Alternatively, optimize for balanced accuracy
model.compile(
optimizer='adamw',
loss=BalancedAccuracyLoss(from_logits=True),
metrics=['accuracy']
)
Unlock insights from graph-structured data with our powerful and configurable GNN implementations.
import keras
from dl_techniques.layers.graphs.graph_neural_network import GraphNeuralNetworkLayer
# Create a Graph Attention Network (GAT) to process relational data
gnn = GraphNeuralNetworkLayer(
concept_dim=256,
num_layers=3,
message_passing='gat', # Use attention for message passing
aggregation='attention',
dropout_rate=0.1
)
# Apply the layer to graph-structured inputs
node_features = keras.Input(shape=(None, 256)) # Variable number of nodes
adjacency_matrix = keras.Input(shape=(None, None))
node_embeddings = gnn([node_features, adjacency_matrix])
This library is engineered to be a living knowledge base, bridging the gap between academia and industry. Our documentation is split into three primary resources: a navigational map of the repository, in-depth theoretical guides, and a comprehensive API reference.
REPO_MAP.mdBefore anything else, read REPO_MAP.md — a path-verified router for the repository. It answers where code of a given kind lives, how the registry and factory dispatch is wired, which trainer trains which model, and which of the many in-tree docs answers your question.
research/)Our research/ directory contains over 120 articles providing the theoretical foundations, implementation details, and best practices behind key components. Highlights include:
research/papers/)Five LaTeX manuscripts written against this codebase live under research/papers/, several with built PDFs: band_rms (band-constrained RMS normalization), bfunet (bias-free denoisers as image priors), cliffordnet_extensions, correlations, and logical_net.
There is no committed documentation directory and no doc generator — generate_docs.py and the make docs target were deleted as deprecated. For detailed documentation on every module, class, and function, browse the source tree directly: each subpackage ships a focused README.md (e.g. src/dl_techniques/analyzer/README.md) and a per-package AGENTS.md describing its conventions, patterns, and components. Every one of the 84 leaf model packages carries its own README.md. The research/ guides above complement these with the underlying theory.
The repository is organized for clarity, maintainability, and ease of contribution.
src/ is the import root: the library is imported as dl_techniques.*, the training pipelines as train.*, and the applications as applications.*.
dl_techniques/
├── src/dl_techniques/ # THE LIBRARY — 13 subpackages
│ ├── models/ # 80 leaf model packages in 11 FAMILY directories
│ │ │ # full catalogue: src/dl_techniques/models/README.md
│ │ ├── vision/ # 35 — resnet, convnext, vit, dino, swin_transformer, beit,
│ │ │ # yolo12, detr, vae, vq_vae, depth_anything, and the
│ │ │ # nested image_restoration/, super_resolution/, keypoints/
│ │ ├── language/ # 17 — bert, gemma, qwen, mamba, modern_bert, gpt2, colbert,
│ │ │ # byte_latent_transformer, hierarchical_reasoning_model, ...
│ │ ├── vision_language/ # 9 — clip, mobile_clip, fastvlm, nano_vlm, sd3_mmdit,
│ │ │ # ideogram4, and the nested sam/{sam1,sam2,sam3}
│ │ ├── time_series/ # 7 — tirex, nbeats, xlstm, deepar, prism, mdn, adaptive_ema
│ │ ├── general_purpose/ # 3 — kan, mothnet, power_mlp
│ │ ├── graph/ # 3 — relgt, graph_energy_transformer, shgcn
│ │ ├── neural_computer/ # 2 — ntm, nam
│ │ ├── common/ # 1 — power_sampling (model-agnostic inference machinery)
│ │ ├── memory/ # 1 — som
│ │ ├── point_cloud/ # 1 — latent_gmm_registration
│ │ └── tabular/ # 1 — tabm
│ ├── layers/ # 275 modules — 200 in 21 themed subpackages,
│ │ │ # 75 loose at the top level
│ │ ├── attention/ # 33 registered attention mechanisms (factory)
│ │ ├── ffn/ # 21 registered feed-forward networks (factory)
│ │ ├── activations/ # 22 registered activations (factory)
│ │ ├── norms/ # 18 registered normalization layers (factory)
│ │ ├── embedding/ # 13 registered positional/semantic embeddings (factory)
│ │ ├── transformers/ # Assembled transformer/encoder/decoder blocks
│ │ ├── heads/ # Task heads dispatched by domain (nlp/vision/vlm)
│ │ ├── time_series/ # Forecasting-specific layers
│ │ ├── fastvit/ # FastViT primitives (+ a committed reference impl)
│ │ ├── graphs/ # Graph neural network components
│ │ ├── moe/ # Mixture of Experts (MoE) system
│ │ ├── memory/ # External / associative memory layers
│ │ ├── statistics/ # Statistical and probabilistic layers (MDN, Flows)
│ │ └── ... # fusion, geometric, logic, mixtures, physics,
│ │ # reasoning, sequence_pooling, tokenizers
│ ├── losses/ # 44 specialized loss modules (AnyLoss, Goodhart, etc.)
│ ├── metrics/ # Custom Keras metrics (PSNR, SSIM, perplexity, Brier)
│ ├── optimization/ # Optimizers (Muon, VSGD, SGLD, Gefen, WW-PGD), LR schedules
│ ├── analyzer/ # Comprehensive model analysis toolkit and visualizers
│ ├── visualization/ # Plotting helpers for training and evaluation
│ ├── datasets/ # Dataset loaders and synthetic generators
│ ├── callbacks/ # Reusable Keras callbacks
│ ├── regularizers/ # Advanced regularization techniques (SRIP, Orthogonal)
│ ├── initializers/ # Structured initializers (Gabor, Haar, orthonormal, KAN)
│ ├── constraints/ # Weight constraints
│ └── utils/ # Core utilities, loggers, masking, and data handlers
├── src/train/ # 70 train_*.py entry points across 47 trainer directories
├── src/applications/ # Deployable apps (today: bias_free_denoiser, Streamlit)
├── research/ # 120+ in-depth articles, guides, and LaTeX papers
├── tests/ # 895 test modules mirroring src/dl_techniques/
└── REPO_MAP.md # Path-verified router — read this first
We welcome contributions from the research community. Whether you are implementing a new technique, improving documentation, or fixing bugs, your input is valuable.
pip install -e ".[dev]".git checkout -b feature/new-technique.dl_techniques.utils.logger (no print)..venv environment and write comprehensive tests using pytest, scoped to the modules you change (make test runs the full ~1.5h suite). Set MPLBACKEND=Agg when running training scripts on headless machines.research/ directory. Docstring style is not uniform across this repo — layers/ is predominantly Sphinx/reST (:param:), models/ is measurably mixed with the Sphinx exemplar models/language/bert/model.py as the model for new packages, and losses//metrics//utils//optimization//analyzer//visualization/ are Google-Args:-majority. Match the package you are editing and never convert a file wholesale; the measured counts and the greps that re-derive them are in src/dl_techniques/AGENTS.md § Core Conventions → Code Style.ModelAnalyzer toolkit.This project is licensed under the GNU General Public License v3.0.
Important Considerations:
See the LICENSE file for complete details.
This library is proudly sponsored and pioneered by Electi Consulting, a premier AI consultancy specializing in enterprise artificial intelligence, blockchain technology, and cryptographic solutions. The practical validation and enterprise-ready nature of these components has been made possible through Electi's extensive experience deploying state-of-the-art AI solutions across diverse industries:
Special recognition is extended to the open-source community and the many researchers whose groundbreaking work forms the foundation of this library.
This library stands on the shoulders of giants. Our implementations are grounded in a rigorous study of the source papers that have defined the field of modern deep learning.
Complete bibliographic information is available in the documentation for individual modules.
Python
99.1%