31 repos
rapidsai/cuml
NVIDIA cuML: GPU-Accelerated Machine Learning
5,281
15,787 commits
NVIDIA/cuml
NVIDIA/cutlass
CUDA Templates and Python DSLs for High-Performance Linear Algebra
10,452
896 commits
flashinfer-ai/flashinfer
FlashInfer: Kernel Library for LLM Serving
6,436
2,957 commits
RightNow-AI/autokernel
Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton…
1,557
43 commits
infervisor/plow
Packet Language for On-device Warps — a GPU inference compiler and runtime
19
431 commits
getainode/ainode
Turn any NVIDIA GPU into a local AI platform. Inference + fine-tuning in your browser. One command…
15
299 commits
m96-chan/PyGPUkit
Minimal GPU runtime for Python - high-performance CUDA kernels, memory management, and LLM…
3
297 commits
NVIDIA/TransformerEngine
A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit…
3,537
2,117 commits
NVIDIA/cuopt
GPU accelerated decision optimization
1,044
1,118 commits
nirw4nna/dsc
Tensor library & inference framework for machine learning
118
164 commits
NVIDIA/cudnn-frontend
cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing…
940
492 commits
vectorch-ai/ScaleLLM
A high-performance inference system for large language models, designed for production environments.
499
769 commits
logxio/picchio
Catch local LLMs spilling into the CPU. One run shows GPU placement, prefill, decode, memory, power…
18
105 commits
pegainfer-project/pegainfer
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
703
786 commits
gittensor-ai-lab/sparkinfer
Fastest MoE/LLM inference runtime for consumer and edge Blackwell GPUs. SN74 on Gittensor.
80
1,463 commits
lupinemachines/lupine
LUPINE is a GPU over IP bridge allowing GPUs on remote machines to be attached to CPU-only machines.
2,430
568 commits
NVIDIA-BioNeMo/BioNeMo-Inference-Runtime
Easy, fast, and memory-efficient structure prediction inference
53
1 commits
BabitMF/bmf
Cross-platform, customizable multimedia/video processing framework. With strong GPU acceleration,…
1,033
700 commits
tenstorrent/tt-metal
:metal: TT-NN operator library, and TT-Metalium low level kernel programming model.
1,684
29,414 commits
parca-dev/parca-agent
eBPF based always-on CPU/GPU profiler auto-discovering targets in Kubernetes and systemd, zero code…
746
3,882 commits
tiagomonteiro0715/fast_trimul
A drop-in, hardware-agnostic library for Fused Triangle Multiplicative Updates across…
5
7 commits
YaronKoresh/definers
A comprehensive Python toolkit for AI, data processing, media manipulation, and system utilities.
1
1,371 commits
LambdaLabsML/distributed-training-guide
Best practices & guides on how to write distributed pytorch training code
631
253 commits
oneapi-src/oneAPI-samples
Samples for Intel® oneAPI Toolkits
1,161
2,434 commits
beam-cloud/beta9
Ultrafast serverless GPU inference, sandboxes, and background jobs
1,781
1,891 commits
catboost/catboost
A fast, scalable, high performance Gradient Boosting on Decision Trees library, used for ranking,…
9,105
34,440 commits
sixteen-miles-labs/sparklab
Run frontier open-weight models privately on NVIDIA DGX Spark.
14
168 commits
notwitcheer/llm-bench-rig
Dual-engine (llama.cpp + vLLM) LLM benchmarking pipeline for GGUF & safetensors on NVIDIA GPUs —…
38
307 commits
skyloevil/llm-scratch-pytorch
lm-scratch-pytorch - The code is designed to be beginner-friendly, with a focus on understanding…
101
146 commits
sowson/darknet
Darknet on OpenCL Convolutional Neural Networks on OpenCL on Intel & NVidia & AMD & Mali GPUs for…
197
453 commits