tensorchord/Awesome-LLMOps

An awesome & curated list of best LLMOps tools for developers

Shell

5,952

319 commits

updated Oct 5, 2026

See the code

README

Awesome LLMOps

discord invitation link

An awesome & curated list of the best LLMOps tools for developers.

[!NOTE] Contributions are most welcome, please adhere to the contribution guidelines.

Table of Contents

Model

Large Language Model

ProjectDetailsRepository
AlpacaCode and documentation to train Stanford's Alpaca models, and generate the data.GitHub Badge
BELLEA 7B Large Language Model fine-tune by 34B Chinese Character Corpus, based on LLaMA and Alpaca.GitHub Badge
BloomBigScience Large Open-science Open-access Multilingual Language ModelGitHub Badge
dollyDatabricks’ Dolly, a large language model trained on the Databricks Machine Learning PlatformGitHub Badge
Falcon 40BFalcon-40B-Instruct is a 40B parameters causal decoder-only model built by TII based on Falcon-40B and finetuned on a mixture of Baize. It is made available under the Apache 2.0 license.
FastChat (Vicuna)An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and FastChat-T5.GitHub Badge
GemmaGemma is a family of lightweight, open models built from the research and technology that Google used to create the Gemini models.
GLM-6B (ChatGLM)An Open Bilingual Pre-Trained Model, quantization of ChatGLM-130B, can run on consumer-level GPUs.GitHub Badge
ChatGLM2-6BChatGLM2-6B is the second-generation version of the open-source bilingual (Chinese-English) chat model ChatGLM-6B.GitHub Badge
GLM-130B (ChatGLM)An Open Bilingual Pre-Trained Model (ICLR 2023)GitHub Badge
GPT-NeoXAn implementation of model parallel autoregressive transformers on GPUs, based on the DeepSpeed library.GitHub Badge
JebadiahOpen decision models that return a probability for each option instead of generating text; available in 4B, 9B and 27B sizes.GitHub Badge
LuotuoA Chinese LLM, Based on LLaMA and fine tune by Stanford Alpaca, Alpaca LoRA, Japanese-Alpaca-LoRA.GitHub Badge
Mixtral-8x7B-v0.1The Mixtral-8x7B Large Language Model (LLM) is a pretrained generative Sparse Mixture of Experts.
StableLMStableLM: Stability AI Language ModelsGitHub Badge

⬆ back to ToC

CV Foundation Model

ProjectDetailsRepository
disco-diffusionA frankensteinian amalgamation of notebooks, models and techniques for the generation of AI Art and Animations.GitHub Badge
midjourneyMidjourney is an independent research lab exploring new mediums of thought and expanding the imaginative powers of the human species.
segment-anything (SAM)produces high quality object masks from input prompts such as points or boxes, and it can be used to generate masks for all objects in an image.GitHub Badge
stable-diffusionA latent text-to-image diffusion modelGitHub Badge

⬆ back to ToC

Audio Foundation Model

ProjectDetailsRepository
barkBark is a transformer-based text-to-audio model created by Suno. Bark can generate highly realistic, multilingual speech as well as other audio - including music, background noise and simple sound effects.GitHub Badge
FunASRSpeech recognition toolkit with pretrained models and tools for voice activity detection, punctuation restoration, and speaker diarization.GitHub Badge
whisperRobust Speech Recognition via Large-Scale Weak SupervisionGitHub Badge

⬆ back to ToC

Robotics Foundation Model

[!NOTE] Emerging Architectures in VLA:

  • Continuous Diffusion Language Models: Integrate diffusion heads or flow-matching to VLMs (e.g., DiVLA, OpenPI), enabling smooth, precise continuous action generation rather than discretized tokens.
  • Recurrent Language Models: Utilize State Space Models (SSMs) like Mamba or recurrent transformers (e.g., RoboMamba, RD-VLA) to reduce inference memory and handle temporal dependencies, allowing iterative reasoning for complex robotic decision-making.
ProjectDetailsRepository
DiVLAA continuous diffusion-based Vision-Language-Action model that integrates diffusion policies into autoregressive VLMs for robust and precise continuous robotic control.GitHub Badge
LeRobotA central community library by Hugging Face for AI in robotics — end-to-end learning tools, data pipelines, and support for training/deploying VLA models.GitHub Badge
OctoA transformer-based generalist robot policy pretrained on 800K+ robot trajectories from the Open X-Embodiment dataset. Supports language instructions, goal images, and fine-tuning to new embodiments.GitHub Badge
OpenPIOpen-source VLA models from Physical Intelligence, including π₀ and π₀.5 — flow-based vision-language-action models pretrained on large-scale robot data with fine-tuning support.GitHub Badge
OpenVLAA 7B-parameter open-source Vision-Language-Action model trained on 970K+ robot demonstrations from the Open X-Embodiment dataset for generalist robotic manipulation.GitHub Badge
RoboMambaAn efficient VLA model leveraging State Space Models (Mamba) instead of standard self-attention, offering linear inference complexity for efficient, recurrent robotic reasoning.GitHub Badge
SmolVLAA compact ~450M parameter VLA by Hugging Face, designed to be computationally efficient and accessible, running on consumer GPUs or CPUs. Part of the LeRobot ecosystem.

Serving

Large Model Serving

ProjectDetailsRepository
Alpaca-LoRA-ServeAlpaca-LoRA as Chatbot serviceGitHub Badge
HiggsRust inference server for Apple Silicon: MLX models behind OpenAI and Anthropic APIs, routing to remote providers, desktop dashboard.GitHub Badge
OneCompFujitsu Research's post-training quantization pipeline for LLMs (QEP, AutoBit, JointQ, rotation) with vLLM plugin (arXiv:2603.28845).GitHub Badge
CTranslate2fast inference engine for Transformer models in C++GitHub Badge
Clip-as-a-serviceserving the OpenAI CLIP modelGitHub Badge
DeepSpeed-MIIMII makes low-latency and high-throughput inference possible, powered by DeepSpeed.GitHub Badge
Faster Whisperfast inference engine for whisper in C++ using CTranslate2.GitHub Badge
FlexGenRunning large language models on a single GPU for throughput-oriented scenarios. (Archived)GitHub Badge
FlowiseDrag & drop UI to build your customized LLM flow using LangchainJS.GitHub Badge
lilbeeSingle-binary local manager and search engine that stands up a llama-server fleet sized to your GPUs by gguf-parser for VRAM-aware multi-GPU placementGitHub Badge
llama.cppPort of Facebook's LLaMA model in C/C++GitHub Badge
LLMKubeKubernetes operator for LLM inference with pluggable runtimes (llama.cpp, PersonaPlex/Moshi, generic), multi-GPU sharding, NVIDIA CUDA and Apple Silicon Metal support, and GGUF/MLX/SafeTensors model formats.GitHub Badge
ShimmyPython-free Rust inference server with OpenAI API compatibility and hot model swappingGitHub Badge
InfinityRest API server for serving text-embeddingsGitHub Badge
Modelz-LLMOpenAI compatible API for LLMs and embeddings (LLaMA, Vicuna, ChatGLM and many others)GitHub Badge
NobodyWhoOn-device LLM inference engine (Rust / llama.cpp) to embed local models directly in games and apps, with Godot / Flutter / React Native / Swift bindings; streaming, embeddings, tool calling, GBNF-structured output, and STT/TTS.GitHub Badge
Off GridOpen-source iOS/Android app running LLMs on-device via llama.cpp. Voice (Whisper), vision, image gen, tool calling — fully offline.GitHub Badge
OllamaServe Llama 2 and other large language models locally from command line or through a browser interface.GitHub Badge
Rapid-MLXOpenAI-compatible LLM inference server for Apple Silicon using MLX. 2-4x faster than Ollama with tool calling and prompt caching.GitHub Badge
TensorRT-LLMInference engine for TensorRT on Nvidia GPUsGitHub Badge
text-generation-inferenceLarge Language Model Text Generation InferenceGitHub Badge
text-embeddings-inferenceInference for text-embedding modelsGitHub Badge
tokenizers💥 Fast State-of-the-Art Tokenizers optimized for Research and ProductionGitHub Badge
vllmA high-throughput and memory-efficient inference and serving engine for LLMs.GitHub stars
whisper-ctranslate2is a 4x faster and low-memory usage drop-in cli replacement that supports word-level timestamps and VAD filterGitHub Badge
whisper.cppPort of OpenAI's Whisper model in C/C++GitHub Badge
x-stable-diffusionReal-time inference for Stable Diffusion - 0.88s latency. Covers AITemplate, nvFuser, TensorRT, FlashAttention. (Archived)GitHub Badge

⬆ back to ToC

Frameworks/Servers for Serving

ProjectDetailsRepository
BentoMLThe Unified Model Serving FrameworkGitHub Badge
BifrostHigh-performance AI gateway written in Go. Unifies 23+ providers behind one OpenAI-compatible endpoint with automatic failover, adaptive load balancing, semantic caching, virtual keys and budgets, MCP tool calling, and native Prometheus observability.GitHub Badge
JinaBuild multimodal AI services via cloud native technologies · Model Serving · Generative AI · Neural Search · Cloud NativeGitHub Badge
MosecA machine learning model serving framework with dynamic batching and pipelined stages, provides an easy-to-use Python interface.GitHub Badge
mcpproxy-goOpen-source MCP proxy with BM25 tool filtering, quarantine security, activity logging, and web UI. Routes multiple MCP servers through single endpoint, reducing context bloat by ~97%.GitHub Badge
SentryNode GatewayOpen-core AI inference governance and FinOps platform providing semantic routing, <40ms agent loop intervention, strict budget caps, and cryptographic audit trails. Includes voice pipeline governance (STT→LLM→TTS).GitHub Badge
TFServingA flexible, high-performance serving system for machine learning models.GitHub Badge
TorchserveServe, optimize and scale PyTorch models in production (Archived)GitHub Badge
Triton Server (TRTIS)The Triton Inference Server provides an optimized cloud and edge inferencing solution.GitHub Badge
langchain-serveServerless LLM apps on Production with Jina AI Cloud (Archived)GitHub Badge
lanarkyFastAPI framework to build production-grade LLM applicationsGitHub Badge
ray-llmLLMs on Ray - RayLLM (Archived)GitHub Badge
SIEOpen-source inference server and production cluster for serving embeddings, reranking, OCR, structured output, content safety, and generative models through one API.GitHub Badge
XinferenceReplace OpenAI GPT with another LLM in your app by changing a single line of code. Xinference gives you the freedom to use any LLM you need. With Xinference, you're empowered to run inference with any open-source language models, speech recognition models, and multimodal models, whether in the cloud, on-premises, or even on your laptop.GitHub Badge
KubeAIDeploy and scale machine learning models on Kubernetes. Built for LLMs, embeddings, and speech-to-text.GitHub Badge
KaitoA Kubernetes operator that simplifies serving and tuning large AI models (e.g. Falcon or phi-3) using container images and GPU auto-provisioning. Includes an OpenAI-compatible server for inference and preset configurations for popular runtimes such as vLLM and transformers.GitHub Badge
Open ResponsesServerless open-source platform for building long-running LLM agents with tool use.GitHub Badge
KubeStellar ConsoleAI-powered multi-cluster Kubernetes dashboard for hybrid edge and cloud. GPU monitoring, LLM inference cluster management, benchmark streaming, and 20+ CNCF integrations. CNCF Sandbox (Apache 2.0).GitHub Badge
laya-decision-apiFastAPI HTTP service for the local Laya System-1 decision model: offline structured decisions (choice / score / yes-no) with probabilities and confidence; MLX and torch backendsGitHub stars

⬆ back to ToC

Security

Frameworks for LLM security

ProjectDetailsRepository
API Relay AuditLocal security audit for AI API relays and LLM proxies; checks prompt injection, model identity drift, tool-call rewriting, error leakage, SSE anomalies, and Web3 wallet probes.GitHub Badge
CordumSafety-first agent orchestration platform with pre-dispatch policy evaluation, output scanning (PII, secrets, injection), job scheduling, workflow engine, and full audit trail.GitHub Badge
brood-boxCLI tool for running coding agents inside hardware-isolated microVMs with snapshot isolation, egress control, and MCP authorization.GitHub Badge
DOS KernelDeterministic trust kernel for AI-agent fleets: verifies agent "done" claims from git evidence rather than self-report, arbitrates concurrent file access between agents, and audits commit messages against their own diffs.GitHub Badge
dstackOpen-source confidential AI framework for secure LLM deployment with data privacy, providing hardware-enforced isolation using Intel TDX and NVIDIA Confidential Computing.GitHub Badge
EmbedGuardCross-layer detection and provenance attestation for adversarial embedding attacks in RAG systems. Flags poisoned or manipulated embeddings before they reach the retrieval layer.GitHub Badge
PlexiglassA Python Machine Learning Pentesting Toolbox for Adversarial Attacks. Works with LLMs, DNNs, and other machine learning algorithms.GitHub Badge
PrivacyScrubberZero-Trust client-side PII and developer secrets sanitizer and MCP server for Cursor, Windsurf, and Claude Desktop with sub-2ms in-memory tokenization.GitHub Badge
PrivAiTeSelf-hosted OpenAI-compatible proxy that reversibly pseudonymizes PII before requests reach the provider, tool-call arguments and multimodal included.GitHub Badge
shimOpen-core AI gateway that redacts PII before prompts reach model providers (Presidio plus Turkish TCKN/VKN recognizers). The source-available enterprise edition adds a tamper-evident hash-chain audit log with daily Merkle anchors as evidence for EU AI Act and KVKK record-keeping.GitHub Badge
wardcatOn-premises PII and sensitive-data detection and masking for LLM and RAG workflows, combining regex, NER, and optional local LLMs with reversible masking.GitHub Badge

⬆ back to ToC

Observability

ProjectDetailsRepository
AgentCostTrack and optimize LLM API costs across OpenAI, Anthropic, and LangChain.GitHub Badge
Azure OpenAI Logger"Batteries included" logging solution for your Azure OpenAI instance.GitHub Badge
budget-guardZero-infra hard spending caps for LLM APIs — blocks before overspending (atomic reservation), per-feature cost attribution, streaming usage, retry-storm detection, and adapters for Vercel AI SDK / LangChain.js / LlamaIndex.TS / Mastra.GitHub Badge
ClevAgentRuntime monitoring for AI agents — heartbeat watchdog, loop detection, cost tracking, auto-restart. Python SDK or HTTP API.
CodeBurnLocal-first cost and token tracking for AI coding agents (Claude Code, Codex, Cursor and 38 more) by project, model and task, with waste detection and plan quota tracking.GitHub Badge
DeepchecksTests for Continuous Validation of ML Models & Data. Deepchecks is a Python package for comprehensively validating your machine learning models and data with minimal effort.GitHub Badge
diglineRegression testing for LLM apps: approved scores, prompt and commit pinned as a versioned baseline in your repo; compare says which case got worse and by how much.GitHub Badge
EvidentlyAn open-source framework to evaluate, test and monitor ML and LLM-powered systems.GitHub Badge
EvalViewRegression testing for AI agents. Snapshot behavior, detect tool-call and output regressions, with golden-baseline diffing and LLM-as-judge scoring. Supports LangGraph, CrewAI, OpenAI, Claude, and any HTTP API.GitHub Badge
Fiddler AIEvaluate, monitor, analyze, and improve machine learning and generative models from pre-production to production. Ship more ML and LLMs into production, and monitor ML and LLM metrics like hallucination, PII, and toxicity.GitHub Badge
FlowlinesBehavioral observability for production AI agents that detects drift, context loss, constraint violations, and user frustration across sessions.
GiskardTesting framework dedicated to ML models, from tabular to LLMs. Detect risks of biases, performance issues and errors in 4 lines of code.GitHub Badge
QWEDDeterministic verification protocol for LLM outputs using 8 formal verification engines (SymPy, Z3, AST, SQLGlot). Prevents hallucinations through mathematical proofs rather than statistical methods.GitHub Badge
Great ExpectationsAlways know what to expect from your data.GitHub Badge
HeliconeOpen source LLM observability platform. One line of code to monitor, evaluate, and experiment with features like prompt management, agent tracing, and evaluations.GitHub Badge
hermes-rubricEvidence-first assessment for agent outputs and applications. Synthesizes, loads, or accepts a rubric; collects citations for each dimension; scores against accepted evidence; reports coverage; and emits reproducibility receipts. Apache-2.0.GitHub Badge
InferrailPer-run dollar budgets for AI agents and jobs: an OpenAI- and Anthropic-compatible gateway (in-process or self-hosted) that reserves each call's worst-case cost before it is sent, refuses calls that don't fit before the provider, and keeps a payload-free cost record per run.GitHub Badge
Traceloop OpenLLMetryOpenTelemetry-based observability and monitoring for LLM and agents workflows.GitHub Badge
Langfuse 🪢Open-source LLM observability platform that helps teams collaboratively debug, analyze, and iterate on their LLM applications.GitHub Badge
LookspanLocal-first observability for AI agents. One command (npx lookspan) runs a dashboard with traces, a timeline/waterfall view, cost tracking, replay & LLM-as-judge, and datasets. MCP-native, with OpenAI/Anthropic drop-ins and an OpenTelemetry receiver — data stays on your machine.GitHub Badge
whylogsThe open standard for data loggingGitHub Badge
Maxim AIPlatform for AI Agent Simulation, Evaluation & Observability
NoveumAI agent tracing, evaluation datasets from production interactions, and model comparisons on quality, latency, and cost.
onWatchLightweight Go CLI that tracks AI API quota usage across 7 providers (Anthropic, OpenAI, GitHub Copilot, MiniMax, and more). Background daemon, <50MB RAM, zero telemetry, SQLite storage.GitHub Badge
OpenGATEDeterministic verification for evidence-grounded AI — gold-anchored grounding checks (required facts, ungrounded numbers, abstention) with no LLM judge, plus CI regression gates, pytest/DeepEval integrations, and an MCP server.GitHub Badge
phoenix2pytestReads Arize Phoenix traces flagged as failures and synthesises pytest regression tests, so production-discovered LLM failures feed back into the test suite without manual translation.GitHub Badge
RagTuneCLI tool for debugging and benchmarking RAG retrieval. EXPLAIN ANALYZE for your retrieval layer.GitHub Badge
runtapeFinds which part of an agent's context caused a bad decision by rerunning it with pieces removed, checks candidate fixes the same way, and writes a pytest regression test. Runs locally with OpenAI, Anthropic, LangGraph and Ollama.GitHub Badge
traceAIOpen-source AI tracing framework built on OpenTelemetry for deep observability across agentic and LLM workflows.GitHub Badge
Future AGIProduction-grade SDK for observability, automated evaluations and prompt management with sub-100ms guardrails for LLM/agent workflows.GitHub Badge
semantic-coverageVisualizes RAG knowledge gaps and "blind spots" using 2D UMAP clustering and density detection.GitHub Badge
Token PoliceSDK (Python, Node) that prices each LLM call before it is made and blocks or reroutes it on per-user, per-session or per-workflow budget rules; prompt text never leaves the app.GitHub Badge
Weco ObserveObservability and debugging tool for AI research agents. Trace multi-step LLM agent runs, visualize decision trees, and identify failure modes in autonomous research workflows. Cloud hosted with open-source agent integration.
witnessA recording cache for model APIs: one Rust binary that proxies any LLM endpoint, dedupes repeat calls, and writes every model call, MCP tool call, and command side effect into a tamper-evident hash-chained journal with Ed25519 agent identity and Merkle proofs anchored to Sigstore Rekor.GitHub Badge
wobblyMetamorphic testing for LLM extraction — catch wrong outputs with no ground-truth labels: perturb inputs in ways that shouldn't change the answer, flag when they do.

⬆ back to ToC

LLMOps

ProjectDetailsRepository
agentaThe LLMOps platform to build robust LLM apps. Easily experiment and evaluate different prompts, models, and workflows to build robust apps.GitHub Badge
AgentMarkType-Safe Markdown-based AgentsGitHub Badge
AgentFieldOpen-source control plane for building and operating AI agents like APIs at scale, with routing, memory, observability, identity, auth, and policy controls.GitHub Badge
AgnosGateway-agnostic, self-hosted LLM control plane (MIT): an OpenAI-compatible governance proxy (auth, budgets, guardrails, cost, observability) that runs LiteLLM, Bifrost or Portkey as a swappable, stateless translation engine, so provider keys stay in your own control plane instead of the gateway.GitHub Badge
AI studioA Reliable Open Source AI studio to build core infrastructure stack for your LLM Applications. It allows you to gain visibility, make your application reliable, and prepare it for production with features such as caching, rate limiting, exponential retry, model fallback, and more.GitHub Badge
AISIXOpen-source AI gateway written in Rust: one OpenAI-compatible API (plus a native Anthropic Messages API) in front of LLM providers, and a gateway for MCP servers and A2A agents, with shared API keys, rate limits, guardrails, caching, and Prometheus/OTLP observability.GitHub Badge
AIWGDeploys project-owned agents, skills, rules, and governed workflows across AI coding platforms.GitHub Badge
Arize-PhoenixML observability for LLMs, vision, language, and tabular models.GitHub Badge
BitRouterAgent-native LLM router that optimizes your agent with every run — zero harness changes, with every model call reliable, traceable, secure, and cost-effective. Routes across OpenAI, Anthropic, Google, OpenRouter, Bedrock, GitHub Copilot, and more through one local endpoint, with cross-protocol translation, an MCP gateway, guardrails, observability, and multi-account failover. Written in Rust.GitHub Badge
boldrouterOpenAI-compatible AI gateway developed and hosted in Switzerland, with automatic routing, failover, and unified prepaid billing.
BudgetMLDeploy a ML inference service on a budget in less than 10 lines of code.GitHub Badge
CauraGoverned shared memory for AI agent fleets. Multi-agent, multi-tenant and MCP-native, with trust tiers, audit trails, a knowledge graph and self-improving retrieval.GitHub Badge
Cheshire Cat AIWeb framework to create vertical AI agents. FastAPI based, plugin system inspired to WordPress, admin panel, vector DB includedGitHub Badge
ChimeraforgeLLM deployment planner and benchmarking CLI: given a model and a GPU it answers will it fit, will it hit your SLO, and what will it cost. Searches model x quantization x backend x GPU count against VRAM, quality, latency, cost, energy and an opt-in safety gate; models tensor/pipeline parallelism, MoE active parameters, FP8, KV-quant and self-host-vs-API break-even. Every number is labeled measured/estimated/unknown. Also ships an MCP server.GitHub Badge
claude-routerLocal prompt router that picks the right Claude model tier and prepends the right scaffold using local embeddings, before the API call.GitHub Badge
CogniCoreStructured experience memory and learning infrastructure for AI agents, providing verified experience reuse, failure learning, environment-aware validity, and cross-session/cross-agent memory.GitHub Badge
ContextoSelf-hosted context engine for AI agents with persistent conversation memory and recall. Works as a drop-in OpenAI-compatible proxy, OpenClaw plugin, or memory SDK — no code changes required.GitHub Badge
CortexGenerates typed MCP servers, interactive API documentation, and SDKs from OpenAPI, AsyncAPI, GraphQL, gRPC, and OpenRPC sources.GitHub Badge
CrewDockSelf-hosted control plane for managing AI agents, with task approvals, activity logs, and cost tracking.GitHub Badge
DataoortsEnjoy unlimited API calls with Serverless AI Workers/LLMs for just $25 per month. No rate or concurrency limits.
deeplakeStream large multimodal datasets to achieve near 100% GPU utilization. Query, visualize, & version control data. Access data w/o the need to recompute the embeddings for the model finetuning.GitHub Badge
depdeskFinds deprecated and retired LLM model identifiers in a codebase and fails CI before a provider retirement reaches production. Parses Python with the stdlib ast module, weights findings by real call volume from a usage export, and ships a hand-transcribed Anthropic and OpenAI catalog carrying the date each provider page was verified. Zero dependencies.GitHub Badge
DifyOpen-source framework aims to enable developers (and even non-developers) to quickly build useful applications based on large language models, ensuring they are visual, operable, and improvable.GitHub Badge
DistilReversible, certified context compression for LLM agents: gates each compression through a statistical non-inferiority test so the agent's decisions don't change. Stdlib-only, zero-dependency proxy for Anthropic/OpenAI/Gemini.GitHub Badge
Doubleword Control LayerThe world's fastest open-source AI model gateway — ~450× less overhead than LiteLLM. Turns any model into a production-ready, OpenAI-compatible API with built-in auth, rate limits, and user controls.GitHub Badge
DstackCost-effective LLM development in any cloud (AWS, GCP, Azure, Lambda, etc).GitHub Badge
EjentumCognitive-harness MCP server with four tools (reasoning, code, anti-deception, memory) returning a structured scaffold (failure pattern, procedure, suppression vectors, falsification test) the agent absorbs before generating. Hosted at api.ejentum.com/mcp and on npm as ejentum-mcp.GitHub Badge
EmbedchainFramework to create ChatGPT like bots over your dataset.GitHub Badge
EpsillaAn all-in-one platform to create vertical AI agents powered by your private data and knowledge.
EvidentlyAn open-source framework to evaluate, test and monitor ML and LLM-powered systems.GitHub Badge
FerryAPILow-cost OpenAI-compatible AI API gateway for production workloads, with prepaid balance, usage billing, customer API keys, and provider account pools.
Fiddler AIEvaluate, monitor, analyze, and improve MLOps and LLMOps from pre-production to production.
FreeLLMAPIOpenAI-compatible proxy that stacks the free tiers of 28 LLM providers (~4B free tokens/month) behind one /v1 endpoint — smart routing, automatic failover, encrypted keys.GitHub Badge
GlideCloud-Native LLM Routing Engine. Improve LLM app resilience and speed.GitHub Badge
GoModelAI gateway exposing a unified OpenAI-compatible API across OpenAI, Anthropic, Gemini, Groq, xAI, Ollama and other providers, with routing, usage tracking, rate limits, and guardrails.GitHub Badge
gotoHumanBring a human into the loop in your LLM-based and agentic workflows. Prompt users to approve actions, select next steps, or review and validate generated results.
GPTCacheCreating semantic cache to store responses from LLM queries.GitHub Badge
GPUStackAn open-source GPU cluster manager for running and managing LLMsGitHub Badge
HaystackQuickly compose applications with LLM Agents, semantic search, question-answering and more.GitHub Badge
HiveOpen-source AI agent framework for building goal-driven, self-improving autonomous agents with auto-generated graphs, evolution loops, and MCP integration.GitHub Badge
HeliconeOpen-source LLM observability platform for logging, monitoring, and debugging AI applications. Simple 1-line integration to get started.GitHub Badge
HumanloopThe LLM evals platform for enterprises, providing tools to develop, evaluate, and observe AI systems.
HypersigilOpen-source prompt lifecycle management and gateway with a Web UI.GitHub Badge
IzloPrompt management tools for teams. Store, improve, test, and deploy your prompts in one unified workspace.
Keywords AIA unified DevOps platform for AI software. Keywords AI makes it easy for developers to build LLM applications.
KubeIntellectLLM-orchestrated multi-agent framework for Kubernetes operations. Investigates a live cluster via kubectl, Prometheus (PromQL) and Loki (LogQL), correlates the evidence into a root cause, and gates every mutating action behind human approval with role-based access control.GitHub Badge
KunavoHosted AI API gateway with OpenAI-compatible and native Anthropic endpoints, one prepaid balance, and per-request cost accounting.
MLflowAn open-source framework for the end-to-end machine learning lifecycle, helping developers track experiments, evaluate models/prompts, deploy models, and add observability with tracing.GitHub Badge
LaminarOpen-source all-in-one platform for engineering AI products. Traces, Evals, Datasets, Labels.GitHub Badge
langchainBuilding applications with LLMs through composabilityGitHub Badge
LangFlowAn effortless way to experiment and prototype LangChain flows with drag-and-drop components and a chat interface.GitHub Badge
LangfuseOpen Source LLM Engineering Platform: Traces, evals, prompt management and metrics to debug and improve your LLM application.GitHub Badge
LangKitOut-of-the-box LLM telemetry collection library that extracts features and profiles prompts, responses and metadata about how your LLM is performing over time to find problems at scale.GitHub Badge
LangWatchLLM Ops platform with Analytics, Monitoring, Evaluations and an LLM Optimization Studio powered by DSPyGitHub Badge
lintlangStatic analysis for AI agent configs, tool descriptions, and system prompts. Catches vague tool descriptions, missing stop conditions, and schema gaps before they reach runtime. Zero-LLM, deterministic checks, built for CI.GitHub Badge
LiteLLM 🚅A simple & light 100 line package to standardize LLM API calls across OpenAI, Azure, Cohere, Anthropic, Replicate API EndpointsGitHub Badge
Literal AIMulti-modal LLM observability and evaluation platform. Create prompt templates, deploy prompts versions, debug LLM runs, create datasets, run evaluations, monitor LLM metrics and collect human feedback.
LlamaIndexProvides a central interface to connect your LLMs with external data.GitHub Badge
LLMAppLLM App is a Python library that helps you build real-time LLM-enabled data pipelines with few lines of code.GitHub Badge
LLMFlowsLLMFlows is a framework for building simple, explicit, and transparent LLM applications such as chatbots, question-answering systems, and agents.GitHub Badge
LLMGraphNo-code visual builder for LLM/AI workflows: turn your docs and models into RAG chatbots and AI agents, then deploy each to a REST API and an embeddable chat widget in one click.
LRMCLI/TUI tool for managing localization files (.resx, JSON, Android, iOS) with LLM-powered translation via Ollama, validation, and code scanning for unused/missing keys.GitHub Badge
LunaryObservability and prompt management for LLM chabots and agents. Debug agents with powerful tracing and logging. Usage analytics and dive deep into the history of your requests. Developer friendly modules with plug-and-play integration into LangChain.GitHub Badge
MengramOpen-source memory infrastructure for AI agents. Provides semantic (entities/facts), episodic (conversations), and procedural (learned behaviors) memory with auto-reflection. Python SDK, JS SDK, MCP server, and REST API.GitHub Badge
magenticSeamlessly integrate LLMs as Python functions. Use type annotations to specify structured output. Mix LLM queries and function calling with regular Python code to create complex LLM-powered functionality.GitHub Badge
Manag.aiYour all-in-one prompt management and observability platform. Craft, track, and perfect your LLM prompts with ease.
MirascopeIntuitive convenience tooling for lightning-fast, efficient development and ensuring quality in LLM-based applicationsGitHub Badge
NativePortInfrastructure for accessing web-search, scraping, browser-automation, voice, and model-inference providers through a single account, API key, and balance. Publishes dated, capability-specific leaderboards and a public methodology.
NeurolinkMulti-provider AI agent framework that unifies 12+ LLM providers (OpenAI, Google, Anthropic, AWS, Azure, Groq, etc.) with workflow orchestration. Production-grade platform for building LLM applications with streaming, tool calling, caching, and enterprise features. Battle-tested at 15M+ requests/month.GitHub Badge
NoraSelf-hosted control plane for deploying and operating OpenClaw and Hermes agent fleets on Docker and Kubernetes, with lifecycle management, observability, cost tracking, secrets, schedules, and REST, CLI, and MCP interfaces.GitHub Badge
OpenLITOpenLIT is an OpenTelemetry-native GenAI and LLM Application Observability tool and provides OpenTelmetry Auto-instrumentation for monitoring LLMs, VectorDBs and Frameworks. It provides valuable insights into token & cost usage, user interaction, and performance related metrics.GitHub Badge
OpikConfidently evaluate, test, and ship LLM applications with a suite of observability tools to calibrate language model outputs across your dev and production lifecycle.GitHub Badge
OrcaRouter LiteSelf-hosted, OpenAI-compatible LLM router. Bring your own provider keys and route across 100+ models from multiple providers with automatic failover and streaming; model="auto" selects a model by cost, latency, or quality. Includes a local analytics dashboard and does not send telemetry.GitHub Badge
Parea AIPlatform and SDK for AI Engineers providing tools for LLM evaluation, observability, and a version-controlled enhanced prompt playground.GitHub Badge
Pezzo 🕹️Pezzo is the open-source LLMOps platform built for developers and teams. In just two lines of code, you can seamlessly troubleshoot your AI operations, collaborate and manage your prompts in one place, and instantly deploy changes to any environment.GitHub Badge
Pilot ProtocolOpen-source overlay network giving AI agents a permanent virtual address, encrypted UDP tunnels with NAT traversal, and an explicit per-peer trust model, plus an app store of installable agent-native capabilities (discover → install → call).GitHub Badge
PraisonAIProduction-ready Multi-AI Agents framework with self-reflection. Fastest agent instantiation (3.77μs), 100+ LLM support via LiteLLM, MCP integration, agentic workflows (route/parallel/loop/repeat), built-in memory, Python & JS SDKs.GitHub Badge
PromptDXA declarative, extensible, and composable approach for developing LLM prompts using Markdown and JSX.GitHub Badge
PromptHubFull stack prompt management tool designed to be usable by technical and non-technical team members. Test, version, collaborate, deploy, and monitor, all from one place.
promptfooOpen-source tool for testing & evaluating prompt quality. Create test cases, automatically check output quality and catch regressions, and reduce evaluation cost.GitHub Badge
PromptFoundryThe simple prompt engineering and evaluation tool designed for developers building AI applications.GitHub Badge
PromptLayer 🍰Prompt Engineering platform. Collaborate, test, evaluate, and monitor your LLM applicationsGithub Badge
PromptMageOpen-source tool to simplify the process of creating and managing LLM workflows and prompts as a self-hosted solution.GitHub Badge
PromptSiteA lightweight Python library for prompt lifecycle management that helps you version control, track, experiment and debug with your LLM prompts with ease. Minimal setup, no servers, databases, or API keys required - works directly with your local filesystem, ideal for data scientists and engineers to easily integrate into existing LLM workflows
PrompteamsPrompt management system. Version, test, collaborate, and retrieve prompts through real-time APIs. Have GitHub style with repos, branches, and commits (and commit history).
prompttoolsOpen-source tools for testing and experimenting with prompts. The core idea is to enable developers to evaluate prompts using familiar interfaces like code and notebooks. In just a few lines of codes, you can test your prompts and parameters across different models (whether you are using OpenAI, Anthropic, or LLaMA models). You can even evaluate the retrieval accuracy of vector databases.GitHub Badge
Puzzlet AIThe Git-Based LLM Engineering Platform. Achieve more from GenAI: Manage, evaluate, and improve your full-stack LLM application - with version control, type-safety, and local development built-in.
QuotaflowAI token and API resource utilization platform that helps teams reduce wasted subscribed quota and improve turnover across controlled internal pools.
systemprompt.ioSystemprompt.io is a Rest API with quality tooling to enable the creation, use and observability of prompts in any AI system. Control every detail of your prompt for a SOTA prompt management experience.
TeamoRouterLLM routing gateway for OpenClaw. One API key to access Claude, GPT-4o, Gemini, DeepSeek, Kimi, MiniMax. Smart routing modes (teamo-best, teamo-balanced, teamo-eco) auto-pick the optimal model. Up to 50% off official prices. 2-second install via skill.md.
TokenMixAI gateway routing 171 LLMs from 14 providers (Claude, GPT, Gemini, DeepSeek, Qwen, and more) through one OpenAI-compatible endpoint. Pass-through pricing, automatic failover, no monthly subscription.
TreeScaleAll In One Dev Platform For LLM Apps. Deploy LLM-enhanced APIs seamlessly using tools for prompt optimization, semantic querying, version management, statistical evaluation, and performance tracking. As a part of the developer friendly API implementation TreeScale offers Elastic LLM product, which makes a unified API Endpoint for all major LLM providers and open source models.
TrueFoundryDeploy LLMOps tools like Vector DBs, Embedding server etc on your own Kubernetes (EKS,AKS,GKE,On-prem) Infra including deploying, Fine-tuning, tracking Prompts and serving Open Source LLM Models with full Data Security and Optimal GPU Management. Train and Launch your LLM Application at Production scale with best Software Engineering practices.
ReliableGPT 💪Handle OpenAI Errors (overloaded OpenAI servers, rotated keys, or context window errors) for your production LLM Applications.GitHub Badge
Registry BrokerUniversal index and routing layer for AI agents. Aggregates agent metadata from multiple registries (NANDA, MCP, Virtuals, OpenRouter, A2A, X402 Bazaar) across web2 and web3, normalizes profiles, and provides protocol translation between agent ecosystems.GitHub Badge
RhesisOpen-source testing infrastructure for LLM and agentic applications. Collaborative platform enabling teams to define quality metrics, run evaluations, and ship confidently with version control and peer review workflows built for AI engineering.GitHub Badge
roteOpen-source CLI that compiles a proven agent skill (a SKILL.md plus references) into a typed, deterministic pipeline. Fixed logic becomes reviewable Python or TypeScript with per-step tests, while judgment steps stay as typed LLM-judge signatures. Emits DBOS, Temporal, Cloudflare Workflows, Inngest, or plain Python/TS, and can serve compiled pipelines as MCP tools.GitHub Badge
RoundtableZero-configuration unified AI assistant management built on the FastMCP framework. Provides seamless integration with Claude, ChatGPT, and other AI assistants through a single MCP interface with session management, logging, and production-ready operations.GitHub Badge
PortkeyControl Panel with an observability suite & an AI gateway — to ship fast, reliable, and cost-efficient apps.
SAVI SDKOpen-source observability SDK for LLM cost, PII masking, carbon, and compliance — drop-in wrapper for OpenAI/Anthropic/Bedrock/Cohere/Mistral/Vertex; works standalone with zero account via local_mode.GitHub Badge
Self-Hosted AI StackDocker Compose stack for local AI with Ollama, a LiteLLM gateway, RAG, voice services, MCP tools, persistent data, and health checks.GitHub Badge
Semantic Cache RouterDistributed semantic cache and stateful routing system that cuts LLM API costs by returning cached responses for semantically similar queries. Uses ANN vector search (cosine ≥ 0.8) and consistent hashing to pin requests to the same worker, achieving ~7× latency reduction on cache hits while scaling horizontally without cache thrash.GitHub Badge
SpendlineFinancial control layer for AI spend — per-customer cost attribution and hierarchical budgets enforced before the provider call.
StatewaveOpen-source memory runtime for AI agents. Compiles events into deterministic, provenance-tagged context bundles instead of query-time retrieval. Apache-2.0, self-hostable on Postgres + pgvector.GitHub Badge
TensorZeroTensorZero is an open-source framework for building production-grade LLM applications. It unifies an LLM gateway, observability, optimization, evaluations, and experimentation.GitHub Badge
ThinkWatch LiteDesktop app for macOS, Windows and Linux that runs a local LLM gateway for Claude Code, Codex and other AI coding clients: routing rules with failover, conversion between the Anthropic, OpenAI and Gemini APIs, a record of each request's route and cost, and redaction of API keys before requests leave. The gateway engine, ThinkWatch Core, also runs on its own on a Linux server.GitHub Badge
TrinitySelf-hosted platform that runs AI coding agents as persistent scheduled services. Each agent runs in a Docker container with its own workspace and git-backed memory, plus scheduling, an approvals queue, an MCP server and a web console.GitHub Badge
UnoRouterOpenAI-compatible LLM gateway with one API key for every major provider and smart routing across models. Drop-in for code, Claude Code, and chat clients like SillyTavern, Janitor.AI, RisuAI, and Chub.
VellumAn AI product development platform to experiment with, evaluate, and deploy advanced LLM apps.
Weights & Biases (Prompts)A suite of LLMOps tools within the developer-first W&B MLOps platform. Utilize W&B Prompts for visualizing and inspecting LLM execution flow, tracking inputs and outputs, viewing intermediate results, securely managing prompts and LLM chain configurations.
WenlanLocal-first AI knowledge base and LLM wiki that distills agent work into source-cited pages with graph context and hybrid retrieval across MCP clients.GitHub Badge
WordwareA web-hosted IDE where non-technical domain experts work with AI Engineers to build task-specific AI agents. It approaches prompting as a new programming language rather than low/no-code blocks.
XiuRouterHosted multi-model API service with OpenAI Chat Completions and Responses, Anthropic Messages, and Gemini GenerateContent routes, scoped API keys, usage-based pricing, and request-level usage and cost records.
xTuringBuild and control your personal LLMs with fast and efficient fine-tuning.GitHub Badge
ZenMLOpen-source framework for orchestrating, experimenting and deploying production-grade ML solutions, with built-in langchain & llama_index integrations.GitHub Badge
SwarmClawSelf-hosted multi-agent AI runtime with 23+ LLM providers, persistent memory, skills, schedules, sub-agent spawning, and MCP client + server support. Ships as desktop app, CLI, or Docker.GitHub Badge
ai-evaluationEvaluation framework for automated, reproducible scoring of LLM, agent, and workflow performance.GitHub Badge
future-agiOpen-source self-hostable end-to-end agent engineering and optimization platform unifying tracing, evals, simulations, datasets, gateway, and guardrails for LLM and AI agent applications.GitHub Badge
ModelglassSourced, versioned pricing and capability data for AI models (image, language, video, audio, plus coding/science/agentic benchmark verticals) to find the cheapest model that clears a capability bar.

⬆ back to ToC

ProjectDetailsRepository
AirweaveAn easy way to turn any app into searchable data for LLMs.GitHub Badge
MemorySyncPersistent multi-tenant memory layer and MCP server for AI coding assistants with sub-50ms hybrid recall.GitHub Badge
ProjectDetailsRepository
AquilaDBAn easy to use Neural Search Engine. Index latent vectors along with JSON metadata and do efficient k-NN search.GitHub Badge
AwadbAI Native database for embedding vectorsGitHub Badge
Chromathe open source embedding databaseGitHub Badge
EpsillaA 10x faster, cheaper, and better vector databaseGitHub Badge
InfinityThe AI-native database built for LLM applications, providing incredibly fast vector and full-text searchGitHub Badge
InfinoEmbedded retrieval engine on Apache Parquet: BM25 full-text, vector, hybrid (RRF), and SQL from one engine over object storage.GitHub Badge
LancedbDeveloper-friendly, serverless vector database for AI applications. Easily add long-term memory to your LLM apps!GitHub Badge
MarqoTensor search for humans.GitHub Badge
MilvusVector database for scalable similarity search and AI applications.GitHub Badge
OmnigraphTyped graph database where agents branch and merge like Git. S3-native, Rust, traversal + vector + BM25 in one runtime.GitHub Badge
ParadeDBThe transactional alternative to Elasticsearch, built on Postgres.GitHub Badge
PineconeThe Pinecone vector database makes it easy to build high-performance vector search applications. Developer-friendly, fully managed, and easily scalable without infrastructure hassles.
pgvectorOpen-source vector similarity search for Postgres.GitHub Badge
RivestackManaged PostgreSQL with pgvector for AI workloads. Built-in SQL editor lets you query your database with natural language (auto-converted to vector embeddings). Free tier includes 2GB storage.
SynapCoresSelf-hosted AI-native database: vector search + Cypher graph + embedded GGUF inference and SQL in one engine. Free Community Edition binary; source proprietary.GitHub Badge
VectorChordScalable, fast, and disk-friendly vector search in Postgres, the successor of pgvecto.rs.GitHub Badge
pgvecto.rsVector database plugin for Postgres, written in Rust, specifically designed for LLM.GitHub Badge
QdrantVector Search Engine and Database for the next generation of AI applications. Also available in the cloudGitHub Badge
txtaiBuild AI-powered semantic search applicationsGitHub Badge
ValdA Highly Scalable Distributed Vector Search EngineGitHub Badge
VearchA distributed system for embedding-based vector retrievalGitHub Badge
VectorDBA Python vector database you just need - no more, no less.GitHub Badge
VellumA managed service for ingesting documents and performing hybrid semantic/keyword search across them. Comes with out-of-box support for OCR, text chunking, embedding model experimentation, metadata filtering, and production-grade APIs.
WeaviateWeaviate is an open source vector search engine that stores both objects and vectors, allowing for combining vector search with structured filtering with the fault-tolerance and scalability of a cloud-native database, all accessible through GraphQL, REST, and various language clients.GitHub Badge

⬆ back to ToC

Code AI

ProjectDetailsRepository
AgentsMeshSelf-hostable AI Agent Workforce Platform. Multi-agent orchestration with remote AI workstations (AgentPods), PTY sandbox + git worktree isolation, built-in Kanban, and per-pod MCP server. Supports Claude Code, Codex CLI, Gemini CLI, Aider, OpenCode.GitHub Badge
Atomic AgentLocal-first CLI and TUI coding agent that runs open-weight models entirely on your machine through a llama.cpp fork. No account or API key required. 56 built-in tools (browser, filesystem, git, memory, vision), MCP support, five-layer local memory, macOS/Linux/Windows.GitHub Badge
BernsteinDeterministic Python orchestrator for 37 CLI coding agents (Claude Code, Codex CLI, Gemini CLI, GitHub Copilot CLI, Cursor, Aider, OpenHands, OpenCode, Goose, Qwen, Ollama, ...) running in parallel git worktrees. First-class MCP server, quality gates, cost tracking with budgets.GitHub Badge
CodeGeeXCodeGeeX: An Open Multilingual Code Generation Model (KDD 2023)GitHub Badge
CodeGenCodeGen is an open-source model for program synthesis. Trained on TPU-v4. Competitive with OpenAI Codex.GitHub Badge
Coder EvalFramework for evaluating, benchmarking, and A/B-testing AI coding agents (Claude Code, Codex CLI, Gemini/Antigravity) and their skills. Runs a real agent in a sandbox against declarative YAML tasks, then scores the resulting files and commands with weighted 0.0-1.0 criteria. Includes 14 criterion types, per-tool token/cost telemetry, dataset fan-out with skill-activation precision/recall gates, and a GitHub Action with JUnit XML output for CI.GitHub Badge
CodeT5Open Code LLMs for Code Understanding and Generation.GitHub Badge
Continue⏩ the open-source autopilot for software development—bring the power of ChatGPT to VS CodeGitHub Badge
CotalOpen pub/sub standard over NATS JetStream for coordinating coding agents. Claude Code, Codex, OpenCode, Hermes, Jcode and pi agents share a space with presence, channels, durable direct messages and role-addressed delivery.GitHub Badge
DSH StudioCross-platform desktop host for installing, configuring, and managing DeepSeek Harness.GitHub Badge
fauxpilotAn open-source alternative to GitHub Copilot serverGitHub Badge
fractalHierarchical coding-agent orchestrator with recursive delegation, per-node Git worktrees, configurable limits, persistent SQLite state, and live terminal monitoring and steering.GitHub Badge
Kolega CodePython terminal coding agent where the model writes its own multi-agent workflows (Gigacode). Provider-agnostic, local-first, 15+ model providers, MCP client, browser agent.GitHub Badge
promptextSmart code context extractor for AI assistants with accurate token counting and budget managementGitHub Badge
RelayUniversal AI API proxy — hot-switch between 14+ LLM providers (GitHub Copilot, OpenAI, Anthropic, DeepSeek, Groq, Ollama, and more) from any AI coding agent without restarting sessions. Local proxy with automatic Anthropic <-> OpenAI protocol translation, account rotation, and real-time usage tracking.GitHub Badge
SuperagentOpen-source macOS desktop app that gives Claude Code and Codex a real browser to drive, an iOS Simulator to install and screenshot apps in, and a phone companion app for remote monitoring.GitHub Badge
tabbySelf-hosted AI coding assistant. An opensource / on-prem alternative to GitHub Copilot.GitHub Badge
AIDEOpen-source ML engineering agent that uses tree search to explore solution spaces. Automates machine learning experimentation from data analysis to model training. Paper.GitHub Badge
KapsoLong-running agents that optimize AI and Data systems, and learn from every experience. #1 open-source on MLE-Bench; ALE-Bench; RelBench.GitHub Badge
webcmdSelf-learning browser infrastructure for AI coding agents that compiles site navigation into deterministic per-site CLI commands.GitHub Badge

Training

IDEs and Workspaces

ProjectDetailsRepository
code serverRun VS Code on any machine anywhere and access it in the browser.GitHub Badge
condaOS-agnostic, system-level binary package manager and ecosystem.GitHub Badge
DockerMoby is an open-source project created by Docker to enable and accelerate software containerization.GitHub Badge
envd🏕️ Reproducible development environment for AI/ML.GitHub Badge
Jupyter NotebooksThe Jupyter notebook is a web-based notebook environment for interactive computing.GitHub Badge
KurtosisA build, packaging, and run system for ephemeral multi-container environments.GitHub Badge
LayerSmithSelf-hosted web UI, TUI, and CLI for building OCI images with Docker or Podman, including LLM training environments and air-gap bundles.GitHub Badge
WordwareA web-hosted IDE where non-technical domain experts work with AI Engineers to build task-specific AI agents. It approaches prompting as a new programming language rather than low/no-code blocks.

⬆ back to ToC

Foundation Model Fine Tuning

ProjectDetailsRepository
alpaca-loraInstruct-tune LLaMA on consumer hardwareGitHub Badge
finetuning-schedulerA PyTorch Lightning extension that accelerates and enhances foundation model experimentation with flexible fine-tuning schedules.GitHub Badge
FlyflowOpen source, high performance fine tuning as a service for GPT4 quality models with 5x lower latency and 3x lower costGitHub Badge
LMFlowAn Extensible Toolkit for Finetuning and Inference of Large Foundation ModelsGitHub Badge
LoraUsing Low-rank adaptation to quickly fine-tune diffusion models.GitHub Badge
peftState-of-the-art Parameter-Efficient Fine-Tuning.GitHub Badge
p-tuning-v2An optimized prompt tuning strategy achieving comparable performance to fine-tuning on small/medium-sized models and sequence tagging challenges. (ACL 2022)GitHub Badge
QLoRAEfficient finetuning approach that reduces memory usage enough to finetune a 65B parameter model on a single 48GB GPU while preserving full 16-bit finetuning task performance.GitHub Badge
TrainJudgeDiagnoses whether fine-tuning fits, then verifies fine-tunes on held-out task metrics and a regression suite, not training loss.GitHub Badge
TRLTrain transformer language models with reinforcement learning.GitHub Badge

⬆ back to ToC

Frameworks for Training

ProjectDetailsRepository
Accelerate🚀 A simple way to train and use PyTorch models with multi-GPU, TPU, mixed-precision.GitHub Badge
Apache MXNetLightweight, Portable, Flexible Distributed/Mobile Deep Learning with Dynamic, Mutation-aware Dataflow Dep Scheduler.GitHub Badge
axolotlA tool designed to streamline the fine-tuning of various AI models, offering support for multiple configurations and architectures.GitHub Badge
CaffeA fast open framework for deep learning.GitHub Badge
CandleMinimalist ML framework for Rust .GitHub Badge
ColossalAIAn integrated large-scale model training system with efficient parallelization techniques.GitHub Badge
DeepSpeedDeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.GitHub Badge
HorovodDistributed training framework for TensorFlow, Keras, PyTorch, and Apache MXNet.GitHub Badge
JaxAutograd and XLA for high-performance machine learning research.GitHub Badge
KedroKedro is an open-source Python framework for creating reproducible, maintainable and modular data science code.GitHub Badge
KerasKeras is a deep learning API written in Python, running on top of the machine learning platform TensorFlow.GitHub Badge
LightGBMA fast, distributed, high performance gradient boosting (GBT, GBDT, GBRT, GBM or MART) framework based on decision tree algorithms, used for ranking, classification and many other machine learning tasks.GitHub Badge
MegEngineMegEngine is a fast, scalable and easy-to-use deep learning framework, with auto-differentiation.GitHub Badge
metric-learnMetric Learning Algorithms in Python.GitHub Badge
MindSporeMindSpore is a new open source deep learning training/inference framework that could be used for mobile, edge and cloud scenarios.GitHub Badge
OneflowOneFlow is a performance-centered and open-source deep learning framework.GitHub Badge
PaddlePaddleMachine Learning Framework from Industrial Practice.GitHub Badge
PyTorchTensors and Dynamic neural networks in Python with strong GPU acceleration.GitHub Badge
PyTorch LightningDeep learning framework to train, deploy, and ship AI products Lightning fast.GitHub Badge
XGBoostScalable, Portable and Distributed Gradient Boosting (GBDT, GBRT or GBM) Library.GitHub Badge
scikit-learnMachine Learning in Python.GitHub Badge
TensorFlowAn Open Source Machine Learning Framework for Everyone.GitHub Badge
VectorFlowA minimalist neural network library optimized for sparse data and single machine environments.GitHub Badge

⬆ back to ToC

Experiment Tracking

ProjectDetailsRepository
Aiman easy-to-use and performant open-source experiment tracker.GitHub Badge
ClearMLAuto-Magical CI/CD to streamline your ML workflow. Experiment Manager, MLOps and Data-ManagementGitHub Badge
CometComet is an MLOps platform that offers experiment tracking, model production management, a model registry, and full data lineage from training straight through to production. Comet plays nicely with all your favorite tools, so you don't have to change your existing workflow. Comet Opik to confidently evaluate, test, and ship LLM applications with a suite of observability tools to calibrate language model outputs across your dev and production lifecycle!GitHub Badge
Guild AIExperiment tracking, ML developer tools.GitHub Badge
MLRunMachine Learning automation and tracking.GitHub Badge
Kedro-VizKedro-Viz is an interactive development tool for building data science pipelines with Kedro. Kedro-Viz also allows users to view and compare different runs in the Kedro project.GitHub Badge
LabNotebookLabNotebook is a tool that allows you to flexibly monitor, record, save, and query all your machine learning experiments.GitHub Badge
SacredSacred is a tool to help you configure, organize, log and reproduce experiments.GitHub Badge
Weights & BiasesA developer first, lightweight, user-friendly experiment tracking and visualization tool for machine learning projects, streamlining collaboration and simplifying MLOps. W&B excels at tracking LLM-powered applications, featuring W&B Prompts for LLM execution flow visualization, input and output monitoring, and secure management of prompts and LLM chain configurations.GitHub Badge

⬆ back to ToC

Visualization

ProjectDetailsRepository
Fiddler AIRich dashboards, reports, and UMAP to perform root cause analysis, pinpoint problem areas, like correctness, safety, and privacy issues, and improve LLM outcomes.
LangWatchVisualize LLM evaluations experiments and DSPy pipeline optimizationsGitHub Badge
ManifordA model-agnostic visual debugging tool for machine learning.GitHub Badge
netronVisualizer for neural network, deep learning, and machine learning models.GitHub Badge
OpenOpsBring multiple data streams into one dashboard.GitHub Badge
TensorBoardTensorFlow's Visualization Toolkit.GitHub Badge
TensorSpaceNeural network 3D visualization framework, build interactive and intuitive model in browsers, support pre-trained deep learning models from TensorFlow, Keras, TensorFlow.js.GitHub Badge
dtreevizA python library for decision tree visualization and model interpretation.GitHub Badge
Zetane ViewerML models and internal tensors 3D visualizer.GitHub Badge
ZenoAI evaluation platform for interactively exploring data and model outputs.GitHub Badge

Model Editing

ProjectDetailsRepository
FastEditFastEdit aims to assist developers with injecting fresh and customized knowledge into large language models efficiently using one single command.GitHub Badge

⬆ back to ToC

Data

Data Management

ProjectDetailsRepository
ArtiVCA version control system to manage large files. Lake is a dataset format with a simple API for creating, storing, and collaborating on AI datasets of any size.GitHub Badge
DoltGit for Data.GitHub Badge
DVCData Version Control - Git for Data & Models - ML Experiments Management.GitHub Badge
Delta-LakeStorage layer that brings scalable, ACID transactions to Apache Spark and other engines.GitHub Badge
PachydermPachyderm is a version control system for data.GitHub Badge
QuiltA self-organizing data hub for S3.GitHub Badge

⬆ back to ToC

Data Storage

ProjectDetailsRepository
JuiceFSA distributed POSIX file system built on top of Redis and S3.GitHub Badge
LakeFSGit-like capabilities for your object storage.GitHub Badge
LanceModern columnar data format for ML implemented in Rust.GitHub Badge
PixeltableDeclarative multimodal AI data engine for versioned tables, computed columns, and vector search.GitHub Badge

⬆ back to ToC

Data Tracking

ProjectDetailsRepository
PiperiderA CLI tool that allows you to build data profiles and write assertion tests for easily evaluating and tracking your data's reliability over time.GitHub Badge
LUXA Python library that facilitates fast and easy data exploration by automating the visualization and data analysis process.GitHub Badge
ragfreshA CLI that detects stale, drifted, and ghost documents in RAG vector indexes by diffing content hashes and embedding drift against the source of truth.GitHub Badge

⬆ back to ToC

Feature Engineering

ProjectDetailsRepository
FeatureformThe Virtual Feature Store. Turn your existing data infrastructure into a feature store.GitHub Badge
FeatureToolsAn open source python framework for automated feature engineeringGitHub Badge

⬆ back to ToC

Data/Feature enrichment

ProjectDetailsRepository
CocoIndexAn ETL framework for AI that transforms data into embeddings and knowledge graphs, with incremental processing to recompute only what changed and keep indexes fresh.GitHub Badge
UpginiFree automated data & feature enrichment library for machine learning: automatically searches through thousands of ready-to-use features from public and community shared data sources and enriches your training dataset with only the accuracy improving featuresGitHub Badge
FeastAn open source feature store for machine learning.GitHub Badge
distilabel⚗️ distilabel is a framework for synthetic data and AI feedback for AI engineers that require high-quality outputs, full data ownership, and overall efficiency.GitHub Badge
FastDatasetsA powerful tool for creating high-quality training datasets for Large Language Models.GitHub Badge
fastdocparseExtracts structured data from documents with an LLM, then grounds every field against the source text to flag hallucinated values automatically.GitHub Badge

⬆ back to ToC

Large Scale Deployment

ML Platforms

ProjectDetailsRepository
CometComet is an MLOps platform that offers experiment tracking, model production management, a model registry, and full data lineage from training straight through to production. Comet plays nicely with all your favorite tools, so you don't have to change your existing workflow. Comet Opik to confidently evaluate, test, and ship LLM applications with a suite of observability tools to calibrate language model outputs across your dev and production lifecycle!GitHub Badge
ClearMLAuto-Magical CI/CD to streamline your ML workflow. Experiment Manager, MLOps and Data-Management.GitHub Badge
dstackOpen-source confidential AI framework for secure LLM deployment with data privacy, providing hardware-enforced isolation for production ML workloads.GitHub Badge
HopsworksHopsworks is a MLOps platform for training and operating large and small ML systems, including fine-tuning and serving LLMs. Hopsworks includes both a feature store and vector database for RAG.GitHub Badge
OpenLLMAn open platform for operating large language models (LLMs) in production. Fine-tune, serve, deploy, and monitor any LLMs with ease.GitHub Badge
MLflowOpen source platform for the machine learning lifecycle.GitHub Badge
MLRunAn open MLOps platform for quickly building and managing continuous ML applications across their lifecycle.GitHub Badge
ModelFoxModelFox is a platform for managing and deploying machine learning models.GitHub Badge
KserveStandardized Serverless ML Inference Platform on KubernetesGitHub Badge
KubeStellar ConsoleOpen source AI-powered multi-cluster Kubernetes dashboard for managing LLM workloads across hybrid edge and cloud environments. GPU monitoring, benchmark streaming, real-time observability with 20+ CNCF integrations, and AI-guided cluster operations. CNCF Sandbox project.GitHub Badge
KubeflowMachine Learning Toolkit for Kubernetes.GitHub Badge
PAIResource scheduling and cluster management for AI.GitHub Badge
piqcOpen-source, read-only GPU waste scanner for Kubernetes inference clusters. Deploys in minutes, no write permissions required.GitHub Badge
PolyaxonMachine Learning Management & Orchestration Platform.GitHub Badge
PrimehubAn effortless infrastructure for machine learning built on the top of Kubernetes.GitHub Badge
OpenModelZOne-click machine learning deployment (LLM, text-to-image and so on) at scale on any cluster (GCP, AWS, Lambda labs, your home lab, or even a single machine).GitHub Badge
Seldon-coreAn MLOps framework to package, deploy, monitor and manage thousands of production machine learning modelsGitHub Badge
StarwhaleAn MLOps/LLMOps platform for model building, evaluation, and fine-tuning.GitHub Badge
TrueFoundryA PaaS to deploy, Fine-tune and serve LLM Models on a company’s own Infrastructure with Data Security and Optimal GPU and Cost Management. Launch your LLM Application at Production scale with best DevSecOps practices.
Weights & BiasesA lightweight and flexible platform for machine learning experiment tracking, dataset versioning, and model management, enhancing collaboration and streamlining MLOps workflows. W&B excels at tracking LLM-powered applications, featuring W&B Prompts for LLM execution flow visualization, input and output monitoring, and secure management of prompts and LLM chain configurations.GitHub Badge

⬆ back to ToC

Workflow

ProjectDetailsRepository
AirflowA platform to programmatically author, schedule and monitor workflows.GitHub Badge
aqueductAn Open-Source Platform for Production Data ScienceGitHub Badge
Argo WorkflowsWorkflow engine for Kubernetes.GitHub Badge
awaithumansOne-function HITL primitive for AI agents. await_human blocks like a Promise; a human reviews via Slack/email/dashboard; the agent resumes with the typed response. Durable across restarts via Stripe-style idempotency keys. Temporal + LangGraph adapters included.GitHub Badge
FlyteKubernetes-native workflow automation platform for complex, mission-critical data and ML processes at scale.GitHub Badge
HamiltonA lightweight framework to represent ML/language model pipelines as a series of python functions.GitHub Badge
HeymSource-available, self-hosted AI workflow automation platform with a visual canvas for agent, RAG, and tool-using workflows. Includes MCP support, evals, traces, and cost tracking.GitHub Badge
KitaruDurable execution layer for AI agents. Checkpoints, replay, resume, and observability primitives that make agent workflows persistent and replayable — no graph DSL required.GitHub Badge
Kubeflow PipelinesMachine Learning Pipelines for Kubeflow.GitHub Badge
LangFlowAn effortless way to experiment and prototype LangChain flows with drag-and-drop components and a chat interface.GitHub Badge
MetaflowBuild and manage real-life data science projects with ease!GitHub Badge
PloomberThe fastest way to build data pipelines. Develop iteratively, deploy anywhere.GitHub Badge
PrefectThe easiest way to automate your data.GitHub Badge
VDPAn open-source unstructured data ETL tool to streamline the end-to-end unstructured data processing pipeline.GitHub Badge
ZenMLMLOps framework to create reproducible pipelines.GitHub Badge
simulate-sdkEnterprise-grade Voice AI simulation SDK for scenario-driven stress testing of multimodal and agentic systems.GitHub Badge

⬆ back to ToC

Scheduling

ProjectDetailsRepository
KueueKubernetes-native Job Queueing.GitHub Badge
PAIResource scheduling and cluster management for AI (Open-sourced by Microsoft).GitHub Badge
SlurmA Highly Scalable Workload Manager.GitHub Badge
VolcanoA Cloud Native Batch System (Project under CNCF).GitHub Badge
YunikornLight-weight, universal resource scheduler for container orchestrator systems.GitHub Badge

⬆ back to ToC

Model Management

ProjectDetailsRepository
CometComet is an MLOps platform that offers Model Production Management, a Model Registry, and full model lineage from training straight through to production. Use Comet for model reproducibility, model debugging, model versioning, model visibility, model auditing, model governance, and model monitoring.GitHub Badge
dvcML Experiments Management - Data Version Control - Git for Data & Models![GitHub Badge](https://img.shields.io/github/stars/iterative/dvc.svg?style=flat-square

Truncated — view the full README on GitHub.

ai-development-tools
awesome-list
llmops
mlops

Significant stargazers

(top 24 of 38)

Marc Klingen

462 followers · starred Nov 2024

Guangdong Liu

102 followers · starred Aug 2024

Utku Demir

159 followers · starred Apr 2023

Yuki Iwai

188 followers · starred May 2022

tensorchord/Awesome-LLMOps

An awesome & curated list of best LLMOps tools for developers

Shell

5,952

319 commits

updated Oct 5, 2026

See the code

README

Awesome LLMOps

discord invitation link

An awesome & curated list of the best LLMOps tools for developers.

[!NOTE] Contributions are most welcome, please adhere to the contribution guidelines.

Table of Contents

Model

Large Language Model

ProjectDetailsRepository
AlpacaCode and documentation to train Stanford's Alpaca models, and generate the data.GitHub Badge
BELLEA 7B Large Language Model fine-tune by 34B Chinese Character Corpus, based on LLaMA and Alpaca.GitHub Badge
BloomBigScience Large Open-science Open-access Multilingual Language ModelGitHub Badge
dollyDatabricks’ Dolly, a large language model trained on the Databricks Machine Learning PlatformGitHub Badge
Falcon 40BFalcon-40B-Instruct is a 40B parameters causal decoder-only model built by TII based on Falcon-40B and finetuned on a mixture of Baize. It is made available under the Apache 2.0 license.
FastChat (Vicuna)An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and FastChat-T5.GitHub Badge
GemmaGemma is a family of lightweight, open models built from the research and technology that Google used to create the Gemini models.
GLM-6B (ChatGLM)An Open Bilingual Pre-Trained Model, quantization of ChatGLM-130B, can run on consumer-level GPUs.GitHub Badge
ChatGLM2-6BChatGLM2-6B is the second-generation version of the open-source bilingual (Chinese-English) chat model ChatGLM-6B.GitHub Badge
GLM-130B (ChatGLM)An Open Bilingual Pre-Trained Model (ICLR 2023)GitHub Badge
GPT-NeoXAn implementation of model parallel autoregressive transformers on GPUs, based on the DeepSpeed library.GitHub Badge
JebadiahOpen decision models that return a probability for each option instead of generating text; available in 4B, 9B and 27B sizes.GitHub Badge
LuotuoA Chinese LLM, Based on LLaMA and fine tune by Stanford Alpaca, Alpaca LoRA, Japanese-Alpaca-LoRA.GitHub Badge
Mixtral-8x7B-v0.1The Mixtral-8x7B Large Language Model (LLM) is a pretrained generative Sparse Mixture of Experts.
StableLMStableLM: Stability AI Language ModelsGitHub Badge

⬆ back to ToC

CV Foundation Model

ProjectDetailsRepository
disco-diffusionA frankensteinian amalgamation of notebooks, models and techniques for the generation of AI Art and Animations.GitHub Badge
midjourneyMidjourney is an independent research lab exploring new mediums of thought and expanding the imaginative powers of the human species.
segment-anything (SAM)produces high quality object masks from input prompts such as points or boxes, and it can be used to generate masks for all objects in an image.GitHub Badge
stable-diffusionA latent text-to-image diffusion modelGitHub Badge

⬆ back to ToC

Audio Foundation Model

ProjectDetailsRepository
barkBark is a transformer-based text-to-audio model created by Suno. Bark can generate highly realistic, multilingual speech as well as other audio - including music, background noise and simple sound effects.GitHub Badge
FunASRSpeech recognition toolkit with pretrained models and tools for voice activity detection, punctuation restoration, and speaker diarization.GitHub Badge
whisperRobust Speech Recognition via Large-Scale Weak SupervisionGitHub Badge

⬆ back to ToC

Robotics Foundation Model

[!NOTE] Emerging Architectures in VLA:

  • Continuous Diffusion Language Models: Integrate diffusion heads or flow-matching to VLMs (e.g., DiVLA, OpenPI), enabling smooth, precise continuous action generation rather than discretized tokens.
  • Recurrent Language Models: Utilize State Space Models (SSMs) like Mamba or recurrent transformers (e.g., RoboMamba, RD-VLA) to reduce inference memory and handle temporal dependencies, allowing iterative reasoning for complex robotic decision-making.
ProjectDetailsRepository
DiVLAA continuous diffusion-based Vision-Language-Action model that integrates diffusion policies into autoregressive VLMs for robust and precise continuous robotic control.GitHub Badge
LeRobotA central community library by Hugging Face for AI in robotics — end-to-end learning tools, data pipelines, and support for training/deploying VLA models.GitHub Badge
OctoA transformer-based generalist robot policy pretrained on 800K+ robot trajectories from the Open X-Embodiment dataset. Supports language instructions, goal images, and fine-tuning to new embodiments.GitHub Badge
OpenPIOpen-source VLA models from Physical Intelligence, including π₀ and π₀.5 — flow-based vision-language-action models pretrained on large-scale robot data with fine-tuning support.GitHub Badge
OpenVLAA 7B-parameter open-source Vision-Language-Action model trained on 970K+ robot demonstrations from the Open X-Embodiment dataset for generalist robotic manipulation.GitHub Badge
RoboMambaAn efficient VLA model leveraging State Space Models (Mamba) instead of standard self-attention, offering linear inference complexity for efficient, recurrent robotic reasoning.GitHub Badge
SmolVLAA compact ~450M parameter VLA by Hugging Face, designed to be computationally efficient and accessible, running on consumer GPUs or CPUs. Part of the LeRobot ecosystem.

Serving

Large Model Serving

ProjectDetailsRepository
Alpaca-LoRA-ServeAlpaca-LoRA as Chatbot serviceGitHub Badge
HiggsRust inference server for Apple Silicon: MLX models behind OpenAI and Anthropic APIs, routing to remote providers, desktop dashboard.GitHub Badge
OneCompFujitsu Research's post-training quantization pipeline for LLMs (QEP, AutoBit, JointQ, rotation) with vLLM plugin (arXiv:2603.28845).GitHub Badge
CTranslate2fast inference engine for Transformer models in C++GitHub Badge
Clip-as-a-serviceserving the OpenAI CLIP modelGitHub Badge
DeepSpeed-MIIMII makes low-latency and high-throughput inference possible, powered by DeepSpeed.GitHub Badge
Faster Whisperfast inference engine for whisper in C++ using CTranslate2.GitHub Badge
FlexGenRunning large language models on a single GPU for throughput-oriented scenarios. (Archived)GitHub Badge
FlowiseDrag & drop UI to build your customized LLM flow using LangchainJS.GitHub Badge
lilbeeSingle-binary local manager and search engine that stands up a llama-server fleet sized to your GPUs by gguf-parser for VRAM-aware multi-GPU placementGitHub Badge
llama.cppPort of Facebook's LLaMA model in C/C++GitHub Badge
LLMKubeKubernetes operator for LLM inference with pluggable runtimes (llama.cpp, PersonaPlex/Moshi, generic), multi-GPU sharding, NVIDIA CUDA and Apple Silicon Metal support, and GGUF/MLX/SafeTensors model formats.GitHub Badge
ShimmyPython-free Rust inference server with OpenAI API compatibility and hot model swappingGitHub Badge
InfinityRest API server for serving text-embeddingsGitHub Badge
Modelz-LLMOpenAI compatible API for LLMs and embeddings (LLaMA, Vicuna, ChatGLM and many others)GitHub Badge
NobodyWhoOn-device LLM inference engine (Rust / llama.cpp) to embed local models directly in games and apps, with Godot / Flutter / React Native / Swift bindings; streaming, embeddings, tool calling, GBNF-structured output, and STT/TTS.GitHub Badge
Off GridOpen-source iOS/Android app running LLMs on-device via llama.cpp. Voice (Whisper), vision, image gen, tool calling — fully offline.GitHub Badge
OllamaServe Llama 2 and other large language models locally from command line or through a browser interface.GitHub Badge
Rapid-MLXOpenAI-compatible LLM inference server for Apple Silicon using MLX. 2-4x faster than Ollama with tool calling and prompt caching.GitHub Badge
TensorRT-LLMInference engine for TensorRT on Nvidia GPUsGitHub Badge
text-generation-inferenceLarge Language Model Text Generation InferenceGitHub Badge
text-embeddings-inferenceInference for text-embedding modelsGitHub Badge
tokenizers💥 Fast State-of-the-Art Tokenizers optimized for Research and ProductionGitHub Badge
vllmA high-throughput and memory-efficient inference and serving engine for LLMs.GitHub stars
whisper-ctranslate2is a 4x faster and low-memory usage drop-in cli replacement that supports word-level timestamps and VAD filterGitHub Badge
whisper.cppPort of OpenAI's Whisper model in C/C++GitHub Badge
x-stable-diffusionReal-time inference for Stable Diffusion - 0.88s latency. Covers AITemplate, nvFuser, TensorRT, FlashAttention. (Archived)GitHub Badge

⬆ back to ToC

Frameworks/Servers for Serving

ProjectDetailsRepository
BentoMLThe Unified Model Serving FrameworkGitHub Badge
BifrostHigh-performance AI gateway written in Go. Unifies 23+ providers behind one OpenAI-compatible endpoint with automatic failover, adaptive load balancing, semantic caching, virtual keys and budgets, MCP tool calling, and native Prometheus observability.GitHub Badge
JinaBuild multimodal AI services via cloud native technologies · Model Serving · Generative AI · Neural Search · Cloud NativeGitHub Badge
MosecA machine learning model serving framework with dynamic batching and pipelined stages, provides an easy-to-use Python interface.GitHub Badge
mcpproxy-goOpen-source MCP proxy with BM25 tool filtering, quarantine security, activity logging, and web UI. Routes multiple MCP servers through single endpoint, reducing context bloat by ~97%.GitHub Badge
SentryNode GatewayOpen-core AI inference governance and FinOps platform providing semantic routing, <40ms agent loop intervention, strict budget caps, and cryptographic audit trails. Includes voice pipeline governance (STT→LLM→TTS).GitHub Badge
TFServingA flexible, high-performance serving system for machine learning models.GitHub Badge
TorchserveServe, optimize and scale PyTorch models in production (Archived)GitHub Badge
Triton Server (TRTIS)The Triton Inference Server provides an optimized cloud and edge inferencing solution.GitHub Badge
langchain-serveServerless LLM apps on Production with Jina AI Cloud (Archived)GitHub Badge
lanarkyFastAPI framework to build production-grade LLM applicationsGitHub Badge
ray-llmLLMs on Ray - RayLLM (Archived)GitHub Badge
SIEOpen-source inference server and production cluster for serving embeddings, reranking, OCR, structured output, content safety, and generative models through one API.GitHub Badge
XinferenceReplace OpenAI GPT with another LLM in your app by changing a single line of code. Xinference gives you the freedom to use any LLM you need. With Xinference, you're empowered to run inference with any open-source language models, speech recognition models, and multimodal models, whether in the cloud, on-premises, or even on your laptop.GitHub Badge
KubeAIDeploy and scale machine learning models on Kubernetes. Built for LLMs, embeddings, and speech-to-text.GitHub Badge
KaitoA Kubernetes operator that simplifies serving and tuning large AI models (e.g. Falcon or phi-3) using container images and GPU auto-provisioning. Includes an OpenAI-compatible server for inference and preset configurations for popular runtimes such as vLLM and transformers.GitHub Badge
Open ResponsesServerless open-source platform for building long-running LLM agents with tool use.GitHub Badge
KubeStellar ConsoleAI-powered multi-cluster Kubernetes dashboard for hybrid edge and cloud. GPU monitoring, LLM inference cluster management, benchmark streaming, and 20+ CNCF integrations. CNCF Sandbox (Apache 2.0).GitHub Badge
laya-decision-apiFastAPI HTTP service for the local Laya System-1 decision model: offline structured decisions (choice / score / yes-no) with probabilities and confidence; MLX and torch backendsGitHub stars

⬆ back to ToC

Security

Frameworks for LLM security

ProjectDetailsRepository
API Relay AuditLocal security audit for AI API relays and LLM proxies; checks prompt injection, model identity drift, tool-call rewriting, error leakage, SSE anomalies, and Web3 wallet probes.GitHub Badge
CordumSafety-first agent orchestration platform with pre-dispatch policy evaluation, output scanning (PII, secrets, injection), job scheduling, workflow engine, and full audit trail.GitHub Badge
brood-boxCLI tool for running coding agents inside hardware-isolated microVMs with snapshot isolation, egress control, and MCP authorization.GitHub Badge
DOS KernelDeterministic trust kernel for AI-agent fleets: verifies agent "done" claims from git evidence rather than self-report, arbitrates concurrent file access between agents, and audits commit messages against their own diffs.GitHub Badge
dstackOpen-source confidential AI framework for secure LLM deployment with data privacy, providing hardware-enforced isolation using Intel TDX and NVIDIA Confidential Computing.GitHub Badge
EmbedGuardCross-layer detection and provenance attestation for adversarial embedding attacks in RAG systems. Flags poisoned or manipulated embeddings before they reach the retrieval layer.GitHub Badge
PlexiglassA Python Machine Learning Pentesting Toolbox for Adversarial Attacks. Works with LLMs, DNNs, and other machine learning algorithms.GitHub Badge
PrivacyScrubberZero-Trust client-side PII and developer secrets sanitizer and MCP server for Cursor, Windsurf, and Claude Desktop with sub-2ms in-memory tokenization.GitHub Badge
PrivAiTeSelf-hosted OpenAI-compatible proxy that reversibly pseudonymizes PII before requests reach the provider, tool-call arguments and multimodal included.GitHub Badge
shimOpen-core AI gateway that redacts PII before prompts reach model providers (Presidio plus Turkish TCKN/VKN recognizers). The source-available enterprise edition adds a tamper-evident hash-chain audit log with daily Merkle anchors as evidence for EU AI Act and KVKK record-keeping.GitHub Badge
wardcatOn-premises PII and sensitive-data detection and masking for LLM and RAG workflows, combining regex, NER, and optional local LLMs with reversible masking.GitHub Badge

⬆ back to ToC

Observability

ProjectDetailsRepository
AgentCostTrack and optimize LLM API costs across OpenAI, Anthropic, and LangChain.GitHub Badge
Azure OpenAI Logger"Batteries included" logging solution for your Azure OpenAI instance.GitHub Badge
budget-guardZero-infra hard spending caps for LLM APIs — blocks before overspending (atomic reservation), per-feature cost attribution, streaming usage, retry-storm detection, and adapters for Vercel AI SDK / LangChain.js / LlamaIndex.TS / Mastra.GitHub Badge
ClevAgentRuntime monitoring for AI agents — heartbeat watchdog, loop detection, cost tracking, auto-restart. Python SDK or HTTP API.
CodeBurnLocal-first cost and token tracking for AI coding agents (Claude Code, Codex, Cursor and 38 more) by project, model and task, with waste detection and plan quota tracking.GitHub Badge
DeepchecksTests for Continuous Validation of ML Models & Data. Deepchecks is a Python package for comprehensively validating your machine learning models and data with minimal effort.GitHub Badge
diglineRegression testing for LLM apps: approved scores, prompt and commit pinned as a versioned baseline in your repo; compare says which case got worse and by how much.GitHub Badge
EvidentlyAn open-source framework to evaluate, test and monitor ML and LLM-powered systems.GitHub Badge
EvalViewRegression testing for AI agents. Snapshot behavior, detect tool-call and output regressions, with golden-baseline diffing and LLM-as-judge scoring. Supports LangGraph, CrewAI, OpenAI, Claude, and any HTTP API.GitHub Badge
Fiddler AIEvaluate, monitor, analyze, and improve machine learning and generative models from pre-production to production. Ship more ML and LLMs into production, and monitor ML and LLM metrics like hallucination, PII, and toxicity.GitHub Badge
FlowlinesBehavioral observability for production AI agents that detects drift, context loss, constraint violations, and user frustration across sessions.
GiskardTesting framework dedicated to ML models, from tabular to LLMs. Detect risks of biases, performance issues and errors in 4 lines of code.GitHub Badge
QWEDDeterministic verification protocol for LLM outputs using 8 formal verification engines (SymPy, Z3, AST, SQLGlot). Prevents hallucinations through mathematical proofs rather than statistical methods.GitHub Badge
Great ExpectationsAlways know what to expect from your data.GitHub Badge
HeliconeOpen source LLM observability platform. One line of code to monitor, evaluate, and experiment with features like prompt management, agent tracing, and evaluations.GitHub Badge
hermes-rubricEvidence-first assessment for agent outputs and applications. Synthesizes, loads, or accepts a rubric; collects citations for each dimension; scores against accepted evidence; reports coverage; and emits reproducibility receipts. Apache-2.0.GitHub Badge
InferrailPer-run dollar budgets for AI agents and jobs: an OpenAI- and Anthropic-compatible gateway (in-process or self-hosted) that reserves each call's worst-case cost before it is sent, refuses calls that don't fit before the provider, and keeps a payload-free cost record per run.GitHub Badge
Traceloop OpenLLMetryOpenTelemetry-based observability and monitoring for LLM and agents workflows.GitHub Badge
Langfuse 🪢Open-source LLM observability platform that helps teams collaboratively debug, analyze, and iterate on their LLM applications.GitHub Badge
LookspanLocal-first observability for AI agents. One command (npx lookspan) runs a dashboard with traces, a timeline/waterfall view, cost tracking, replay & LLM-as-judge, and datasets. MCP-native, with OpenAI/Anthropic drop-ins and an OpenTelemetry receiver — data stays on your machine.GitHub Badge
whylogsThe open standard for data loggingGitHub Badge
Maxim AIPlatform for AI Agent Simulation, Evaluation & Observability
NoveumAI agent tracing, evaluation datasets from production interactions, and model comparisons on quality, latency, and cost.
onWatchLightweight Go CLI that tracks AI API quota usage across 7 providers (Anthropic, OpenAI, GitHub Copilot, MiniMax, and more). Background daemon, <50MB RAM, zero telemetry, SQLite storage.GitHub Badge
OpenGATEDeterministic verification for evidence-grounded AI — gold-anchored grounding checks (required facts, ungrounded numbers, abstention) with no LLM judge, plus CI regression gates, pytest/DeepEval integrations, and an MCP server.GitHub Badge
phoenix2pytestReads Arize Phoenix traces flagged as failures and synthesises pytest regression tests, so production-discovered LLM failures feed back into the test suite without manual translation.GitHub Badge
RagTuneCLI tool for debugging and benchmarking RAG retrieval. EXPLAIN ANALYZE for your retrieval layer.GitHub Badge
runtapeFinds which part of an agent's context caused a bad decision by rerunning it with pieces removed, checks candidate fixes the same way, and writes a pytest regression test. Runs locally with OpenAI, Anthropic, LangGraph and Ollama.GitHub Badge
traceAIOpen-source AI tracing framework built on OpenTelemetry for deep observability across agentic and LLM workflows.GitHub Badge
Future AGIProduction-grade SDK for observability, automated evaluations and prompt management with sub-100ms guardrails for LLM/agent workflows.GitHub Badge
semantic-coverageVisualizes RAG knowledge gaps and "blind spots" using 2D UMAP clustering and density detection.GitHub Badge
Token PoliceSDK (Python, Node) that prices each LLM call before it is made and blocks or reroutes it on per-user, per-session or per-workflow budget rules; prompt text never leaves the app.GitHub Badge
Weco ObserveObservability and debugging tool for AI research agents. Trace multi-step LLM agent runs, visualize decision trees, and identify failure modes in autonomous research workflows. Cloud hosted with open-source agent integration.
witnessA recording cache for model APIs: one Rust binary that proxies any LLM endpoint, dedupes repeat calls, and writes every model call, MCP tool call, and command side effect into a tamper-evident hash-chained journal with Ed25519 agent identity and Merkle proofs anchored to Sigstore Rekor.GitHub Badge
wobblyMetamorphic testing for LLM extraction — catch wrong outputs with no ground-truth labels: perturb inputs in ways that shouldn't change the answer, flag when they do.

⬆ back to ToC

LLMOps

ProjectDetailsRepository
agentaThe LLMOps platform to build robust LLM apps. Easily experiment and evaluate different prompts, models, and workflows to build robust apps.GitHub Badge
AgentMarkType-Safe Markdown-based AgentsGitHub Badge
AgentFieldOpen-source control plane for building and operating AI agents like APIs at scale, with routing, memory, observability, identity, auth, and policy controls.GitHub Badge
AgnosGateway-agnostic, self-hosted LLM control plane (MIT): an OpenAI-compatible governance proxy (auth, budgets, guardrails, cost, observability) that runs LiteLLM, Bifrost or Portkey as a swappable, stateless translation engine, so provider keys stay in your own control plane instead of the gateway.GitHub Badge
AI studioA Reliable Open Source AI studio to build core infrastructure stack for your LLM Applications. It allows you to gain visibility, make your application reliable, and prepare it for production with features such as caching, rate limiting, exponential retry, model fallback, and more.GitHub Badge
AISIXOpen-source AI gateway written in Rust: one OpenAI-compatible API (plus a native Anthropic Messages API) in front of LLM providers, and a gateway for MCP servers and A2A agents, with shared API keys, rate limits, guardrails, caching, and Prometheus/OTLP observability.GitHub Badge
AIWGDeploys project-owned agents, skills, rules, and governed workflows across AI coding platforms.GitHub Badge
Arize-PhoenixML observability for LLMs, vision, language, and tabular models.GitHub Badge
BitRouterAgent-native LLM router that optimizes your agent with every run — zero harness changes, with every model call reliable, traceable, secure, and cost-effective. Routes across OpenAI, Anthropic, Google, OpenRouter, Bedrock, GitHub Copilot, and more through one local endpoint, with cross-protocol translation, an MCP gateway, guardrails, observability, and multi-account failover. Written in Rust.GitHub Badge
boldrouterOpenAI-compatible AI gateway developed and hosted in Switzerland, with automatic routing, failover, and unified prepaid billing.
BudgetMLDeploy a ML inference service on a budget in less than 10 lines of code.GitHub Badge
CauraGoverned shared memory for AI agent fleets. Multi-agent, multi-tenant and MCP-native, with trust tiers, audit trails, a knowledge graph and self-improving retrieval.GitHub Badge
Cheshire Cat AIWeb framework to create vertical AI agents. FastAPI based, plugin system inspired to WordPress, admin panel, vector DB includedGitHub Badge
ChimeraforgeLLM deployment planner and benchmarking CLI: given a model and a GPU it answers will it fit, will it hit your SLO, and what will it cost. Searches model x quantization x backend x GPU count against VRAM, quality, latency, cost, energy and an opt-in safety gate; models tensor/pipeline parallelism, MoE active parameters, FP8, KV-quant and self-host-vs-API break-even. Every number is labeled measured/estimated/unknown. Also ships an MCP server.GitHub Badge
claude-routerLocal prompt router that picks the right Claude model tier and prepends the right scaffold using local embeddings, before the API call.GitHub Badge
CogniCoreStructured experience memory and learning infrastructure for AI agents, providing verified experience reuse, failure learning, environment-aware validity, and cross-session/cross-agent memory.GitHub Badge
ContextoSelf-hosted context engine for AI agents with persistent conversation memory and recall. Works as a drop-in OpenAI-compatible proxy, OpenClaw plugin, or memory SDK — no code changes required.GitHub Badge
CortexGenerates typed MCP servers, interactive API documentation, and SDKs from OpenAPI, AsyncAPI, GraphQL, gRPC, and OpenRPC sources.GitHub Badge
CrewDockSelf-hosted control plane for managing AI agents, with task approvals, activity logs, and cost tracking.GitHub Badge
DataoortsEnjoy unlimited API calls with Serverless AI Workers/LLMs for just $25 per month. No rate or concurrency limits.
deeplakeStream large multimodal datasets to achieve near 100% GPU utilization. Query, visualize, & version control data. Access data w/o the need to recompute the embeddings for the model finetuning.GitHub Badge
depdeskFinds deprecated and retired LLM model identifiers in a codebase and fails CI before a provider retirement reaches production. Parses Python with the stdlib ast module, weights findings by real call volume from a usage export, and ships a hand-transcribed Anthropic and OpenAI catalog carrying the date each provider page was verified. Zero dependencies.GitHub Badge
DifyOpen-source framework aims to enable developers (and even non-developers) to quickly build useful applications based on large language models, ensuring they are visual, operable, and improvable.GitHub Badge
DistilReversible, certified context compression for LLM agents: gates each compression through a statistical non-inferiority test so the agent's decisions don't change. Stdlib-only, zero-dependency proxy for Anthropic/OpenAI/Gemini.GitHub Badge
Doubleword Control LayerThe world's fastest open-source AI model gateway — ~450× less overhead than LiteLLM. Turns any model into a production-ready, OpenAI-compatible API with built-in auth, rate limits, and user controls.GitHub Badge
DstackCost-effective LLM development in any cloud (AWS, GCP, Azure, Lambda, etc).GitHub Badge
EjentumCognitive-harness MCP server with four tools (reasoning, code, anti-deception, memory) returning a structured scaffold (failure pattern, procedure, suppression vectors, falsification test) the agent absorbs before generating. Hosted at api.ejentum.com/mcp and on npm as ejentum-mcp.GitHub Badge
EmbedchainFramework to create ChatGPT like bots over your dataset.GitHub Badge
EpsillaAn all-in-one platform to create vertical AI agents powered by your private data and knowledge.
EvidentlyAn open-source framework to evaluate, test and monitor ML and LLM-powered systems.GitHub Badge
FerryAPILow-cost OpenAI-compatible AI API gateway for production workloads, with prepaid balance, usage billing, customer API keys, and provider account pools.
Fiddler AIEvaluate, monitor, analyze, and improve MLOps and LLMOps from pre-production to production.
FreeLLMAPIOpenAI-compatible proxy that stacks the free tiers of 28 LLM providers (~4B free tokens/month) behind one /v1 endpoint — smart routing, automatic failover, encrypted keys.GitHub Badge
GlideCloud-Native LLM Routing Engine. Improve LLM app resilience and speed.GitHub Badge
GoModelAI gateway exposing a unified OpenAI-compatible API across OpenAI, Anthropic, Gemini, Groq, xAI, Ollama and other providers, with routing, usage tracking, rate limits, and guardrails.GitHub Badge
gotoHumanBring a human into the loop in your LLM-based and agentic workflows. Prompt users to approve actions, select next steps, or review and validate generated results.
GPTCacheCreating semantic cache to store responses from LLM queries.GitHub Badge
GPUStackAn open-source GPU cluster manager for running and managing LLMsGitHub Badge
HaystackQuickly compose applications with LLM Agents, semantic search, question-answering and more.GitHub Badge
HiveOpen-source AI agent framework for building goal-driven, self-improving autonomous agents with auto-generated graphs, evolution loops, and MCP integration.GitHub Badge
HeliconeOpen-source LLM observability platform for logging, monitoring, and debugging AI applications. Simple 1-line integration to get started.GitHub Badge
HumanloopThe LLM evals platform for enterprises, providing tools to develop, evaluate, and observe AI systems.
HypersigilOpen-source prompt lifecycle management and gateway with a Web UI.GitHub Badge
IzloPrompt management tools for teams. Store, improve, test, and deploy your prompts in one unified workspace.
Keywords AIA unified DevOps platform for AI software. Keywords AI makes it easy for developers to build LLM applications.
KubeIntellectLLM-orchestrated multi-agent framework for Kubernetes operations. Investigates a live cluster via kubectl, Prometheus (PromQL) and Loki (LogQL), correlates the evidence into a root cause, and gates every mutating action behind human approval with role-based access control.GitHub Badge
KunavoHosted AI API gateway with OpenAI-compatible and native Anthropic endpoints, one prepaid balance, and per-request cost accounting.
MLflowAn open-source framework for the end-to-end machine learning lifecycle, helping developers track experiments, evaluate models/prompts, deploy models, and add observability with tracing.GitHub Badge
LaminarOpen-source all-in-one platform for engineering AI products. Traces, Evals, Datasets, Labels.GitHub Badge
langchainBuilding applications with LLMs through composabilityGitHub Badge
LangFlowAn effortless way to experiment and prototype LangChain flows with drag-and-drop components and a chat interface.GitHub Badge
LangfuseOpen Source LLM Engineering Platform: Traces, evals, prompt management and metrics to debug and improve your LLM application.GitHub Badge
LangKitOut-of-the-box LLM telemetry collection library that extracts features and profiles prompts, responses and metadata about how your LLM is performing over time to find problems at scale.GitHub Badge
LangWatchLLM Ops platform with Analytics, Monitoring, Evaluations and an LLM Optimization Studio powered by DSPyGitHub Badge
lintlangStatic analysis for AI agent configs, tool descriptions, and system prompts. Catches vague tool descriptions, missing stop conditions, and schema gaps before they reach runtime. Zero-LLM, deterministic checks, built for CI.GitHub Badge
LiteLLM 🚅A simple & light 100 line package to standardize LLM API calls across OpenAI, Azure, Cohere, Anthropic, Replicate API EndpointsGitHub Badge
Literal AIMulti-modal LLM observability and evaluation platform. Create prompt templates, deploy prompts versions, debug LLM runs, create datasets, run evaluations, monitor LLM metrics and collect human feedback.
LlamaIndexProvides a central interface to connect your LLMs with external data.GitHub Badge
LLMAppLLM App is a Python library that helps you build real-time LLM-enabled data pipelines with few lines of code.GitHub Badge
LLMFlowsLLMFlows is a framework for building simple, explicit, and transparent LLM applications such as chatbots, question-answering systems, and agents.GitHub Badge
LLMGraphNo-code visual builder for LLM/AI workflows: turn your docs and models into RAG chatbots and AI agents, then deploy each to a REST API and an embeddable chat widget in one click.
LRMCLI/TUI tool for managing localization files (.resx, JSON, Android, iOS) with LLM-powered translation via Ollama, validation, and code scanning for unused/missing keys.GitHub Badge
LunaryObservability and prompt management for LLM chabots and agents. Debug agents with powerful tracing and logging. Usage analytics and dive deep into the history of your requests. Developer friendly modules with plug-and-play integration into LangChain.GitHub Badge
MengramOpen-source memory infrastructure for AI agents. Provides semantic (entities/facts), episodic (conversations), and procedural (learned behaviors) memory with auto-reflection. Python SDK, JS SDK, MCP server, and REST API.GitHub Badge
magenticSeamlessly integrate LLMs as Python functions. Use type annotations to specify structured output. Mix LLM queries and function calling with regular Python code to create complex LLM-powered functionality.GitHub Badge
Manag.aiYour all-in-one prompt management and observability platform. Craft, track, and perfect your LLM prompts with ease.
MirascopeIntuitive convenience tooling for lightning-fast, efficient development and ensuring quality in LLM-based applicationsGitHub Badge
NativePortInfrastructure for accessing web-search, scraping, browser-automation, voice, and model-inference providers through a single account, API key, and balance. Publishes dated, capability-specific leaderboards and a public methodology.
NeurolinkMulti-provider AI agent framework that unifies 12+ LLM providers (OpenAI, Google, Anthropic, AWS, Azure, Groq, etc.) with workflow orchestration. Production-grade platform for building LLM applications with streaming, tool calling, caching, and enterprise features. Battle-tested at 15M+ requests/month.GitHub Badge
NoraSelf-hosted control plane for deploying and operating OpenClaw and Hermes agent fleets on Docker and Kubernetes, with lifecycle management, observability, cost tracking, secrets, schedules, and REST, CLI, and MCP interfaces.GitHub Badge
OpenLITOpenLIT is an OpenTelemetry-native GenAI and LLM Application Observability tool and provides OpenTelmetry Auto-instrumentation for monitoring LLMs, VectorDBs and Frameworks. It provides valuable insights into token & cost usage, user interaction, and performance related metrics.GitHub Badge
OpikConfidently evaluate, test, and ship LLM applications with a suite of observability tools to calibrate language model outputs across your dev and production lifecycle.GitHub Badge
OrcaRouter LiteSelf-hosted, OpenAI-compatible LLM router. Bring your own provider keys and route across 100+ models from multiple providers with automatic failover and streaming; model="auto" selects a model by cost, latency, or quality. Includes a local analytics dashboard and does not send telemetry.GitHub Badge
Parea AIPlatform and SDK for AI Engineers providing tools for LLM evaluation, observability, and a version-controlled enhanced prompt playground.GitHub Badge
Pezzo 🕹️Pezzo is the open-source LLMOps platform built for developers and teams. In just two lines of code, you can seamlessly troubleshoot your AI operations, collaborate and manage your prompts in one place, and instantly deploy changes to any environment.GitHub Badge
Pilot ProtocolOpen-source overlay network giving AI agents a permanent virtual address, encrypted UDP tunnels with NAT traversal, and an explicit per-peer trust model, plus an app store of installable agent-native capabilities (discover → install → call).GitHub Badge
PraisonAIProduction-ready Multi-AI Agents framework with self-reflection. Fastest agent instantiation (3.77μs), 100+ LLM support via LiteLLM, MCP integration, agentic workflows (route/parallel/loop/repeat), built-in memory, Python & JS SDKs.GitHub Badge
PromptDXA declarative, extensible, and composable approach for developing LLM prompts using Markdown and JSX.GitHub Badge
PromptHubFull stack prompt management tool designed to be usable by technical and non-technical team members. Test, version, collaborate, deploy, and monitor, all from one place.
promptfooOpen-source tool for testing & evaluating prompt quality. Create test cases, automatically check output quality and catch regressions, and reduce evaluation cost.GitHub Badge
PromptFoundryThe simple prompt engineering and evaluation tool designed for developers building AI applications.GitHub Badge
PromptLayer 🍰Prompt Engineering platform. Collaborate, test, evaluate, and monitor your LLM applicationsGithub Badge
PromptMageOpen-source tool to simplify the process of creating and managing LLM workflows and prompts as a self-hosted solution.GitHub Badge
PromptSiteA lightweight Python library for prompt lifecycle management that helps you version control, track, experiment and debug with your LLM prompts with ease. Minimal setup, no servers, databases, or API keys required - works directly with your local filesystem, ideal for data scientists and engineers to easily integrate into existing LLM workflows
PrompteamsPrompt management system. Version, test, collaborate, and retrieve prompts through real-time APIs. Have GitHub style with repos, branches, and commits (and commit history).
prompttoolsOpen-source tools for testing and experimenting with prompts. The core idea is to enable developers to evaluate prompts using familiar interfaces like code and notebooks. In just a few lines of codes, you can test your prompts and parameters across different models (whether you are using OpenAI, Anthropic, or LLaMA models). You can even evaluate the retrieval accuracy of vector databases.GitHub Badge
Puzzlet AIThe Git-Based LLM Engineering Platform. Achieve more from GenAI: Manage, evaluate, and improve your full-stack LLM application - with version control, type-safety, and local development built-in.
QuotaflowAI token and API resource utilization platform that helps teams reduce wasted subscribed quota and improve turnover across controlled internal pools.
systemprompt.ioSystemprompt.io is a Rest API with quality tooling to enable the creation, use and observability of prompts in any AI system. Control every detail of your prompt for a SOTA prompt management experience.
TeamoRouterLLM routing gateway for OpenClaw. One API key to access Claude, GPT-4o, Gemini, DeepSeek, Kimi, MiniMax. Smart routing modes (teamo-best, teamo-balanced, teamo-eco) auto-pick the optimal model. Up to 50% off official prices. 2-second install via skill.md.
TokenMixAI gateway routing 171 LLMs from 14 providers (Claude, GPT, Gemini, DeepSeek, Qwen, and more) through one OpenAI-compatible endpoint. Pass-through pricing, automatic failover, no monthly subscription.
TreeScaleAll In One Dev Platform For LLM Apps. Deploy LLM-enhanced APIs seamlessly using tools for prompt optimization, semantic querying, version management, statistical evaluation, and performance tracking. As a part of the developer friendly API implementation TreeScale offers Elastic LLM product, which makes a unified API Endpoint for all major LLM providers and open source models.
TrueFoundryDeploy LLMOps tools like Vector DBs, Embedding server etc on your own Kubernetes (EKS,AKS,GKE,On-prem) Infra including deploying, Fine-tuning, tracking Prompts and serving Open Source LLM Models with full Data Security and Optimal GPU Management. Train and Launch your LLM Application at Production scale with best Software Engineering practices.
ReliableGPT 💪Handle OpenAI Errors (overloaded OpenAI servers, rotated keys, or context window errors) for your production LLM Applications.GitHub Badge
Registry BrokerUniversal index and routing layer for AI agents. Aggregates agent metadata from multiple registries (NANDA, MCP, Virtuals, OpenRouter, A2A, X402 Bazaar) across web2 and web3, normalizes profiles, and provides protocol translation between agent ecosystems.GitHub Badge
RhesisOpen-source testing infrastructure for LLM and agentic applications. Collaborative platform enabling teams to define quality metrics, run evaluations, and ship confidently with version control and peer review workflows built for AI engineering.GitHub Badge
roteOpen-source CLI that compiles a proven agent skill (a SKILL.md plus references) into a typed, deterministic pipeline. Fixed logic becomes reviewable Python or TypeScript with per-step tests, while judgment steps stay as typed LLM-judge signatures. Emits DBOS, Temporal, Cloudflare Workflows, Inngest, or plain Python/TS, and can serve compiled pipelines as MCP tools.GitHub Badge
RoundtableZero-configuration unified AI assistant management built on the FastMCP framework. Provides seamless integration with Claude, ChatGPT, and other AI assistants through a single MCP interface with session management, logging, and production-ready operations.GitHub Badge
PortkeyControl Panel with an observability suite & an AI gateway — to ship fast, reliable, and cost-efficient apps.
SAVI SDKOpen-source observability SDK for LLM cost, PII masking, carbon, and compliance — drop-in wrapper for OpenAI/Anthropic/Bedrock/Cohere/Mistral/Vertex; works standalone with zero account via local_mode.GitHub Badge
Self-Hosted AI StackDocker Compose stack for local AI with Ollama, a LiteLLM gateway, RAG, voice services, MCP tools, persistent data, and health checks.GitHub Badge
Semantic Cache RouterDistributed semantic cache and stateful routing system that cuts LLM API costs by returning cached responses for semantically similar queries. Uses ANN vector search (cosine ≥ 0.8) and consistent hashing to pin requests to the same worker, achieving ~7× latency reduction on cache hits while scaling horizontally without cache thrash.GitHub Badge
SpendlineFinancial control layer for AI spend — per-customer cost attribution and hierarchical budgets enforced before the provider call.
StatewaveOpen-source memory runtime for AI agents. Compiles events into deterministic, provenance-tagged context bundles instead of query-time retrieval. Apache-2.0, self-hostable on Postgres + pgvector.GitHub Badge
TensorZeroTensorZero is an open-source framework for building production-grade LLM applications. It unifies an LLM gateway, observability, optimization, evaluations, and experimentation.GitHub Badge
ThinkWatch LiteDesktop app for macOS, Windows and Linux that runs a local LLM gateway for Claude Code, Codex and other AI coding clients: routing rules with failover, conversion between the Anthropic, OpenAI and Gemini APIs, a record of each request's route and cost, and redaction of API keys before requests leave. The gateway engine, ThinkWatch Core, also runs on its own on a Linux server.GitHub Badge
TrinitySelf-hosted platform that runs AI coding agents as persistent scheduled services. Each agent runs in a Docker container with its own workspace and git-backed memory, plus scheduling, an approvals queue, an MCP server and a web console.GitHub Badge
UnoRouterOpenAI-compatible LLM gateway with one API key for every major provider and smart routing across models. Drop-in for code, Claude Code, and chat clients like SillyTavern, Janitor.AI, RisuAI, and Chub.
VellumAn AI product development platform to experiment with, evaluate, and deploy advanced LLM apps.
Weights & Biases (Prompts)A suite of LLMOps tools within the developer-first W&B MLOps platform. Utilize W&B Prompts for visualizing and inspecting LLM execution flow, tracking inputs and outputs, viewing intermediate results, securely managing prompts and LLM chain configurations.
WenlanLocal-first AI knowledge base and LLM wiki that distills agent work into source-cited pages with graph context and hybrid retrieval across MCP clients.GitHub Badge
WordwareA web-hosted IDE where non-technical domain experts work with AI Engineers to build task-specific AI agents. It approaches prompting as a new programming language rather than low/no-code blocks.
XiuRouterHosted multi-model API service with OpenAI Chat Completions and Responses, Anthropic Messages, and Gemini GenerateContent routes, scoped API keys, usage-based pricing, and request-level usage and cost records.
xTuringBuild and control your personal LLMs with fast and efficient fine-tuning.GitHub Badge
ZenMLOpen-source framework for orchestrating, experimenting and deploying production-grade ML solutions, with built-in langchain & llama_index integrations.GitHub Badge
SwarmClawSelf-hosted multi-agent AI runtime with 23+ LLM providers, persistent memory, skills, schedules, sub-agent spawning, and MCP client + server support. Ships as desktop app, CLI, or Docker.GitHub Badge
ai-evaluationEvaluation framework for automated, reproducible scoring of LLM, agent, and workflow performance.GitHub Badge
future-agiOpen-source self-hostable end-to-end agent engineering and optimization platform unifying tracing, evals, simulations, datasets, gateway, and guardrails for LLM and AI agent applications.GitHub Badge
ModelglassSourced, versioned pricing and capability data for AI models (image, language, video, audio, plus coding/science/agentic benchmark verticals) to find the cheapest model that clears a capability bar.

⬆ back to ToC

ProjectDetailsRepository
AirweaveAn easy way to turn any app into searchable data for LLMs.GitHub Badge
MemorySyncPersistent multi-tenant memory layer and MCP server for AI coding assistants with sub-50ms hybrid recall.GitHub Badge
ProjectDetailsRepository
AquilaDBAn easy to use Neural Search Engine. Index latent vectors along with JSON metadata and do efficient k-NN search.GitHub Badge
AwadbAI Native database for embedding vectorsGitHub Badge
Chromathe open source embedding databaseGitHub Badge
EpsillaA 10x faster, cheaper, and better vector databaseGitHub Badge
InfinityThe AI-native database built for LLM applications, providing incredibly fast vector and full-text searchGitHub Badge
InfinoEmbedded retrieval engine on Apache Parquet: BM25 full-text, vector, hybrid (RRF), and SQL from one engine over object storage.GitHub Badge
LancedbDeveloper-friendly, serverless vector database for AI applications. Easily add long-term memory to your LLM apps!GitHub Badge
MarqoTensor search for humans.GitHub Badge
MilvusVector database for scalable similarity search and AI applications.GitHub Badge
OmnigraphTyped graph database where agents branch and merge like Git. S3-native, Rust, traversal + vector + BM25 in one runtime.GitHub Badge
ParadeDBThe transactional alternative to Elasticsearch, built on Postgres.GitHub Badge
PineconeThe Pinecone vector database makes it easy to build high-performance vector search applications. Developer-friendly, fully managed, and easily scalable without infrastructure hassles.
pgvectorOpen-source vector similarity search for Postgres.GitHub Badge
RivestackManaged PostgreSQL with pgvector for AI workloads. Built-in SQL editor lets you query your database with natural language (auto-converted to vector embeddings). Free tier includes 2GB storage.
SynapCoresSelf-hosted AI-native database: vector search + Cypher graph + embedded GGUF inference and SQL in one engine. Free Community Edition binary; source proprietary.GitHub Badge
VectorChordScalable, fast, and disk-friendly vector search in Postgres, the successor of pgvecto.rs.GitHub Badge
pgvecto.rsVector database plugin for Postgres, written in Rust, specifically designed for LLM.GitHub Badge
QdrantVector Search Engine and Database for the next generation of AI applications. Also available in the cloudGitHub Badge
txtaiBuild AI-powered semantic search applicationsGitHub Badge
ValdA Highly Scalable Distributed Vector Search EngineGitHub Badge
VearchA distributed system for embedding-based vector retrievalGitHub Badge
VectorDBA Python vector database you just need - no more, no less.GitHub Badge
VellumA managed service for ingesting documents and performing hybrid semantic/keyword search across them. Comes with out-of-box support for OCR, text chunking, embedding model experimentation, metadata filtering, and production-grade APIs.
WeaviateWeaviate is an open source vector search engine that stores both objects and vectors, allowing for combining vector search with structured filtering with the fault-tolerance and scalability of a cloud-native database, all accessible through GraphQL, REST, and various language clients.GitHub Badge

⬆ back to ToC

Code AI

ProjectDetailsRepository
AgentsMeshSelf-hostable AI Agent Workforce Platform. Multi-agent orchestration with remote AI workstations (AgentPods), PTY sandbox + git worktree isolation, built-in Kanban, and per-pod MCP server. Supports Claude Code, Codex CLI, Gemini CLI, Aider, OpenCode.GitHub Badge
Atomic AgentLocal-first CLI and TUI coding agent that runs open-weight models entirely on your machine through a llama.cpp fork. No account or API key required. 56 built-in tools (browser, filesystem, git, memory, vision), MCP support, five-layer local memory, macOS/Linux/Windows.GitHub Badge
BernsteinDeterministic Python orchestrator for 37 CLI coding agents (Claude Code, Codex CLI, Gemini CLI, GitHub Copilot CLI, Cursor, Aider, OpenHands, OpenCode, Goose, Qwen, Ollama, ...) running in parallel git worktrees. First-class MCP server, quality gates, cost tracking with budgets.GitHub Badge
CodeGeeXCodeGeeX: An Open Multilingual Code Generation Model (KDD 2023)GitHub Badge
CodeGenCodeGen is an open-source model for program synthesis. Trained on TPU-v4. Competitive with OpenAI Codex.GitHub Badge
Coder EvalFramework for evaluating, benchmarking, and A/B-testing AI coding agents (Claude Code, Codex CLI, Gemini/Antigravity) and their skills. Runs a real agent in a sandbox against declarative YAML tasks, then scores the resulting files and commands with weighted 0.0-1.0 criteria. Includes 14 criterion types, per-tool token/cost telemetry, dataset fan-out with skill-activation precision/recall gates, and a GitHub Action with JUnit XML output for CI.GitHub Badge
CodeT5Open Code LLMs for Code Understanding and Generation.GitHub Badge
Continue⏩ the open-source autopilot for software development—bring the power of ChatGPT to VS CodeGitHub Badge
CotalOpen pub/sub standard over NATS JetStream for coordinating coding agents. Claude Code, Codex, OpenCode, Hermes, Jcode and pi agents share a space with presence, channels, durable direct messages and role-addressed delivery.GitHub Badge
DSH StudioCross-platform desktop host for installing, configuring, and managing DeepSeek Harness.GitHub Badge
fauxpilotAn open-source alternative to GitHub Copilot serverGitHub Badge
fractalHierarchical coding-agent orchestrator with recursive delegation, per-node Git worktrees, configurable limits, persistent SQLite state, and live terminal monitoring and steering.GitHub Badge
Kolega CodePython terminal coding agent where the model writes its own multi-agent workflows (Gigacode). Provider-agnostic, local-first, 15+ model providers, MCP client, browser agent.GitHub Badge
promptextSmart code context extractor for AI assistants with accurate token counting and budget managementGitHub Badge
RelayUniversal AI API proxy — hot-switch between 14+ LLM providers (GitHub Copilot, OpenAI, Anthropic, DeepSeek, Groq, Ollama, and more) from any AI coding agent without restarting sessions. Local proxy with automatic Anthropic <-> OpenAI protocol translation, account rotation, and real-time usage tracking.GitHub Badge
SuperagentOpen-source macOS desktop app that gives Claude Code and Codex a real browser to drive, an iOS Simulator to install and screenshot apps in, and a phone companion app for remote monitoring.GitHub Badge
tabbySelf-hosted AI coding assistant. An opensource / on-prem alternative to GitHub Copilot.GitHub Badge
AIDEOpen-source ML engineering agent that uses tree search to explore solution spaces. Automates machine learning experimentation from data analysis to model training. Paper.GitHub Badge
KapsoLong-running agents that optimize AI and Data systems, and learn from every experience. #1 open-source on MLE-Bench; ALE-Bench; RelBench.GitHub Badge
webcmdSelf-learning browser infrastructure for AI coding agents that compiles site navigation into deterministic per-site CLI commands.GitHub Badge

Training

IDEs and Workspaces

ProjectDetailsRepository
code serverRun VS Code on any machine anywhere and access it in the browser.GitHub Badge
condaOS-agnostic, system-level binary package manager and ecosystem.GitHub Badge
DockerMoby is an open-source project created by Docker to enable and accelerate software containerization.GitHub Badge
envd🏕️ Reproducible development environment for AI/ML.GitHub Badge
Jupyter NotebooksThe Jupyter notebook is a web-based notebook environment for interactive computing.GitHub Badge
KurtosisA build, packaging, and run system for ephemeral multi-container environments.GitHub Badge
LayerSmithSelf-hosted web UI, TUI, and CLI for building OCI images with Docker or Podman, including LLM training environments and air-gap bundles.GitHub Badge
WordwareA web-hosted IDE where non-technical domain experts work with AI Engineers to build task-specific AI agents. It approaches prompting as a new programming language rather than low/no-code blocks.

⬆ back to ToC

Foundation Model Fine Tuning

ProjectDetailsRepository
alpaca-loraInstruct-tune LLaMA on consumer hardwareGitHub Badge
finetuning-schedulerA PyTorch Lightning extension that accelerates and enhances foundation model experimentation with flexible fine-tuning schedules.GitHub Badge
FlyflowOpen source, high performance fine tuning as a service for GPT4 quality models with 5x lower latency and 3x lower costGitHub Badge
LMFlowAn Extensible Toolkit for Finetuning and Inference of Large Foundation ModelsGitHub Badge
LoraUsing Low-rank adaptation to quickly fine-tune diffusion models.GitHub Badge
peftState-of-the-art Parameter-Efficient Fine-Tuning.GitHub Badge
p-tuning-v2An optimized prompt tuning strategy achieving comparable performance to fine-tuning on small/medium-sized models and sequence tagging challenges. (ACL 2022)GitHub Badge
QLoRAEfficient finetuning approach that reduces memory usage enough to finetune a 65B parameter model on a single 48GB GPU while preserving full 16-bit finetuning task performance.GitHub Badge
TrainJudgeDiagnoses whether fine-tuning fits, then verifies fine-tunes on held-out task metrics and a regression suite, not training loss.GitHub Badge
TRLTrain transformer language models with reinforcement learning.GitHub Badge

⬆ back to ToC

Frameworks for Training

ProjectDetailsRepository
Accelerate🚀 A simple way to train and use PyTorch models with multi-GPU, TPU, mixed-precision.GitHub Badge
Apache MXNetLightweight, Portable, Flexible Distributed/Mobile Deep Learning with Dynamic, Mutation-aware Dataflow Dep Scheduler.GitHub Badge
axolotlA tool designed to streamline the fine-tuning of various AI models, offering support for multiple configurations and architectures.GitHub Badge
CaffeA fast open framework for deep learning.GitHub Badge
CandleMinimalist ML framework for Rust .GitHub Badge
ColossalAIAn integrated large-scale model training system with efficient parallelization techniques.GitHub Badge
DeepSpeedDeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.GitHub Badge
HorovodDistributed training framework for TensorFlow, Keras, PyTorch, and Apache MXNet.GitHub Badge
JaxAutograd and XLA for high-performance machine learning research.GitHub Badge
KedroKedro is an open-source Python framework for creating reproducible, maintainable and modular data science code.GitHub Badge
KerasKeras is a deep learning API written in Python, running on top of the machine learning platform TensorFlow.GitHub Badge
LightGBMA fast, distributed, high performance gradient boosting (GBT, GBDT, GBRT, GBM or MART) framework based on decision tree algorithms, used for ranking, classification and many other machine learning tasks.GitHub Badge
MegEngineMegEngine is a fast, scalable and easy-to-use deep learning framework, with auto-differentiation.GitHub Badge
metric-learnMetric Learning Algorithms in Python.GitHub Badge
MindSporeMindSpore is a new open source deep learning training/inference framework that could be used for mobile, edge and cloud scenarios.GitHub Badge
OneflowOneFlow is a performance-centered and open-source deep learning framework.GitHub Badge
PaddlePaddleMachine Learning Framework from Industrial Practice.GitHub Badge
PyTorchTensors and Dynamic neural networks in Python with strong GPU acceleration.GitHub Badge
PyTorch LightningDeep learning framework to train, deploy, and ship AI products Lightning fast.GitHub Badge
XGBoostScalable, Portable and Distributed Gradient Boosting (GBDT, GBRT or GBM) Library.GitHub Badge
scikit-learnMachine Learning in Python.GitHub Badge
TensorFlowAn Open Source Machine Learning Framework for Everyone.GitHub Badge
VectorFlowA minimalist neural network library optimized for sparse data and single machine environments.GitHub Badge

⬆ back to ToC

Experiment Tracking

ProjectDetailsRepository
Aiman easy-to-use and performant open-source experiment tracker.GitHub Badge
ClearMLAuto-Magical CI/CD to streamline your ML workflow. Experiment Manager, MLOps and Data-ManagementGitHub Badge
CometComet is an MLOps platform that offers experiment tracking, model production management, a model registry, and full data lineage from training straight through to production. Comet plays nicely with all your favorite tools, so you don't have to change your existing workflow. Comet Opik to confidently evaluate, test, and ship LLM applications with a suite of observability tools to calibrate language model outputs across your dev and production lifecycle!GitHub Badge
Guild AIExperiment tracking, ML developer tools.GitHub Badge
MLRunMachine Learning automation and tracking.GitHub Badge
Kedro-VizKedro-Viz is an interactive development tool for building data science pipelines with Kedro. Kedro-Viz also allows users to view and compare different runs in the Kedro project.GitHub Badge
LabNotebookLabNotebook is a tool that allows you to flexibly monitor, record, save, and query all your machine learning experiments.GitHub Badge
SacredSacred is a tool to help you configure, organize, log and reproduce experiments.GitHub Badge
Weights & BiasesA developer first, lightweight, user-friendly experiment tracking and visualization tool for machine learning projects, streamlining collaboration and simplifying MLOps. W&B excels at tracking LLM-powered applications, featuring W&B Prompts for LLM execution flow visualization, input and output monitoring, and secure management of prompts and LLM chain configurations.GitHub Badge

⬆ back to ToC

Visualization

ProjectDetailsRepository
Fiddler AIRich dashboards, reports, and UMAP to perform root cause analysis, pinpoint problem areas, like correctness, safety, and privacy issues, and improve LLM outcomes.
LangWatchVisualize LLM evaluations experiments and DSPy pipeline optimizationsGitHub Badge
ManifordA model-agnostic visual debugging tool for machine learning.GitHub Badge
netronVisualizer for neural network, deep learning, and machine learning models.GitHub Badge
OpenOpsBring multiple data streams into one dashboard.GitHub Badge
TensorBoardTensorFlow's Visualization Toolkit.GitHub Badge
TensorSpaceNeural network 3D visualization framework, build interactive and intuitive model in browsers, support pre-trained deep learning models from TensorFlow, Keras, TensorFlow.js.GitHub Badge
dtreevizA python library for decision tree visualization and model interpretation.GitHub Badge
Zetane ViewerML models and internal tensors 3D visualizer.GitHub Badge
ZenoAI evaluation platform for interactively exploring data and model outputs.GitHub Badge

Model Editing

ProjectDetailsRepository
FastEditFastEdit aims to assist developers with injecting fresh and customized knowledge into large language models efficiently using one single command.GitHub Badge

⬆ back to ToC

Data

Data Management

ProjectDetailsRepository
ArtiVCA version control system to manage large files. Lake is a dataset format with a simple API for creating, storing, and collaborating on AI datasets of any size.GitHub Badge
DoltGit for Data.GitHub Badge
DVCData Version Control - Git for Data & Models - ML Experiments Management.GitHub Badge
Delta-LakeStorage layer that brings scalable, ACID transactions to Apache Spark and other engines.GitHub Badge
PachydermPachyderm is a version control system for data.GitHub Badge
QuiltA self-organizing data hub for S3.GitHub Badge

⬆ back to ToC

Data Storage

ProjectDetailsRepository
JuiceFSA distributed POSIX file system built on top of Redis and S3.GitHub Badge
LakeFSGit-like capabilities for your object storage.GitHub Badge
LanceModern columnar data format for ML implemented in Rust.GitHub Badge
PixeltableDeclarative multimodal AI data engine for versioned tables, computed columns, and vector search.GitHub Badge

⬆ back to ToC

Data Tracking

ProjectDetailsRepository
PiperiderA CLI tool that allows you to build data profiles and write assertion tests for easily evaluating and tracking your data's reliability over time.GitHub Badge
LUXA Python library that facilitates fast and easy data exploration by automating the visualization and data analysis process.GitHub Badge
ragfreshA CLI that detects stale, drifted, and ghost documents in RAG vector indexes by diffing content hashes and embedding drift against the source of truth.GitHub Badge

⬆ back to ToC

Feature Engineering

ProjectDetailsRepository
FeatureformThe Virtual Feature Store. Turn your existing data infrastructure into a feature store.GitHub Badge
FeatureToolsAn open source python framework for automated feature engineeringGitHub Badge

⬆ back to ToC

Data/Feature enrichment

ProjectDetailsRepository
CocoIndexAn ETL framework for AI that transforms data into embeddings and knowledge graphs, with incremental processing to recompute only what changed and keep indexes fresh.GitHub Badge
UpginiFree automated data & feature enrichment library for machine learning: automatically searches through thousands of ready-to-use features from public and community shared data sources and enriches your training dataset with only the accuracy improving featuresGitHub Badge
FeastAn open source feature store for machine learning.GitHub Badge
distilabel⚗️ distilabel is a framework for synthetic data and AI feedback for AI engineers that require high-quality outputs, full data ownership, and overall efficiency.GitHub Badge
FastDatasetsA powerful tool for creating high-quality training datasets for Large Language Models.GitHub Badge
fastdocparseExtracts structured data from documents with an LLM, then grounds every field against the source text to flag hallucinated values automatically.GitHub Badge

⬆ back to ToC

Large Scale Deployment

ML Platforms

ProjectDetailsRepository
CometComet is an MLOps platform that offers experiment tracking, model production management, a model registry, and full data lineage from training straight through to production. Comet plays nicely with all your favorite tools, so you don't have to change your existing workflow. Comet Opik to confidently evaluate, test, and ship LLM applications with a suite of observability tools to calibrate language model outputs across your dev and production lifecycle!GitHub Badge
ClearMLAuto-Magical CI/CD to streamline your ML workflow. Experiment Manager, MLOps and Data-Management.GitHub Badge
dstackOpen-source confidential AI framework for secure LLM deployment with data privacy, providing hardware-enforced isolation for production ML workloads.GitHub Badge
HopsworksHopsworks is a MLOps platform for training and operating large and small ML systems, including fine-tuning and serving LLMs. Hopsworks includes both a feature store and vector database for RAG.GitHub Badge
OpenLLMAn open platform for operating large language models (LLMs) in production. Fine-tune, serve, deploy, and monitor any LLMs with ease.GitHub Badge
MLflowOpen source platform for the machine learning lifecycle.GitHub Badge
MLRunAn open MLOps platform for quickly building and managing continuous ML applications across their lifecycle.GitHub Badge
ModelFoxModelFox is a platform for managing and deploying machine learning models.GitHub Badge
KserveStandardized Serverless ML Inference Platform on KubernetesGitHub Badge
KubeStellar ConsoleOpen source AI-powered multi-cluster Kubernetes dashboard for managing LLM workloads across hybrid edge and cloud environments. GPU monitoring, benchmark streaming, real-time observability with 20+ CNCF integrations, and AI-guided cluster operations. CNCF Sandbox project.GitHub Badge
KubeflowMachine Learning Toolkit for Kubernetes.GitHub Badge
PAIResource scheduling and cluster management for AI.GitHub Badge
piqcOpen-source, read-only GPU waste scanner for Kubernetes inference clusters. Deploys in minutes, no write permissions required.GitHub Badge
PolyaxonMachine Learning Management & Orchestration Platform.GitHub Badge
PrimehubAn effortless infrastructure for machine learning built on the top of Kubernetes.GitHub Badge
OpenModelZOne-click machine learning deployment (LLM, text-to-image and so on) at scale on any cluster (GCP, AWS, Lambda labs, your home lab, or even a single machine).GitHub Badge
Seldon-coreAn MLOps framework to package, deploy, monitor and manage thousands of production machine learning modelsGitHub Badge
StarwhaleAn MLOps/LLMOps platform for model building, evaluation, and fine-tuning.GitHub Badge
TrueFoundryA PaaS to deploy, Fine-tune and serve LLM Models on a company’s own Infrastructure with Data Security and Optimal GPU and Cost Management. Launch your LLM Application at Production scale with best DevSecOps practices.
Weights & BiasesA lightweight and flexible platform for machine learning experiment tracking, dataset versioning, and model management, enhancing collaboration and streamlining MLOps workflows. W&B excels at tracking LLM-powered applications, featuring W&B Prompts for LLM execution flow visualization, input and output monitoring, and secure management of prompts and LLM chain configurations.GitHub Badge

⬆ back to ToC

Workflow

ProjectDetailsRepository
AirflowA platform to programmatically author, schedule and monitor workflows.GitHub Badge
aqueductAn Open-Source Platform for Production Data ScienceGitHub Badge
Argo WorkflowsWorkflow engine for Kubernetes.GitHub Badge
awaithumansOne-function HITL primitive for AI agents. await_human blocks like a Promise; a human reviews via Slack/email/dashboard; the agent resumes with the typed response. Durable across restarts via Stripe-style idempotency keys. Temporal + LangGraph adapters included.GitHub Badge
FlyteKubernetes-native workflow automation platform for complex, mission-critical data and ML processes at scale.GitHub Badge
HamiltonA lightweight framework to represent ML/language model pipelines as a series of python functions.GitHub Badge
HeymSource-available, self-hosted AI workflow automation platform with a visual canvas for agent, RAG, and tool-using workflows. Includes MCP support, evals, traces, and cost tracking.GitHub Badge
KitaruDurable execution layer for AI agents. Checkpoints, replay, resume, and observability primitives that make agent workflows persistent and replayable — no graph DSL required.GitHub Badge
Kubeflow PipelinesMachine Learning Pipelines for Kubeflow.GitHub Badge
LangFlowAn effortless way to experiment and prototype LangChain flows with drag-and-drop components and a chat interface.GitHub Badge
MetaflowBuild and manage real-life data science projects with ease!GitHub Badge
PloomberThe fastest way to build data pipelines. Develop iteratively, deploy anywhere.GitHub Badge
PrefectThe easiest way to automate your data.GitHub Badge
VDPAn open-source unstructured data ETL tool to streamline the end-to-end unstructured data processing pipeline.GitHub Badge
ZenMLMLOps framework to create reproducible pipelines.GitHub Badge
simulate-sdkEnterprise-grade Voice AI simulation SDK for scenario-driven stress testing of multimodal and agentic systems.GitHub Badge

⬆ back to ToC

Scheduling

ProjectDetailsRepository
KueueKubernetes-native Job Queueing.GitHub Badge
PAIResource scheduling and cluster management for AI (Open-sourced by Microsoft).GitHub Badge
SlurmA Highly Scalable Workload Manager.GitHub Badge
VolcanoA Cloud Native Batch System (Project under CNCF).GitHub Badge
YunikornLight-weight, universal resource scheduler for container orchestrator systems.GitHub Badge

⬆ back to ToC

Model Management

ProjectDetailsRepository
CometComet is an MLOps platform that offers Model Production Management, a Model Registry, and full model lineage from training straight through to production. Use Comet for model reproducibility, model debugging, model versioning, model visibility, model auditing, model governance, and model monitoring.GitHub Badge
dvcML Experiments Management - Data Version Control - Git for Data & Models![GitHub Badge](https://img.shields.io/github/stars/iterative/dvc.svg?style=flat-square

Truncated — view the full README on GitHub.

ai-development-tools
awesome-list
llmops
mlops

Significant stargazers

(top 24 of 38)

Marc Klingen

462 followers · starred Nov 2024

Guangdong Liu

102 followers · starred Aug 2024

Utku Demir

159 followers · starred Apr 2023

Yuki Iwai

188 followers · starred May 2022