LLM Operations & Evaluation

10 repos

Tools and platforms for managing, monitoring, and evaluating large language model applications in production. This cluster spans observability systems (MLflow, Langfuse, Helicone), evaluation frameworks (Opik, muteval), and prompt engineering infrastructure for LLMOps workflows. Repositories here support practitioners in tracking model performance, debugging LLM behavior, running automated evaluations, and optimizing prompts across development and production environments.

Python · 4
TypeScript · 4
Clojure · 1
Jupyter Notebook · 1
prompt-engineering ·54,458
llmops ·54,454
llm ·51,044
evaluation ·50,397
llm-evaluation ·47,364
prompts ·28,285
rag ·25,883
prompt-testing ·25,247
testing ·25,247
evaluation-framework ·25,243