10 repos
Tools and platforms for managing, monitoring, and evaluating large language model applications in production. This cluster spans observability systems (MLflow, Langfuse, Helicone), evaluation frameworks (Opik, muteval), and prompt engineering infrastructure for LLMOps workflows. Repositories here support practitioners in tracking model performance, debugging LLM behavior, running automated evaluations, and optimizing prompts across development and production environments.