LLM Application Frameworks & Evaluation

11 repos

Libraries and frameworks for building production applications with large language models, including retrieval-augmented generation (RAG), agent orchestration, and prompt caching. The cluster emphasizes practical tooling for LLM integration—caching responses for cost and latency optimization, structuring multi-step agent workflows, and evaluating model outputs—with most repos written in Python and centered around OpenAI and LangChain ecosystems. Central repos like GPTCache and LLPhant represent specialized layers (caching and agent frameworks respectively), while VLMEvalKit brings systematic evaluation methodology to the broader LLM application stack.

Python · 8
PHP · 1
Rust · 1
openai ·108,802
llm ·101,642
langchain ·82,870
agent ·82,733
rag ·81,025
prompt-engineering ·77,393
anthropic ·73,552
mcp ·73,338
token-optimization ·72,804
tokens ·72,804