Sriti Core is an open-source, ultra-efficient intelligent routing proxy and cascading engine for Large Language Models.
Python
2
6 commits
updated Sep 17, 2026
Intelligent, Reliability-Aware Model Routing & Cascade Engine for LLMs
Key Features • Architecture • Quickstart • How It Works • Configuration • API Reference • Testing • Troubleshooting • License
Sriti Core is an open-source, ultra-efficient intelligent routing proxy and cascading engine for Large Language Models.
Instead of routing every single request to expensive frontier models (like Claude 3.5 Sonnet or GPT-4o), Sriti dynamically categorizes the request, checks an encrypted semantic cache, and cascades through a hierarchy of models:
Sriti continuously learns model reliability, evaluates output quality gates, and guarantees strict latency SLOs while dramatically slashing cloud API bills.
fastembed BAAI/bge-small-en-v1.5) mapping prompts across categories (e.g., Code, Extraction, Summarization, QA, Math, Reasoning). User / Application Request
│
▼
┌──────────────────────────────────┐
│ 1. Fast ONNX Classifier │ ──> Task Category & Embeddings
└──────────────────────────────────┘
│
▼
┌──────────────────────────────────┐
│ 2. Semantic Cache (Valkey) │ ──[Hit]──> Return Cached Response
└──────────────────────────────────┘
│ [Miss]
▼
┌──────────────────────────────────┐
│ 3. Model Cascading Engine │
│ │
│ ┌────────────────────────────┐ │
│ │ Tier 3: Local Edge Models │ │ ──[Success]──> Accept & Learn
│ └────────────────────────────┘ │
│ │ [Fail/Timeout] │
│ ┌────────────────────────────┐ │
│ │ Tier 2: Cost-Effective LLM │ │ ──[Success]──> Accept & Learn
│ └────────────────────────────┘ │
│ │ [Fail/Timeout] │
│ ┌────────────────────────────┐ │
│ │ Tier 1: Frontier Models │ │ ──[Guaranteed Execution]
│ └────────────────────────────┘ │
└──────────────────────────────────┘
│
▼
┌──────────────────────────────────┐
│ 4. Update Reliability & Cache │
└──────────────────────────────────┘
Python 3.11+ is required.
Via pip (recommended):
pip install sriti-core
From source:
git clone https://github.com/sriti-ai/sriti-core.git
cd sriti-core
pip install -e ".[dev]"
Create a .env file with your credentials and configuration:
SRITI_REDIS_URL=redis://localhost:6379/0
# Optional: Provider API keys (only configure what you use)
ANTHROPIC_API_KEY=sk-ant-...
OPENAI_API_KEY=sk-...
GROQ_API_KEY=gsk_...
# Optional: AES-256 GCM key for cache encryption (base64-encoded 32 bytes)
# CACHE_ENCRYPTION_KEY=...
uvicorn sriti.app:app --host 0.0.0.0 --port 8100
curl -X POST http://localhost:8100/complete \
-H "Content-Type: application/json" \
-d '{
"prompt": "Summarise the following paragraph in one sentence: ...",
"minimum_tier": "tier_3",
"cacheable": true,
"requires_json": false
}'
You can also use the bundled async Python client:
import asyncio
from sriti.client import SritiModelClient
async def main():
client = SritiModelClient(base_url="http://localhost:8100")
response = await client.complete(
prompt="Explain how semantic caching reduces LLM inference costs.",
minimum_tier="tier_3",
cacheable=True
)
print(f"Response: {response.text}")
print(f"Tier Used: {response.tier_used}")
print(f"Latency: {response.latency_ms}ms")
print(f"Cost: ${response.cost_usd}")
asyncio.run(main())
Prompts are embedded using fastembed and compared against predefined semantic task anchors (or an optional lightweight classification head). The classifier runs completely locally in under 10ms.
policy.yaml. For example, conversational questions allow a similarity threshold of 0.75, while structured data extraction requires 0.92.models.yaml are evaluated based on their Pareto score: a composite metric factoring in latency, cost, and historical success probability.Sriti's configuration is managed through two clean YAML files in sriti/config/:
models.yaml: Defines model providers, context windows, cost per token, and tier mappings (default_tier: 1 | 2 | 3).policy.yaml: Defines cache TTLs, per-task similarity thresholds, latency SLOs, and quality gate cutoffs.The test suite lives in the source repository and is not included in the published wheel. Clone the repo first, then run:
git clone https://github.com/sriti-ai/sriti-core.git
cd sriti-core
pip install -e ".[dev]"
python -m pytest
All 39 unit tests run without requiring live LLM API keys or a running Redis instance (mocked in-memory).
uvicorn: command not foundThe uvicorn binary is not on your PATH. Run it via Python instead:
python -m uvicorn sriti.app:app --host 0.0.0.0 --port 8100
ModuleNotFoundError: No module named 'fastapi' (or other missing modules)You installed sriti-core but the server dependencies are not in your active environment. Install them:
pip install sriti-core[dev]
# or from source:
pip install -e ".[dev]"
Could not find a version that satisfies the requirement sriti-core (from versions: none)Your Python version is below 3.11. Verify:
python --version
sriti-core requires Python 3.11+. If your default Python is older (e.g. Anaconda 3.8), create a compatible environment:
conda create -n sriti-env python=3.11 -y
conda activate sriti-env
pip install sriti-core
pip install -e ".[dev]" fails with setup.py not foundYour pip version is too old to support pyproject.toml-based editable installs. Upgrade pip first:
pip install --upgrade pip
pip install -e ".[dev]"
pytest collects 0 items or fails with ModuleNotFoundErrorYour shell is picking up the system/Anaconda pytest instead of the one in your active environment. Use:
python -m pytest
conda activateIf which python still points to Anaconda's base Python, your shell may not have conda initialized. Run:
source /opt/anaconda3/etc/profile.d/conda.sh
conda activate sriti-env
which python # should now show the sriti-env path
This project is licensed under the Apache License 2.0.
Python
100.0%
Sriti Core is an open-source, ultra-efficient intelligent routing proxy and cascading engine for Large Language Models.
Python
2
6 commits
updated Sep 17, 2026
Intelligent, Reliability-Aware Model Routing & Cascade Engine for LLMs
Key Features • Architecture • Quickstart • How It Works • Configuration • API Reference • Testing • Troubleshooting • License
Sriti Core is an open-source, ultra-efficient intelligent routing proxy and cascading engine for Large Language Models.
Instead of routing every single request to expensive frontier models (like Claude 3.5 Sonnet or GPT-4o), Sriti dynamically categorizes the request, checks an encrypted semantic cache, and cascades through a hierarchy of models:
Sriti continuously learns model reliability, evaluates output quality gates, and guarantees strict latency SLOs while dramatically slashing cloud API bills.
fastembed BAAI/bge-small-en-v1.5) mapping prompts across categories (e.g., Code, Extraction, Summarization, QA, Math, Reasoning). User / Application Request
│
▼
┌──────────────────────────────────┐
│ 1. Fast ONNX Classifier │ ──> Task Category & Embeddings
└──────────────────────────────────┘
│
▼
┌──────────────────────────────────┐
│ 2. Semantic Cache (Valkey) │ ──[Hit]──> Return Cached Response
└──────────────────────────────────┘
│ [Miss]
▼
┌──────────────────────────────────┐
│ 3. Model Cascading Engine │
│ │
│ ┌────────────────────────────┐ │
│ │ Tier 3: Local Edge Models │ │ ──[Success]──> Accept & Learn
│ └────────────────────────────┘ │
│ │ [Fail/Timeout] │
│ ┌────────────────────────────┐ │
│ │ Tier 2: Cost-Effective LLM │ │ ──[Success]──> Accept & Learn
│ └────────────────────────────┘ │
│ │ [Fail/Timeout] │
│ ┌────────────────────────────┐ │
│ │ Tier 1: Frontier Models │ │ ──[Guaranteed Execution]
│ └────────────────────────────┘ │
└──────────────────────────────────┘
│
▼
┌──────────────────────────────────┐
│ 4. Update Reliability & Cache │
└──────────────────────────────────┘
Python 3.11+ is required.
Via pip (recommended):
pip install sriti-core
From source:
git clone https://github.com/sriti-ai/sriti-core.git
cd sriti-core
pip install -e ".[dev]"
Create a .env file with your credentials and configuration:
SRITI_REDIS_URL=redis://localhost:6379/0
# Optional: Provider API keys (only configure what you use)
ANTHROPIC_API_KEY=sk-ant-...
OPENAI_API_KEY=sk-...
GROQ_API_KEY=gsk_...
# Optional: AES-256 GCM key for cache encryption (base64-encoded 32 bytes)
# CACHE_ENCRYPTION_KEY=...
uvicorn sriti.app:app --host 0.0.0.0 --port 8100
curl -X POST http://localhost:8100/complete \
-H "Content-Type: application/json" \
-d '{
"prompt": "Summarise the following paragraph in one sentence: ...",
"minimum_tier": "tier_3",
"cacheable": true,
"requires_json": false
}'
You can also use the bundled async Python client:
import asyncio
from sriti.client import SritiModelClient
async def main():
client = SritiModelClient(base_url="http://localhost:8100")
response = await client.complete(
prompt="Explain how semantic caching reduces LLM inference costs.",
minimum_tier="tier_3",
cacheable=True
)
print(f"Response: {response.text}")
print(f"Tier Used: {response.tier_used}")
print(f"Latency: {response.latency_ms}ms")
print(f"Cost: ${response.cost_usd}")
asyncio.run(main())
Prompts are embedded using fastembed and compared against predefined semantic task anchors (or an optional lightweight classification head). The classifier runs completely locally in under 10ms.
policy.yaml. For example, conversational questions allow a similarity threshold of 0.75, while structured data extraction requires 0.92.models.yaml are evaluated based on their Pareto score: a composite metric factoring in latency, cost, and historical success probability.Sriti's configuration is managed through two clean YAML files in sriti/config/:
models.yaml: Defines model providers, context windows, cost per token, and tier mappings (default_tier: 1 | 2 | 3).policy.yaml: Defines cache TTLs, per-task similarity thresholds, latency SLOs, and quality gate cutoffs.The test suite lives in the source repository and is not included in the published wheel. Clone the repo first, then run:
git clone https://github.com/sriti-ai/sriti-core.git
cd sriti-core
pip install -e ".[dev]"
python -m pytest
All 39 unit tests run without requiring live LLM API keys or a running Redis instance (mocked in-memory).
uvicorn: command not foundThe uvicorn binary is not on your PATH. Run it via Python instead:
python -m uvicorn sriti.app:app --host 0.0.0.0 --port 8100
ModuleNotFoundError: No module named 'fastapi' (or other missing modules)You installed sriti-core but the server dependencies are not in your active environment. Install them:
pip install sriti-core[dev]
# or from source:
pip install -e ".[dev]"
Could not find a version that satisfies the requirement sriti-core (from versions: none)Your Python version is below 3.11. Verify:
python --version
sriti-core requires Python 3.11+. If your default Python is older (e.g. Anaconda 3.8), create a compatible environment:
conda create -n sriti-env python=3.11 -y
conda activate sriti-env
pip install sriti-core
pip install -e ".[dev]" fails with setup.py not foundYour pip version is too old to support pyproject.toml-based editable installs. Upgrade pip first:
pip install --upgrade pip
pip install -e ".[dev]"
pytest collects 0 items or fails with ModuleNotFoundErrorYour shell is picking up the system/Anaconda pytest instead of the one in your active environment. Use:
python -m pytest
conda activateIf which python still points to Anaconda's base Python, your shell may not have conda initialized. Run:
source /opt/anaconda3/etc/profile.d/conda.sh
conda activate sriti-env
which python # should now show the sriti-env path
This project is licensed under the Apache License 2.0.
Python
100.0%