emanuelecz/rag_project

Python

1

8 commits

updated Jun 15, 2026

See the code

README

NTP Safety Assistant — RAG over Spanish Workplace Safety Regulations

A production-style Retrieval-Augmented Generation system that answers questions about Spanish occupational safety regulations (INSST Notas Técnicas de Prevención), with grounded citations, streaming responses, and a full LLM-as-judge evaluation suite — including a meta-evaluation that validates the judge itself against human labels.

Python FastAPI Next.js TypeScript Docker License

How it works

PDFs (16 NTP regulations)
   │  ingestion + OCR (poppler / tesseract)
   ▼
Chunking ──► Voyage AI embeddings ──► ChromaDB
                                         │
User question ──► Hybrid retrieval (vector + BM25) ──► Voyage rerank-2.5
                                         │
                              Claude (Anthropic) ──► streamed answer + [1][2] citations
  • Hybrid retrieval: dense vector search and BM25 keyword search combined in a 50/50 ensemble, then reranked with Voyage rerank-2.5 to select the top-k chunks.
  • Grounded generation: Claude answers strictly from retrieved context, cites sources inline, and refuses when the context is insufficient.
  • Streaming API: answers stream token-by-token over Server-Sent Events.
  • Bilingual: Spanish and English prompts and UI.

Evaluation (the part most RAG demos skip)

The system is evaluated end-to-end on a 50-question dataset with human-verified reference answers (backend/eval/):

MetricResult
Retrieval hit rate @80.98
Retrieval MRR0.955
Generation pass rate (LLM judge)0.88
Faithfulness / Relevance / Completeness4.8 / 4.86 / 4.58 (out of 5)

Two evaluation layers:

  1. Pipeline eval (run_eval.py) — retrieval metrics (hit@k, precision@k, MRR) plus a Claude judge that scores every answer on faithfulness, relevance, and completeness with structured fault reporting (NOT_SUPPORTED, CONTRADICTS_REFERENCE, MISSING_FACT).
  2. Meta-eval (meta_eval.py) — validates the judge itself: a second, independent GPT-4o judge is compared against human labels (including adversarial negatives — deliberately corrupted answers) and gated on TPR/TNR ≥ 85%, so hallucinations can't slip past an over-lenient judge. Exits non-zero on failure to block CI.

Full results: backend/eval/data/results.json

Tech stack

LayerTechnology
BackendFastAPI, LangChain, Python 3.11
LLMClaude (Anthropic API)
Embeddings + rerankingVoyage AI
Vector storeChromaDB
Eval judgesClaude + GPT-4o (structured outputs)
FrontendNext.js 14, React 18, TypeScript, Tailwind CSS
InfraDocker, Docker Compose

Quick start

# 1. Configure API keys
cp backend/.env.example backend/.env   # then fill in your keys

# 2. Build and run
docker compose up --build

# 3. Index the regulations (first run only)
curl -X POST http://localhost:8000/ingest

Open http://localhost:3000 and ask away.

Environment variables (backend/.env)

VariablePurpose
ANTHROPIC_API_KEYAnswer generation + suggestions (Claude)
VOYAGE_API_KEYEmbeddings + reranking
OPENAI_API_KEYMeta-evaluation judge only (GPT-4o)

API

EndpointDescription
POST /askAsk a question — streams the answer and sources via SSE
GET /suggestionsLLM-generated suggested questions based on the indexed corpus
POST /ingestParse, chunk, embed and store the PDF corpus
GET /healthHealth check

Project structure

├── backend/
│   ├── main.py              # FastAPI app (ask / suggestions / ingest / health)
│   ├── ingest.py            # CLI ingestion entry point
│   ├── services/
│   │   ├── ingestion.py     # PDF parsing + OCR
│   │   ├── chunking.py      # Chunking + ChromaDB vector store
│   │   ├── retrieval.py     # Hybrid search (vector + BM25) + reranking
│   │   └── generate.py      # Claude generation with citations + streaming
│   └── eval/
│       ├── run_eval.py      # Full pipeline eval: retrieval + LLM judge
│       ├── retrieval_eval.py# hit@k, precision@k, MRR
│       ├── judge.py         # Claude judge (1–5 rubric, fault categories)
│       ├── meta_eval.py     # Judge-vs-human validation (TPR/TNR gates)
│       ├── meta_judge.py    # Independent GPT-4o binary judge
│       ├── rag_runner.py    # Builds the eval dataset answers
│       ├── grade.py         # Interactive human labeling CLI
│       └── data/            # Datasets, human labels, results
├── frontend/                # Next.js chat UI (streaming, bilingual)
├── ntp_regulations/         # Source PDF corpus (INSST NTPs)
└── docker-compose.yml

Running the evaluations

python -m backend.eval.run_eval    # retrieval + generation eval → data/results.json
python -m backend.eval.meta_eval   # validate the judge against human labels
python -m backend.eval.grade       # label answers interactively

License

MIT

emanuelecz/rag_project

Python

1

8 commits

updated Jun 15, 2026

See the code

README

NTP Safety Assistant — RAG over Spanish Workplace Safety Regulations

A production-style Retrieval-Augmented Generation system that answers questions about Spanish occupational safety regulations (INSST Notas Técnicas de Prevención), with grounded citations, streaming responses, and a full LLM-as-judge evaluation suite — including a meta-evaluation that validates the judge itself against human labels.

Python FastAPI Next.js TypeScript Docker License

How it works

PDFs (16 NTP regulations)
   │  ingestion + OCR (poppler / tesseract)
   ▼
Chunking ──► Voyage AI embeddings ──► ChromaDB
                                         │
User question ──► Hybrid retrieval (vector + BM25) ──► Voyage rerank-2.5
                                         │
                              Claude (Anthropic) ──► streamed answer + [1][2] citations
  • Hybrid retrieval: dense vector search and BM25 keyword search combined in a 50/50 ensemble, then reranked with Voyage rerank-2.5 to select the top-k chunks.
  • Grounded generation: Claude answers strictly from retrieved context, cites sources inline, and refuses when the context is insufficient.
  • Streaming API: answers stream token-by-token over Server-Sent Events.
  • Bilingual: Spanish and English prompts and UI.

Evaluation (the part most RAG demos skip)

The system is evaluated end-to-end on a 50-question dataset with human-verified reference answers (backend/eval/):

MetricResult
Retrieval hit rate @80.98
Retrieval MRR0.955
Generation pass rate (LLM judge)0.88
Faithfulness / Relevance / Completeness4.8 / 4.86 / 4.58 (out of 5)

Two evaluation layers:

  1. Pipeline eval (run_eval.py) — retrieval metrics (hit@k, precision@k, MRR) plus a Claude judge that scores every answer on faithfulness, relevance, and completeness with structured fault reporting (NOT_SUPPORTED, CONTRADICTS_REFERENCE, MISSING_FACT).
  2. Meta-eval (meta_eval.py) — validates the judge itself: a second, independent GPT-4o judge is compared against human labels (including adversarial negatives — deliberately corrupted answers) and gated on TPR/TNR ≥ 85%, so hallucinations can't slip past an over-lenient judge. Exits non-zero on failure to block CI.

Full results: backend/eval/data/results.json

Tech stack

LayerTechnology
BackendFastAPI, LangChain, Python 3.11
LLMClaude (Anthropic API)
Embeddings + rerankingVoyage AI
Vector storeChromaDB
Eval judgesClaude + GPT-4o (structured outputs)
FrontendNext.js 14, React 18, TypeScript, Tailwind CSS
InfraDocker, Docker Compose

Quick start

# 1. Configure API keys
cp backend/.env.example backend/.env   # then fill in your keys

# 2. Build and run
docker compose up --build

# 3. Index the regulations (first run only)
curl -X POST http://localhost:8000/ingest

Open http://localhost:3000 and ask away.

Environment variables (backend/.env)

VariablePurpose
ANTHROPIC_API_KEYAnswer generation + suggestions (Claude)
VOYAGE_API_KEYEmbeddings + reranking
OPENAI_API_KEYMeta-evaluation judge only (GPT-4o)

API

EndpointDescription
POST /askAsk a question — streams the answer and sources via SSE
GET /suggestionsLLM-generated suggested questions based on the indexed corpus
POST /ingestParse, chunk, embed and store the PDF corpus
GET /healthHealth check

Project structure

├── backend/
│   ├── main.py              # FastAPI app (ask / suggestions / ingest / health)
│   ├── ingest.py            # CLI ingestion entry point
│   ├── services/
│   │   ├── ingestion.py     # PDF parsing + OCR
│   │   ├── chunking.py      # Chunking + ChromaDB vector store
│   │   ├── retrieval.py     # Hybrid search (vector + BM25) + reranking
│   │   └── generate.py      # Claude generation with citations + streaming
│   └── eval/
│       ├── run_eval.py      # Full pipeline eval: retrieval + LLM judge
│       ├── retrieval_eval.py# hit@k, precision@k, MRR
│       ├── judge.py         # Claude judge (1–5 rubric, fault categories)
│       ├── meta_eval.py     # Judge-vs-human validation (TPR/TNR gates)
│       ├── meta_judge.py    # Independent GPT-4o binary judge
│       ├── rag_runner.py    # Builds the eval dataset answers
│       ├── grade.py         # Interactive human labeling CLI
│       └── data/            # Datasets, human labels, results
├── frontend/                # Next.js chat UI (streaming, bilingual)
├── ntp_regulations/         # Source PDF corpus (INSST NTPs)
└── docker-compose.yml

Running the evaluations

python -m backend.eval.run_eval    # retrieval + generation eval → data/results.json
python -m backend.eval.meta_eval   # validate the judge against human labels
python -m backend.eval.grade       # label answers interactively

License

MIT

Languages

Python

53.0%

TypeScript

43.9%

CSS

1.6%

Dockerfile

1.4%