routsom/causeval

Statistically rigorous, causal evaluation for LLM apps on top of DeepEval: confidence intervals, causal interventions (RAG grounding, perturbations, agent attribution), and judge validity (bias audits, calibration, PPI).

Python

0

26 commits

updated Oct 1, 2026

See the code

See what people are saying

README

causeval - turn LLM evaluation scores into defensible evidence

causeval

A statistically rigorous, causal evaluation layer for LLM apps, built on top of DeepEval.

Python License Built on DeepEval Status

GitHub followers GitHub stars

Created and maintained by @routsom.


DeepEval measures. It gives you a score. But a single score can't tell you whether it's real (or just noise), why your app produced an output, or whether the LLM judge that produced the score can be trusted.

causeval wraps DeepEval's metrics - they stay the measurement instrument - and adds the three things a score alone can't give you:

  • 📊 Uncertainty - Is the score real? Repeated sampling, variance decomposition, clustered-bootstrap confidence intervals, paired comparisons, and statistically valid gates.
  • 🔬 Causality - What caused this output or failure? RAG context ablation and counterfactual context, input perturbations with metamorphic relations, and agent step-level counterfactual replay.
  • ⚖️ Judge validity - Can we trust the judge? Bias audits, calibration against human labels, prediction-powered inference (PPI), and conformal abstention.

The one rule that drives everything: causeval never reports a bare score. Every public result carries an item count, a repeat count, a confidence interval, the method used, and full provenance (versions, seed, dataset hash, git SHA).

Table of contents

Why causeval?

DeepEval is an excellent measurement library. causeval is not a competitor - it is the layer that turns DeepEval's measurements into decisions you can defend in a code review, a launch meeting, or a paper. Here is precisely what it adds:

Gap when you use metrics alonecauseval's answer
Single-sample scores, threshold pass/fail, flaky results run to runRepeated sampling, clustered-bootstrap CIs, three-valued gates (pass / regression / inconclusive)
No paired tests, no multiple-comparison controlPaired bootstrap of Δ, Wilcoxon, Holm-adjusted CIs across metrics
Faithfulness checks entailment, not dependence on the contextContext Reliance + Counterfactual Adherence: does the answer actually change when you change the evidence?
The judge is never itself evaluatedPosition/verbosity/formatting bias probes, human calibration, PPI, conformal abstention
No robustness or fairness causal testsPerturbation engine with metamorphic relations + counterfactual fairness gaps
Agent failures give you no step-level blameCounterfactual replay: which step, if fixed, would have saved the run?
Wasteful test sets - every item costs the same2PL IRT difficulty/discrimination + information-based pruning
Shared metric objects corrupt scores under concurrency (issue #3356)A fresh metric instance per measurement, enforced by a factory + a test
Telemetry on by default; dotenv loaded at importAn import guard sets the opt-outs before the first import deepeval

Install

causeval targets Python 3.10-3.12.

# from source (recommended while in alpha)
git clone https://github.com/routsom/causeval.git
cd causeval
uv sync                      # core install
uv sync --all-extras         # + optional backends (see below)

Optional extras, each pulling in a heavier dependency only where you need it:

ExtraEnablesPulls in
causeval[nli]NLI-based equivalence checks for perturbation invariancetransformers, torch
causeval[ppi]numerical cross-check of the PPI estimatorppi-python
causeval[irt]cross-check of the 2PL IRT fitgirth
causeval[observational]reference for observational causal estimationdowhy, econml

The core estimators (bootstrap CIs, PPI, IRT, AIPW) are implemented natively on numpy/scipy - the extras are for optional cross-checks and NLI, not for the math.

Quickstart

1. Get a score you can trust (uncertainty)

The core object is Experiment: run each metric over each item R times and get back a RunResult whose estimates carry CIs and variance components.

from causeval import Experiment, MetricSpec

exp = Experiment(
    dataset=goldens,  # DeepEval Goldens or plain dicts
    metrics=[MetricSpec("FaithfulnessMetric", {"threshold": 0.7})],
    repeats=5,  # resample the judge 5x per item
    seed=0,
)
run = exp.run()  # or: await exp.a_run()
print(run.summary())
Run 'run'  (seed=0)
judge=gpt-4o-mini  causeval=0.0.1
  Faithfulness [base]: 0.812 [95% CI 0.771, 0.849] (n=50, R=5, icc=0.34, flaky=6)
      method=cluster_bootstrap_studentized_B=2000

You immediately learn what a bare score hides: the CI, that ~34% of the variance is between items (the rest is judge noise), and that 6 items are flaky (pass rate between 0.2 and 0.8).

2. Decide if B is really better than A (paired gate)

from causeval.stats import compare, gate

comparisons = compare(  # Holm-adjusted across metrics
    baseline_run.measurements,
    candidate_run.measurements,
    margin=0.02,
)
report = gate(comparisons)
print(report.overall)  # "pass" | "regression" | "inconclusive"
raise SystemExit(report.exit_code)  # 0 / 1 / 2 for CI

The gate is three-valued on purpose: "inconclusive" (the CI straddles the margin) is a first-class outcome, so you never ship a regression that hid behind noise, and never block a release on a difference you didn't have the power to detect. Plan that power up front with causeval.stats.plan_power.

3. Ask whether the answer actually used the retrieved context (causality)

DeepEval's Faithfulness tells you the answer is entailed by the context. It cannot tell you the model would have said the same thing without it. causeval intervenes on the context and watches the answer move:

from causeval.interventions import ground, GroundingItem

result = ground(my_rag_app, dataset, repeats=8)
print(result.summary())
RAG grounding (tau=0.5, policy=follow_context)
  context_reliance:          +0.463 [95% CI +0.311, +0.621] (n=40, R=8)
  counterfactual_adherence:  +0.506 [95% CI +0.372, +0.643] (n=40, R=8)
  classes: grounded=20, parametric=20

Counterfactual Adherence edits one supporting fact to a plausible-but-false value and checks whether the answer follows the edit. Items are labelled grounded / parametric / confabulating / mixed, and any answer that passes Faithfulness while being parametric is flagged false-faithful.

4. Audit the judge before you believe it (judge validity)

from causeval.judge_audit import calibrate, ppi_mean_ci

report = calibrate(judge_scores, human_labels)  # Spearman, kappa, isotonic map, ECE + CI
effect = ppi_mean_ci(human_labels, judge_on_labeled, judge_on_unlabeled)  # human mean, debiased

PPI gives you the human-quality mean with a CI that is unbiased (unlike averaging the judge) yet far tighter than using your few human labels alone.

What it adds, feature by feature

ModuleWhat you getKey estimand
stats/Repeated-sampling estimates, cluster-bootstrap CIs (studentized / percentile / BCa), variance decomposition (ICC), paired Δ + Wilcoxon + McNemar, Holm, three-valued gates, power planningμ = E_item[E_repeat[y]] with valid uncertainty
interventions/rag.pyContext Reliance, Counterfactual Adherence, per-chunk leave-one-out effects, grounded/parametric classification, false-faithful flagP(answer follows a counterfactual edit)
interventions/perturb.pyMetamorphic perturbations (paraphrase, reorder, distractor, typo, attribute swap) with invariance rate, paired metric effect, Holm-adjusted fairness gapsΔ metric under a meaning-preserving change
interventions/cot.pyChain-of-thought faithfulness: early-answering curve + area-over-curve, mistake-insertion sensitivitydoes the stated reasoning drive the answer?
judge_audit/Bias probes (position/verbosity/formatting/authorship), calibration (isotonic + ECE), sampling plans, PPI/PPI++, conformal abstention (Learn-then-Test), jury + Krippendorff's αbias effect, E[human score], error rate ≤ α w.p. ≥ 1-δ
attribution/Agent trace schema, record/replay cassettes, step-level counterfactual attribution, decisive-step detectionP(success | do(step_k = oracle)) - P(success | natural)
stats/irt.py2PL item-response theory, Fisher-information pruning, system rankingitem difficulty / discrimination, ability
stats/observational.pyDoubly-robust AIPW effect from production logs, with DoWhy-style refutersATE of a version on a metric
checks/Deterministic pre-judge checks (JSON schema, regex, verifier callables, NLI) + the metamorphic-relation registryfail fast before spending judge tokens
core/Pydantic result schemas, content-hashed LLM cache (SQLite), cost planner, provenancereproducibility & cost control

Does it actually work? (validation benchmarks)

Every statistical claim has a benchmark with a known ground truth, run in CI on synthetic data (fakes, no network), with a live variant for real models. These are the committed offline numbers (bench/results/):

BenchmarkClaimTargetResult
B2 RAG groundingCounterfactual Adherence separates grounded from parametric answersAUROC ≥ 0.91.000
B3 Judge bias recoveryProbes recover an injected position (0.15) + verbosity (0.08) bias inside their CIs; non-injected ≈ 0~95% coveragerecovered 0.156 / 0.070; formatting & authorship not flagged
B4 PPIPPI covers the true human mean; the naive judge-mean does not~95% coverage, narrower than human-onlyPPI 0.99 @ width 0.061; naive 0.00; human-only 0.98 @ width 0.112
B5 Agent attributionThe decisive step equals the injected fault step≥ 90% of tasks100%
B6 IRT2PL parameter recovery; pruning to 50% preserves system rankingcorr ≥ 0.9; Kendall τ ≥ 0.9b=0.99, a=0.95, θ=0.98; τ=1.00

(B1 - CI coverage, false-regression rate, and power - is verified by simulation tests under tests/.)

Live-model benchmarks

The numbers above are the offline benchmarks: synthetic fakes with a known ground truth, run in CI on every commit so the statistics are provably correct without spending a token. Each benchmark also has a live variant that swaps the fake for a real judge/model, so you can publish the same claims against, say, gpt-4o-mini or claude-haiku.

🚧 Status: partial. The baseline row below is a real live run against Claude Haiku 4.5; B2-B6 live runs are still pending a larger budget (tracked in PROGRESS.md). Run any of them yourself with the commands underneath.

BenchmarkLive claim to verifyJudge modelResult
Baseline varianceDeepEval single-sample scores are flaky; repeats + CIs quantify itclaude-haiku-4-5clean items → 1.000, flip 0; borderline items → AnswerRelevancy 0.833 [0.633, 1.000], ContextualRelevancy 0.173 [0.000, 0.373] ¹
B2 RAG groundingCounterfactual Adherence separates grounded vs parametric answers (AUROC ≥ 0.9); Faithfulness cannotclaude-haiku-4-5CA AUROC 0.833 (n=12: 6 fictional vs 6 well-known) ³
B3 Judge biasReal judges show measurable position/verbosity bias with CIsclaude-haiku-4-5position −0.292 [−0.422, −0.126] (flagged), verbosity −0.061 (flagged), formatting/authorship n.s. ²
B4 PPIPPI covers the human mean and beats the naive judge-mean at equal labelsclaude-haiku-4-5PPI 0.510 [0.353, 0.667] covers truth (0.5); eff. n ≈ 41 from 12 labels, CI half of human-only ⁴
B5 Agent attributionDecisive step = injected fault step in ≥ 90% of tasks with a real modelclaude-haiku-4-5decisive step = injected fault in 4/5 valid tasks (0.80) ⁶
B6 IRTOn ≥ 5 real systems, pruning 30-50% of items keeps ranking Kendall τ ≥ 0.96 Claude modelspruned to 50% → Kendall τ 1.000 (systems near ceiling; narrow spread) ⁵

¹ Two real live runs against Claude Haiku 4.5. Clean run (n=6, R=5, temp 0, baseline_live.md): every item a confident 1.000 with zero flip - the reassuring baseline, and an end-to-end pipeline confirmation. Borderline run (n=5, R=5, temp 1.0, deliberately partial/ambiguous items, baseline_live_borderline.md) surfaces what a bare score hides: AnswerRelevancy drops to 0.833 with a wide CI [0.63, 1.00] (a single item's score is genuinely uncertain across the set), ContextualRelevancy correctly falls to 0.17 on the off-topic contexts, and one item wobbled across repeats. Two findings you can only get by measuring: within-item flip is ≈ 0 even at temperature 1, so Haiku 4.5 is a stable judge; but it rated deliberately unsupported claims (invented patent counts, a made-up budget) as fully Faithful = 1.000 - a real judge-leniency signal that causeval's judge audit (bias probes, calibration, PPI) exists to catch.

² Live judge-bias probes against Claude Haiku 4.5 (n=12 neutral answer pairs judged in both orders; n=12 pointwise items; b3_judge_bias_live.md). The headline is real and significant: a position effect of −0.292 [−0.422, −0.126] means that, on answer pairs of equal quality, Haiku picks the second-presented option ~79% of the time - a textbook LLM-judge position bias, and exactly the kind of thing you must correct for before trusting a pairwise judge. It also mildly penalizes padded/verbose answers (−0.061), and shows no significant formatting or authorship-label effect. Unlike the offline B3 (which injects a known 0.15/0.08 bias to prove the probes recover it), the live run measures whatever bias the real judge actually has.

³ Live RAG grounding against a real Claude Haiku 4.5 RAG app (b2_rag_grounding_live.md), using the SPEC's fictional-vs-well-known design. Counterfactual Adherence (does the answer follow a false edit to a supporting fact?) separates the two groups with AUROC 0.833: all 6 fictional items score CA 1.0 with high Context Reliance (the model must use the context), while the well-known items mostly resist the false edit (CA ≈ 0, CR ≈ 0 - the model already knows the answer). It lands below the offline 1.000 / the 0.9 target for an honest reason on real data: on 2 of 12 items the model was swayed by the false context (e.g. it accepted "Romeo and Juliet was written by Dickens") - a real sycophancy signal the metric surfaces. DeepEval Faithfulness would rate every edited-context answer "faithful" and could not make this grounded-vs-parametric distinction at all.

⁴ Live PPI (b4_ppi_live.md) over 40 factual-QA items (20 correct, 20 with a plausible-but-wrong answer) where objective 0/1 correctness is the "human label" and Claude's pointwise score is the predictor f. From a random 12-item labeled subset, PPI estimates the true mean correctness as 0.510 [0.353, 0.667] (truth = 0.500) with an effective sample size ≈ 41 - i.e. 12 human labels bought the precision of ~41, and the CI is half the width of the human-only estimate (0.31 vs 0.58). Honest caveat: on these clear-cut items Claude was a well-calibrated grader (naive judge-mean 0.503, essentially unbiased), so there was little bias to correct here - unlike the offline synthetic judge. PPI's win on this run is label efficiency; its bias-correction matters most on the subtler tasks where judges drift (see the faithfulness leniency in the baseline footnote).

⁵ Live IRT (b6_irt_live.md) using 6 real Claude models as the systems (haiku-4-5, sonnet-4-5, sonnet-5, opus-4-5, opus-4-8, fable-5) on 30 hard short-answer items. Fitting 2PL and pruning to the most-informative 15 items preserved the system ranking exactly (Kendall τ = 1.000). Honest caveat: these are all frontier models, so they cluster near ceiling (accuracy 0.93-1.00) and the true ability spread is narrow - the ranking is close, so preserving it is a lighter test than the offline B6, which validates pruning across a wide simulated ability range. The live run is a real end-to-end confirmation that Fisher-information pruning doesn't scramble the ranking; the offline B6 is the rigorous one.

⁶ Live agent attribution (b5_agent_attribution_live.md) against a real Claude Haiku tool-agent (price → multiply → add-tax tasks) with a wrong tool argument injected at a known step. Counterfactual replay localized the injected fault as the decisive step in 4 of 5 valid tasks (0.80). It lands below the offline 100% for two honest, interesting reasons: (1) real Claude agents often self-correct an injected fault (one task was dropped as "no persistent failure" because the agent noticed and redid the step - a genuine robustness finding), and (2) this environment's Anthropic SDK build rejects temperature=0, so the replay rollouts are noisy. The offline B5 (a deterministic scripted agent) validates the attribution engine at 100% over 40 tasks.

Reproduce a live run (needs an API key for the provider you name; nothing is hardcoded):

export OPENAI_API_KEY=...                # or your provider's key
uv sync

# baseline flakiness with a real judge
uv run python -m causeval.bench.baseline_variance --model gpt-4o-mini

# the full live test suite (opt-in; excluded from the default offline run)
uv run pytest -m live

Live results are written under bench/results/ alongside the offline ones; open a PR with your table and we'll add it here.

CLI

Every command writes a JSON result and prints a human-readable summary; none of them will ever print a bare number.

causeval run          --config eval.yaml   --out runs/   # repeated-sampling run
causeval compare      --baseline a.json --candidate b.json --margin 0.02
causeval gate         --baseline a.json --candidate b.json --margin 0.02  # exit 0/1/2
causeval plan         --pilot-baseline a.json --pilot-candidate b.json --metric Faithfulness --detect 0.03
causeval ground       --config rag.yaml    --out runs/   # RAG causal grounding
causeval audit-judge  --config judge.yaml  --out runs/   # calibration + PPI
causeval attribute    --config agent.yaml  --out runs/   # agent step-level blame

causeval gate sets the process exit code (0 pass, 1 regression, 2 inconclusive), so it drops straight into CI as a release gate.

Design principles

These are enforced by tests, not just documented (see CLAUDE.md):

  1. Wrap DeepEval, never fork it. Depend on deepeval>=4.2,<5; never patch its internals.
  2. All DeepEval imports go through one guarded module that sets telemetry/dotenv opt-outs before the first import. A test enforces that no other module imports it directly.
  3. A fresh metric object per measurement - never share an instance across concurrent calls (DeepEval issue #3356).
  4. Never report a bare score. Every result carries n_items, n_repeats, a CI, the method, and provenance - including CLI output.
  5. No telemetry, no network, no side effects at import time.
  6. No hardcoded model prices or names in logic - they come from user config.
  7. Offline tests never call real LLMs - they use configurable fakes; live tests are opt-in.
  8. Every statistical method has a simulation test proving its coverage or error rate on data with a known answer.

Project status & roadmap

Alpha. The full estimator suite (Phases 0-6) and the causal extensions of Phase 7 are implemented, with 180+ offline tests and strict typing. causeval is a working name.

  • ✅ Phase 0-1: scaffold, core schemas, adapters, statistics (CIs, gates, planning)
  • ✅ Phase 2: RAG causal grounding
  • ✅ Phase 3: judge audit (bias, calibration, PPI, conformal, jury)
  • ✅ Phase 4: perturbations + deterministic checks
  • ✅ Phase 5: agent step-level attribution
  • ✅ Phase 6: IRT pruning + adaptive sampling
  • ✅ Phase 7 (partial): CoT faithfulness, observational AIPW, OpenTelemetry trace import
  • ⏳ Planned: framework harness adapters (LangGraph / OpenAI Agents / Pydantic AI), an HTML report, a pytest plugin, and published live-model benchmark tables

See SPEC.md for the full design and PROGRESS.md for the decision log.

Development

uv sync --all-extras
uv run pytest -m "not live"                       # fast offline suite (must always pass)
uv run pytest -m live                             # real LLM calls; needs API keys
uv run ruff check . && uv run ruff format --check .
uv run mypy src/causeval                          # strict
uv run python -m causeval.bench.rag_grounding     # regenerate a benchmark report

Contributing

Contributions are welcome. See CONTRIBUTING.md for setup, the checks your PR must pass, and the non-negotiable rules (especially: never report a bare score, and every statistical method needs a simulation test). Use the issue templates to report a bug or request a feature.

Author

causeval is created, designed, and maintained by @routsom.

If this project is useful to you, please ⭐ star the repo and follow @routsom for more work on rigorous LLM evaluation. Issues, ideas, and pull requests are welcome.

License

Apache 2.0, matching DeepEval. Portions of DeepEval, where copied, retain their Apache 2.0 headers and are recorded in NOTICE.

causeval is an independent project and is not affiliated with or endorsed by Confident AI, the maintainers of DeepEval.

ab-testing
ai-evaluation
bootstrap
causal-inference
deepeval
evaluation
item-response-theory
llm
llm-evaluation
llm-judge
machine-learning
prediction-powered-inference
python
rag
statistics
uncertainty-quantification

routsom/causeval

Statistically rigorous, causal evaluation for LLM apps on top of DeepEval: confidence intervals, causal interventions (RAG grounding, perturbations, agent attribution), and judge validity (bias audits, calibration, PPI).

Python

0

26 commits

updated Oct 1, 2026

See the code

See what people are saying

README

causeval - turn LLM evaluation scores into defensible evidence

causeval

A statistically rigorous, causal evaluation layer for LLM apps, built on top of DeepEval.

Python License Built on DeepEval Status

GitHub followers GitHub stars

Created and maintained by @routsom.


DeepEval measures. It gives you a score. But a single score can't tell you whether it's real (or just noise), why your app produced an output, or whether the LLM judge that produced the score can be trusted.

causeval wraps DeepEval's metrics - they stay the measurement instrument - and adds the three things a score alone can't give you:

  • 📊 Uncertainty - Is the score real? Repeated sampling, variance decomposition, clustered-bootstrap confidence intervals, paired comparisons, and statistically valid gates.
  • 🔬 Causality - What caused this output or failure? RAG context ablation and counterfactual context, input perturbations with metamorphic relations, and agent step-level counterfactual replay.
  • ⚖️ Judge validity - Can we trust the judge? Bias audits, calibration against human labels, prediction-powered inference (PPI), and conformal abstention.

The one rule that drives everything: causeval never reports a bare score. Every public result carries an item count, a repeat count, a confidence interval, the method used, and full provenance (versions, seed, dataset hash, git SHA).

Table of contents

Why causeval?

DeepEval is an excellent measurement library. causeval is not a competitor - it is the layer that turns DeepEval's measurements into decisions you can defend in a code review, a launch meeting, or a paper. Here is precisely what it adds:

Gap when you use metrics alonecauseval's answer
Single-sample scores, threshold pass/fail, flaky results run to runRepeated sampling, clustered-bootstrap CIs, three-valued gates (pass / regression / inconclusive)
No paired tests, no multiple-comparison controlPaired bootstrap of Δ, Wilcoxon, Holm-adjusted CIs across metrics
Faithfulness checks entailment, not dependence on the contextContext Reliance + Counterfactual Adherence: does the answer actually change when you change the evidence?
The judge is never itself evaluatedPosition/verbosity/formatting bias probes, human calibration, PPI, conformal abstention
No robustness or fairness causal testsPerturbation engine with metamorphic relations + counterfactual fairness gaps
Agent failures give you no step-level blameCounterfactual replay: which step, if fixed, would have saved the run?
Wasteful test sets - every item costs the same2PL IRT difficulty/discrimination + information-based pruning
Shared metric objects corrupt scores under concurrency (issue #3356)A fresh metric instance per measurement, enforced by a factory + a test
Telemetry on by default; dotenv loaded at importAn import guard sets the opt-outs before the first import deepeval

Install

causeval targets Python 3.10-3.12.

# from source (recommended while in alpha)
git clone https://github.com/routsom/causeval.git
cd causeval
uv sync                      # core install
uv sync --all-extras         # + optional backends (see below)

Optional extras, each pulling in a heavier dependency only where you need it:

ExtraEnablesPulls in
causeval[nli]NLI-based equivalence checks for perturbation invariancetransformers, torch
causeval[ppi]numerical cross-check of the PPI estimatorppi-python
causeval[irt]cross-check of the 2PL IRT fitgirth
causeval[observational]reference for observational causal estimationdowhy, econml

The core estimators (bootstrap CIs, PPI, IRT, AIPW) are implemented natively on numpy/scipy - the extras are for optional cross-checks and NLI, not for the math.

Quickstart

1. Get a score you can trust (uncertainty)

The core object is Experiment: run each metric over each item R times and get back a RunResult whose estimates carry CIs and variance components.

from causeval import Experiment, MetricSpec

exp = Experiment(
    dataset=goldens,  # DeepEval Goldens or plain dicts
    metrics=[MetricSpec("FaithfulnessMetric", {"threshold": 0.7})],
    repeats=5,  # resample the judge 5x per item
    seed=0,
)
run = exp.run()  # or: await exp.a_run()
print(run.summary())
Run 'run'  (seed=0)
judge=gpt-4o-mini  causeval=0.0.1
  Faithfulness [base]: 0.812 [95% CI 0.771, 0.849] (n=50, R=5, icc=0.34, flaky=6)
      method=cluster_bootstrap_studentized_B=2000

You immediately learn what a bare score hides: the CI, that ~34% of the variance is between items (the rest is judge noise), and that 6 items are flaky (pass rate between 0.2 and 0.8).

2. Decide if B is really better than A (paired gate)

from causeval.stats import compare, gate

comparisons = compare(  # Holm-adjusted across metrics
    baseline_run.measurements,
    candidate_run.measurements,
    margin=0.02,
)
report = gate(comparisons)
print(report.overall)  # "pass" | "regression" | "inconclusive"
raise SystemExit(report.exit_code)  # 0 / 1 / 2 for CI

The gate is three-valued on purpose: "inconclusive" (the CI straddles the margin) is a first-class outcome, so you never ship a regression that hid behind noise, and never block a release on a difference you didn't have the power to detect. Plan that power up front with causeval.stats.plan_power.

3. Ask whether the answer actually used the retrieved context (causality)

DeepEval's Faithfulness tells you the answer is entailed by the context. It cannot tell you the model would have said the same thing without it. causeval intervenes on the context and watches the answer move:

from causeval.interventions import ground, GroundingItem

result = ground(my_rag_app, dataset, repeats=8)
print(result.summary())
RAG grounding (tau=0.5, policy=follow_context)
  context_reliance:          +0.463 [95% CI +0.311, +0.621] (n=40, R=8)
  counterfactual_adherence:  +0.506 [95% CI +0.372, +0.643] (n=40, R=8)
  classes: grounded=20, parametric=20

Counterfactual Adherence edits one supporting fact to a plausible-but-false value and checks whether the answer follows the edit. Items are labelled grounded / parametric / confabulating / mixed, and any answer that passes Faithfulness while being parametric is flagged false-faithful.

4. Audit the judge before you believe it (judge validity)

from causeval.judge_audit import calibrate, ppi_mean_ci

report = calibrate(judge_scores, human_labels)  # Spearman, kappa, isotonic map, ECE + CI
effect = ppi_mean_ci(human_labels, judge_on_labeled, judge_on_unlabeled)  # human mean, debiased

PPI gives you the human-quality mean with a CI that is unbiased (unlike averaging the judge) yet far tighter than using your few human labels alone.

What it adds, feature by feature

ModuleWhat you getKey estimand
stats/Repeated-sampling estimates, cluster-bootstrap CIs (studentized / percentile / BCa), variance decomposition (ICC), paired Δ + Wilcoxon + McNemar, Holm, three-valued gates, power planningμ = E_item[E_repeat[y]] with valid uncertainty
interventions/rag.pyContext Reliance, Counterfactual Adherence, per-chunk leave-one-out effects, grounded/parametric classification, false-faithful flagP(answer follows a counterfactual edit)
interventions/perturb.pyMetamorphic perturbations (paraphrase, reorder, distractor, typo, attribute swap) with invariance rate, paired metric effect, Holm-adjusted fairness gapsΔ metric under a meaning-preserving change
interventions/cot.pyChain-of-thought faithfulness: early-answering curve + area-over-curve, mistake-insertion sensitivitydoes the stated reasoning drive the answer?
judge_audit/Bias probes (position/verbosity/formatting/authorship), calibration (isotonic + ECE), sampling plans, PPI/PPI++, conformal abstention (Learn-then-Test), jury + Krippendorff's αbias effect, E[human score], error rate ≤ α w.p. ≥ 1-δ
attribution/Agent trace schema, record/replay cassettes, step-level counterfactual attribution, decisive-step detectionP(success | do(step_k = oracle)) - P(success | natural)
stats/irt.py2PL item-response theory, Fisher-information pruning, system rankingitem difficulty / discrimination, ability
stats/observational.pyDoubly-robust AIPW effect from production logs, with DoWhy-style refutersATE of a version on a metric
checks/Deterministic pre-judge checks (JSON schema, regex, verifier callables, NLI) + the metamorphic-relation registryfail fast before spending judge tokens
core/Pydantic result schemas, content-hashed LLM cache (SQLite), cost planner, provenancereproducibility & cost control

Does it actually work? (validation benchmarks)

Every statistical claim has a benchmark with a known ground truth, run in CI on synthetic data (fakes, no network), with a live variant for real models. These are the committed offline numbers (bench/results/):

BenchmarkClaimTargetResult
B2 RAG groundingCounterfactual Adherence separates grounded from parametric answersAUROC ≥ 0.91.000
B3 Judge bias recoveryProbes recover an injected position (0.15) + verbosity (0.08) bias inside their CIs; non-injected ≈ 0~95% coveragerecovered 0.156 / 0.070; formatting & authorship not flagged
B4 PPIPPI covers the true human mean; the naive judge-mean does not~95% coverage, narrower than human-onlyPPI 0.99 @ width 0.061; naive 0.00; human-only 0.98 @ width 0.112
B5 Agent attributionThe decisive step equals the injected fault step≥ 90% of tasks100%
B6 IRT2PL parameter recovery; pruning to 50% preserves system rankingcorr ≥ 0.9; Kendall τ ≥ 0.9b=0.99, a=0.95, θ=0.98; τ=1.00

(B1 - CI coverage, false-regression rate, and power - is verified by simulation tests under tests/.)

Live-model benchmarks

The numbers above are the offline benchmarks: synthetic fakes with a known ground truth, run in CI on every commit so the statistics are provably correct without spending a token. Each benchmark also has a live variant that swaps the fake for a real judge/model, so you can publish the same claims against, say, gpt-4o-mini or claude-haiku.

🚧 Status: partial. The baseline row below is a real live run against Claude Haiku 4.5; B2-B6 live runs are still pending a larger budget (tracked in PROGRESS.md). Run any of them yourself with the commands underneath.

BenchmarkLive claim to verifyJudge modelResult
Baseline varianceDeepEval single-sample scores are flaky; repeats + CIs quantify itclaude-haiku-4-5clean items → 1.000, flip 0; borderline items → AnswerRelevancy 0.833 [0.633, 1.000], ContextualRelevancy 0.173 [0.000, 0.373] ¹
B2 RAG groundingCounterfactual Adherence separates grounded vs parametric answers (AUROC ≥ 0.9); Faithfulness cannotclaude-haiku-4-5CA AUROC 0.833 (n=12: 6 fictional vs 6 well-known) ³
B3 Judge biasReal judges show measurable position/verbosity bias with CIsclaude-haiku-4-5position −0.292 [−0.422, −0.126] (flagged), verbosity −0.061 (flagged), formatting/authorship n.s. ²
B4 PPIPPI covers the human mean and beats the naive judge-mean at equal labelsclaude-haiku-4-5PPI 0.510 [0.353, 0.667] covers truth (0.5); eff. n ≈ 41 from 12 labels, CI half of human-only ⁴
B5 Agent attributionDecisive step = injected fault step in ≥ 90% of tasks with a real modelclaude-haiku-4-5decisive step = injected fault in 4/5 valid tasks (0.80) ⁶
B6 IRTOn ≥ 5 real systems, pruning 30-50% of items keeps ranking Kendall τ ≥ 0.96 Claude modelspruned to 50% → Kendall τ 1.000 (systems near ceiling; narrow spread) ⁵

¹ Two real live runs against Claude Haiku 4.5. Clean run (n=6, R=5, temp 0, baseline_live.md): every item a confident 1.000 with zero flip - the reassuring baseline, and an end-to-end pipeline confirmation. Borderline run (n=5, R=5, temp 1.0, deliberately partial/ambiguous items, baseline_live_borderline.md) surfaces what a bare score hides: AnswerRelevancy drops to 0.833 with a wide CI [0.63, 1.00] (a single item's score is genuinely uncertain across the set), ContextualRelevancy correctly falls to 0.17 on the off-topic contexts, and one item wobbled across repeats. Two findings you can only get by measuring: within-item flip is ≈ 0 even at temperature 1, so Haiku 4.5 is a stable judge; but it rated deliberately unsupported claims (invented patent counts, a made-up budget) as fully Faithful = 1.000 - a real judge-leniency signal that causeval's judge audit (bias probes, calibration, PPI) exists to catch.

² Live judge-bias probes against Claude Haiku 4.5 (n=12 neutral answer pairs judged in both orders; n=12 pointwise items; b3_judge_bias_live.md). The headline is real and significant: a position effect of −0.292 [−0.422, −0.126] means that, on answer pairs of equal quality, Haiku picks the second-presented option ~79% of the time - a textbook LLM-judge position bias, and exactly the kind of thing you must correct for before trusting a pairwise judge. It also mildly penalizes padded/verbose answers (−0.061), and shows no significant formatting or authorship-label effect. Unlike the offline B3 (which injects a known 0.15/0.08 bias to prove the probes recover it), the live run measures whatever bias the real judge actually has.

³ Live RAG grounding against a real Claude Haiku 4.5 RAG app (b2_rag_grounding_live.md), using the SPEC's fictional-vs-well-known design. Counterfactual Adherence (does the answer follow a false edit to a supporting fact?) separates the two groups with AUROC 0.833: all 6 fictional items score CA 1.0 with high Context Reliance (the model must use the context), while the well-known items mostly resist the false edit (CA ≈ 0, CR ≈ 0 - the model already knows the answer). It lands below the offline 1.000 / the 0.9 target for an honest reason on real data: on 2 of 12 items the model was swayed by the false context (e.g. it accepted "Romeo and Juliet was written by Dickens") - a real sycophancy signal the metric surfaces. DeepEval Faithfulness would rate every edited-context answer "faithful" and could not make this grounded-vs-parametric distinction at all.

⁴ Live PPI (b4_ppi_live.md) over 40 factual-QA items (20 correct, 20 with a plausible-but-wrong answer) where objective 0/1 correctness is the "human label" and Claude's pointwise score is the predictor f. From a random 12-item labeled subset, PPI estimates the true mean correctness as 0.510 [0.353, 0.667] (truth = 0.500) with an effective sample size ≈ 41 - i.e. 12 human labels bought the precision of ~41, and the CI is half the width of the human-only estimate (0.31 vs 0.58). Honest caveat: on these clear-cut items Claude was a well-calibrated grader (naive judge-mean 0.503, essentially unbiased), so there was little bias to correct here - unlike the offline synthetic judge. PPI's win on this run is label efficiency; its bias-correction matters most on the subtler tasks where judges drift (see the faithfulness leniency in the baseline footnote).

⁵ Live IRT (b6_irt_live.md) using 6 real Claude models as the systems (haiku-4-5, sonnet-4-5, sonnet-5, opus-4-5, opus-4-8, fable-5) on 30 hard short-answer items. Fitting 2PL and pruning to the most-informative 15 items preserved the system ranking exactly (Kendall τ = 1.000). Honest caveat: these are all frontier models, so they cluster near ceiling (accuracy 0.93-1.00) and the true ability spread is narrow - the ranking is close, so preserving it is a lighter test than the offline B6, which validates pruning across a wide simulated ability range. The live run is a real end-to-end confirmation that Fisher-information pruning doesn't scramble the ranking; the offline B6 is the rigorous one.

⁶ Live agent attribution (b5_agent_attribution_live.md) against a real Claude Haiku tool-agent (price → multiply → add-tax tasks) with a wrong tool argument injected at a known step. Counterfactual replay localized the injected fault as the decisive step in 4 of 5 valid tasks (0.80). It lands below the offline 100% for two honest, interesting reasons: (1) real Claude agents often self-correct an injected fault (one task was dropped as "no persistent failure" because the agent noticed and redid the step - a genuine robustness finding), and (2) this environment's Anthropic SDK build rejects temperature=0, so the replay rollouts are noisy. The offline B5 (a deterministic scripted agent) validates the attribution engine at 100% over 40 tasks.

Reproduce a live run (needs an API key for the provider you name; nothing is hardcoded):

export OPENAI_API_KEY=...                # or your provider's key
uv sync

# baseline flakiness with a real judge
uv run python -m causeval.bench.baseline_variance --model gpt-4o-mini

# the full live test suite (opt-in; excluded from the default offline run)
uv run pytest -m live

Live results are written under bench/results/ alongside the offline ones; open a PR with your table and we'll add it here.

CLI

Every command writes a JSON result and prints a human-readable summary; none of them will ever print a bare number.

causeval run          --config eval.yaml   --out runs/   # repeated-sampling run
causeval compare      --baseline a.json --candidate b.json --margin 0.02
causeval gate         --baseline a.json --candidate b.json --margin 0.02  # exit 0/1/2
causeval plan         --pilot-baseline a.json --pilot-candidate b.json --metric Faithfulness --detect 0.03
causeval ground       --config rag.yaml    --out runs/   # RAG causal grounding
causeval audit-judge  --config judge.yaml  --out runs/   # calibration + PPI
causeval attribute    --config agent.yaml  --out runs/   # agent step-level blame

causeval gate sets the process exit code (0 pass, 1 regression, 2 inconclusive), so it drops straight into CI as a release gate.

Design principles

These are enforced by tests, not just documented (see CLAUDE.md):

  1. Wrap DeepEval, never fork it. Depend on deepeval>=4.2,<5; never patch its internals.
  2. All DeepEval imports go through one guarded module that sets telemetry/dotenv opt-outs before the first import. A test enforces that no other module imports it directly.
  3. A fresh metric object per measurement - never share an instance across concurrent calls (DeepEval issue #3356).
  4. Never report a bare score. Every result carries n_items, n_repeats, a CI, the method, and provenance - including CLI output.
  5. No telemetry, no network, no side effects at import time.
  6. No hardcoded model prices or names in logic - they come from user config.
  7. Offline tests never call real LLMs - they use configurable fakes; live tests are opt-in.
  8. Every statistical method has a simulation test proving its coverage or error rate on data with a known answer.

Project status & roadmap

Alpha. The full estimator suite (Phases 0-6) and the causal extensions of Phase 7 are implemented, with 180+ offline tests and strict typing. causeval is a working name.

  • ✅ Phase 0-1: scaffold, core schemas, adapters, statistics (CIs, gates, planning)
  • ✅ Phase 2: RAG causal grounding
  • ✅ Phase 3: judge audit (bias, calibration, PPI, conformal, jury)
  • ✅ Phase 4: perturbations + deterministic checks
  • ✅ Phase 5: agent step-level attribution
  • ✅ Phase 6: IRT pruning + adaptive sampling
  • ✅ Phase 7 (partial): CoT faithfulness, observational AIPW, OpenTelemetry trace import
  • ⏳ Planned: framework harness adapters (LangGraph / OpenAI Agents / Pydantic AI), an HTML report, a pytest plugin, and published live-model benchmark tables

See SPEC.md for the full design and PROGRESS.md for the decision log.

Development

uv sync --all-extras
uv run pytest -m "not live"                       # fast offline suite (must always pass)
uv run pytest -m live                             # real LLM calls; needs API keys
uv run ruff check . && uv run ruff format --check .
uv run mypy src/causeval                          # strict
uv run python -m causeval.bench.rag_grounding     # regenerate a benchmark report

Contributing

Contributions are welcome. See CONTRIBUTING.md for setup, the checks your PR must pass, and the non-negotiable rules (especially: never report a bare score, and every statistical method needs a simulation test). Use the issue templates to report a bug or request a feature.

Author

causeval is created, designed, and maintained by @routsom.

If this project is useful to you, please ⭐ star the repo and follow @routsom for more work on rigorous LLM evaluation. Issues, ideas, and pull requests are welcome.

License

Apache 2.0, matching DeepEval. Portions of DeepEval, where copied, retain their Apache 2.0 headers and are recorded in NOTICE.

causeval is an independent project and is not affiliated with or endorsed by Confident AI, the maintainers of DeepEval.

ab-testing
ai-evaluation
bootstrap
causal-inference
deepeval
evaluation
item-response-theory
llm
llm-evaluation
llm-judge
machine-learning
prediction-powered-inference
python
rag
statistics
uncertainty-quantification

Languages

Python

100.0%