Statistically rigorous, causal evaluation for LLM apps on top of DeepEval: confidence intervals, causal interventions (RAG grounding, perturbations, agent attribution), and judge validity (bias audits, calibration, PPI).
Python
0
26 commits
updated Oct 1, 2026
A statistically rigorous, causal evaluation layer for LLM apps, built on top of DeepEval.
Created and maintained by @routsom.
DeepEval measures. It gives you a score. But a single score can't tell you whether it's real (or just noise), why your app produced an output, or whether the LLM judge that produced the score can be trusted.
causeval wraps DeepEval's metrics - they stay the measurement instrument - and adds the three things a score alone can't give you:
The one rule that drives everything: causeval never reports a bare score. Every public result carries an item count, a repeat count, a confidence interval, the method used, and full provenance (versions, seed, dataset hash, git SHA).
DeepEval is an excellent measurement library. causeval is not a competitor - it is the layer that turns DeepEval's measurements into decisions you can defend in a code review, a launch meeting, or a paper. Here is precisely what it adds:
| Gap when you use metrics alone | causeval's answer |
|---|---|
| Single-sample scores, threshold pass/fail, flaky results run to run | Repeated sampling, clustered-bootstrap CIs, three-valued gates (pass / regression / inconclusive) |
| No paired tests, no multiple-comparison control | Paired bootstrap of Δ, Wilcoxon, Holm-adjusted CIs across metrics |
| Faithfulness checks entailment, not dependence on the context | Context Reliance + Counterfactual Adherence: does the answer actually change when you change the evidence? |
| The judge is never itself evaluated | Position/verbosity/formatting bias probes, human calibration, PPI, conformal abstention |
| No robustness or fairness causal tests | Perturbation engine with metamorphic relations + counterfactual fairness gaps |
| Agent failures give you no step-level blame | Counterfactual replay: which step, if fixed, would have saved the run? |
| Wasteful test sets - every item costs the same | 2PL IRT difficulty/discrimination + information-based pruning |
| Shared metric objects corrupt scores under concurrency (issue #3356) | A fresh metric instance per measurement, enforced by a factory + a test |
| Telemetry on by default; dotenv loaded at import | An import guard sets the opt-outs before the first import deepeval |
causeval targets Python 3.10-3.12.
# from source (recommended while in alpha)
git clone https://github.com/routsom/causeval.git
cd causeval
uv sync # core install
uv sync --all-extras # + optional backends (see below)
Optional extras, each pulling in a heavier dependency only where you need it:
| Extra | Enables | Pulls in |
|---|---|---|
causeval[nli] | NLI-based equivalence checks for perturbation invariance | transformers, torch |
causeval[ppi] | numerical cross-check of the PPI estimator | ppi-python |
causeval[irt] | cross-check of the 2PL IRT fit | girth |
causeval[observational] | reference for observational causal estimation | dowhy, econml |
The core estimators (bootstrap CIs, PPI, IRT, AIPW) are implemented natively on numpy/scipy - the extras are for optional cross-checks and NLI, not for the math.
The core object is Experiment: run each metric over each item R times and get back a
RunResult whose estimates carry CIs and variance components.
from causeval import Experiment, MetricSpec
exp = Experiment(
dataset=goldens, # DeepEval Goldens or plain dicts
metrics=[MetricSpec("FaithfulnessMetric", {"threshold": 0.7})],
repeats=5, # resample the judge 5x per item
seed=0,
)
run = exp.run() # or: await exp.a_run()
print(run.summary())
Run 'run' (seed=0)
judge=gpt-4o-mini causeval=0.0.1
Faithfulness [base]: 0.812 [95% CI 0.771, 0.849] (n=50, R=5, icc=0.34, flaky=6)
method=cluster_bootstrap_studentized_B=2000
You immediately learn what a bare score hides: the CI, that ~34% of the variance is between items (the rest is judge noise), and that 6 items are flaky (pass rate between 0.2 and 0.8).
from causeval.stats import compare, gate
comparisons = compare( # Holm-adjusted across metrics
baseline_run.measurements,
candidate_run.measurements,
margin=0.02,
)
report = gate(comparisons)
print(report.overall) # "pass" | "regression" | "inconclusive"
raise SystemExit(report.exit_code) # 0 / 1 / 2 for CI
The gate is three-valued on purpose: "inconclusive" (the CI straddles the margin) is a
first-class outcome, so you never ship a regression that hid behind noise, and never block a
release on a difference you didn't have the power to detect. Plan that power up front with
causeval.stats.plan_power.
DeepEval's Faithfulness tells you the answer is entailed by the context. It cannot tell you the model would have said the same thing without it. causeval intervenes on the context and watches the answer move:
from causeval.interventions import ground, GroundingItem
result = ground(my_rag_app, dataset, repeats=8)
print(result.summary())
RAG grounding (tau=0.5, policy=follow_context)
context_reliance: +0.463 [95% CI +0.311, +0.621] (n=40, R=8)
counterfactual_adherence: +0.506 [95% CI +0.372, +0.643] (n=40, R=8)
classes: grounded=20, parametric=20
Counterfactual Adherence edits one supporting fact to a plausible-but-false value and checks whether the answer follows the edit. Items are labelled grounded / parametric / confabulating / mixed, and any answer that passes Faithfulness while being parametric is flagged false-faithful.
from causeval.judge_audit import calibrate, ppi_mean_ci
report = calibrate(judge_scores, human_labels) # Spearman, kappa, isotonic map, ECE + CI
effect = ppi_mean_ci(human_labels, judge_on_labeled, judge_on_unlabeled) # human mean, debiased
PPI gives you the human-quality mean with a CI that is unbiased (unlike averaging the judge) yet far tighter than using your few human labels alone.
| Module | What you get | Key estimand |
|---|---|---|
stats/ | Repeated-sampling estimates, cluster-bootstrap CIs (studentized / percentile / BCa), variance decomposition (ICC), paired Δ + Wilcoxon + McNemar, Holm, three-valued gates, power planning | μ = E_item[E_repeat[y]] with valid uncertainty |
interventions/rag.py | Context Reliance, Counterfactual Adherence, per-chunk leave-one-out effects, grounded/parametric classification, false-faithful flag | P(answer follows a counterfactual edit) |
interventions/perturb.py | Metamorphic perturbations (paraphrase, reorder, distractor, typo, attribute swap) with invariance rate, paired metric effect, Holm-adjusted fairness gaps | Δ metric under a meaning-preserving change |
interventions/cot.py | Chain-of-thought faithfulness: early-answering curve + area-over-curve, mistake-insertion sensitivity | does the stated reasoning drive the answer? |
judge_audit/ | Bias probes (position/verbosity/formatting/authorship), calibration (isotonic + ECE), sampling plans, PPI/PPI++, conformal abstention (Learn-then-Test), jury + Krippendorff's α | bias effect, E[human score], error rate ≤ α w.p. ≥ 1-δ |
attribution/ | Agent trace schema, record/replay cassettes, step-level counterfactual attribution, decisive-step detection | P(success | do(step_k = oracle)) - P(success | natural) |
stats/irt.py | 2PL item-response theory, Fisher-information pruning, system ranking | item difficulty / discrimination, ability |
stats/observational.py | Doubly-robust AIPW effect from production logs, with DoWhy-style refuters | ATE of a version on a metric |
checks/ | Deterministic pre-judge checks (JSON schema, regex, verifier callables, NLI) + the metamorphic-relation registry | fail fast before spending judge tokens |
core/ | Pydantic result schemas, content-hashed LLM cache (SQLite), cost planner, provenance | reproducibility & cost control |
Every statistical claim has a benchmark with a known ground truth, run in CI on synthetic
data (fakes, no network), with a live variant for real models. These are the committed offline
numbers (bench/results/):
| Benchmark | Claim | Target | Result |
|---|---|---|---|
| B2 RAG grounding | Counterfactual Adherence separates grounded from parametric answers | AUROC ≥ 0.9 | 1.000 |
| B3 Judge bias recovery | Probes recover an injected position (0.15) + verbosity (0.08) bias inside their CIs; non-injected ≈ 0 | ~95% coverage | recovered 0.156 / 0.070; formatting & authorship not flagged |
| B4 PPI | PPI covers the true human mean; the naive judge-mean does not | ~95% coverage, narrower than human-only | PPI 0.99 @ width 0.061; naive 0.00; human-only 0.98 @ width 0.112 |
| B5 Agent attribution | The decisive step equals the injected fault step | ≥ 90% of tasks | 100% |
| B6 IRT | 2PL parameter recovery; pruning to 50% preserves system ranking | corr ≥ 0.9; Kendall τ ≥ 0.9 | b=0.99, a=0.95, θ=0.98; τ=1.00 |
(B1 - CI coverage, false-regression rate, and power - is verified by simulation tests under
tests/.)
The numbers above are the offline benchmarks: synthetic fakes with a known ground truth,
run in CI on every commit so the statistics are provably correct without spending a token.
Each benchmark also has a live variant that swaps the fake for a real judge/model, so you
can publish the same claims against, say, gpt-4o-mini or claude-haiku.
🚧 Status: partial. The baseline row below is a real live run against Claude Haiku 4.5; B2-B6 live runs are still pending a larger budget (tracked in
PROGRESS.md). Run any of them yourself with the commands underneath.
| Benchmark | Live claim to verify | Judge model | Result |
|---|---|---|---|
| Baseline variance | DeepEval single-sample scores are flaky; repeats + CIs quantify it | claude-haiku-4-5 | clean items → 1.000, flip 0; borderline items → AnswerRelevancy 0.833 [0.633, 1.000], ContextualRelevancy 0.173 [0.000, 0.373] ¹ |
| B2 RAG grounding | Counterfactual Adherence separates grounded vs parametric answers (AUROC ≥ 0.9); Faithfulness cannot | claude-haiku-4-5 | CA AUROC 0.833 (n=12: 6 fictional vs 6 well-known) ³ |
| B3 Judge bias | Real judges show measurable position/verbosity bias with CIs | claude-haiku-4-5 | position −0.292 [−0.422, −0.126] (flagged), verbosity −0.061 (flagged), formatting/authorship n.s. ² |
| B4 PPI | PPI covers the human mean and beats the naive judge-mean at equal labels | claude-haiku-4-5 | PPI 0.510 [0.353, 0.667] covers truth (0.5); eff. n ≈ 41 from 12 labels, CI half of human-only ⁴ |
| B5 Agent attribution | Decisive step = injected fault step in ≥ 90% of tasks with a real model | claude-haiku-4-5 | decisive step = injected fault in 4/5 valid tasks (0.80) ⁶ |
| B6 IRT | On ≥ 5 real systems, pruning 30-50% of items keeps ranking Kendall τ ≥ 0.9 | 6 Claude models | pruned to 50% → Kendall τ 1.000 (systems near ceiling; narrow spread) ⁵ |
¹ Two real live runs against Claude Haiku 4.5. Clean run (n=6, R=5, temp 0,
baseline_live.md): every item a confident 1.000 with zero flip - the reassuring baseline, and an end-to-end pipeline confirmation. Borderline run (n=5, R=5, temp 1.0, deliberately partial/ambiguous items,baseline_live_borderline.md) surfaces what a bare score hides: AnswerRelevancy drops to 0.833 with a wide CI [0.63, 1.00] (a single item's score is genuinely uncertain across the set), ContextualRelevancy correctly falls to 0.17 on the off-topic contexts, and one item wobbled across repeats. Two findings you can only get by measuring: within-item flip is ≈ 0 even at temperature 1, so Haiku 4.5 is a stable judge; but it rated deliberately unsupported claims (invented patent counts, a made-up budget) as fully Faithful = 1.000 - a real judge-leniency signal that causeval's judge audit (bias probes, calibration, PPI) exists to catch.² Live judge-bias probes against Claude Haiku 4.5 (n=12 neutral answer pairs judged in both orders; n=12 pointwise items;
b3_judge_bias_live.md). The headline is real and significant: a position effect of −0.292 [−0.422, −0.126] means that, on answer pairs of equal quality, Haiku picks the second-presented option ~79% of the time - a textbook LLM-judge position bias, and exactly the kind of thing you must correct for before trusting a pairwise judge. It also mildly penalizes padded/verbose answers (−0.061), and shows no significant formatting or authorship-label effect. Unlike the offline B3 (which injects a known 0.15/0.08 bias to prove the probes recover it), the live run measures whatever bias the real judge actually has.³ Live RAG grounding against a real Claude Haiku 4.5 RAG app (
b2_rag_grounding_live.md), using the SPEC's fictional-vs-well-known design. Counterfactual Adherence (does the answer follow a false edit to a supporting fact?) separates the two groups with AUROC 0.833: all 6 fictional items score CA 1.0 with high Context Reliance (the model must use the context), while the well-known items mostly resist the false edit (CA ≈ 0, CR ≈ 0 - the model already knows the answer). It lands below the offline 1.000 / the 0.9 target for an honest reason on real data: on 2 of 12 items the model was swayed by the false context (e.g. it accepted "Romeo and Juliet was written by Dickens") - a real sycophancy signal the metric surfaces. DeepEval Faithfulness would rate every edited-context answer "faithful" and could not make this grounded-vs-parametric distinction at all.⁴ Live PPI (
b4_ppi_live.md) over 40 factual-QA items (20 correct, 20 with a plausible-but-wrong answer) where objective 0/1 correctness is the "human label" and Claude's pointwise score is the predictorf. From a random 12-item labeled subset, PPI estimates the true mean correctness as 0.510 [0.353, 0.667] (truth = 0.500) with an effective sample size ≈ 41 - i.e. 12 human labels bought the precision of ~41, and the CI is half the width of the human-only estimate (0.31 vs 0.58). Honest caveat: on these clear-cut items Claude was a well-calibrated grader (naive judge-mean 0.503, essentially unbiased), so there was little bias to correct here - unlike the offline synthetic judge. PPI's win on this run is label efficiency; its bias-correction matters most on the subtler tasks where judges drift (see the faithfulness leniency in the baseline footnote).⁵ Live IRT (
b6_irt_live.md) using 6 real Claude models as the systems (haiku-4-5, sonnet-4-5, sonnet-5, opus-4-5, opus-4-8, fable-5) on 30 hard short-answer items. Fitting 2PL and pruning to the most-informative 15 items preserved the system ranking exactly (Kendall τ = 1.000). Honest caveat: these are all frontier models, so they cluster near ceiling (accuracy 0.93-1.00) and the true ability spread is narrow - the ranking is close, so preserving it is a lighter test than the offline B6, which validates pruning across a wide simulated ability range. The live run is a real end-to-end confirmation that Fisher-information pruning doesn't scramble the ranking; the offline B6 is the rigorous one.⁶ Live agent attribution (
b5_agent_attribution_live.md) against a real Claude Haiku tool-agent (price → multiply → add-tax tasks) with a wrong tool argument injected at a known step. Counterfactual replay localized the injected fault as the decisive step in 4 of 5 valid tasks (0.80). It lands below the offline 100% for two honest, interesting reasons: (1) real Claude agents often self-correct an injected fault (one task was dropped as "no persistent failure" because the agent noticed and redid the step - a genuine robustness finding), and (2) this environment's Anthropic SDK build rejectstemperature=0, so the replay rollouts are noisy. The offline B5 (a deterministic scripted agent) validates the attribution engine at 100% over 40 tasks.
Reproduce a live run (needs an API key for the provider you name; nothing is hardcoded):
export OPENAI_API_KEY=... # or your provider's key
uv sync
# baseline flakiness with a real judge
uv run python -m causeval.bench.baseline_variance --model gpt-4o-mini
# the full live test suite (opt-in; excluded from the default offline run)
uv run pytest -m live
Live results are written under bench/results/ alongside the offline ones;
open a PR with your table and we'll add it here.
Every command writes a JSON result and prints a human-readable summary; none of them will ever print a bare number.
causeval run --config eval.yaml --out runs/ # repeated-sampling run
causeval compare --baseline a.json --candidate b.json --margin 0.02
causeval gate --baseline a.json --candidate b.json --margin 0.02 # exit 0/1/2
causeval plan --pilot-baseline a.json --pilot-candidate b.json --metric Faithfulness --detect 0.03
causeval ground --config rag.yaml --out runs/ # RAG causal grounding
causeval audit-judge --config judge.yaml --out runs/ # calibration + PPI
causeval attribute --config agent.yaml --out runs/ # agent step-level blame
causeval gate sets the process exit code (0 pass, 1 regression, 2 inconclusive), so it
drops straight into CI as a release gate.
These are enforced by tests, not just documented (see CLAUDE.md):
deepeval>=4.2,<5; never patch its internals.n_items, n_repeats, a CI, the
method, and provenance - including CLI output.Alpha. The full estimator suite (Phases 0-6) and the causal extensions of Phase 7 are
implemented, with 180+ offline tests and strict typing. causeval is a working name.
See SPEC.md for the full design and PROGRESS.md for the
decision log.
uv sync --all-extras
uv run pytest -m "not live" # fast offline suite (must always pass)
uv run pytest -m live # real LLM calls; needs API keys
uv run ruff check . && uv run ruff format --check .
uv run mypy src/causeval # strict
uv run python -m causeval.bench.rag_grounding # regenerate a benchmark report
Contributions are welcome. See CONTRIBUTING.md for setup, the checks your
PR must pass, and the non-negotiable rules (especially: never report a bare score, and every
statistical method needs a simulation test). Use the issue templates to
report a bug or
request a feature.
causeval is created, designed, and maintained by @routsom.
If this project is useful to you, please ⭐ star the repo and follow @routsom for more work on rigorous LLM evaluation. Issues, ideas, and pull requests are welcome.
Apache 2.0, matching DeepEval. Portions of DeepEval, where copied, retain their
Apache 2.0 headers and are recorded in NOTICE.
causeval is an independent project and is not affiliated with or endorsed by Confident AI, the maintainers of DeepEval.
Python
100.0%
Statistically rigorous, causal evaluation for LLM apps on top of DeepEval: confidence intervals, causal interventions (RAG grounding, perturbations, agent attribution), and judge validity (bias audits, calibration, PPI).
Python
0
26 commits
updated Oct 1, 2026
A statistically rigorous, causal evaluation layer for LLM apps, built on top of DeepEval.
Created and maintained by @routsom.
DeepEval measures. It gives you a score. But a single score can't tell you whether it's real (or just noise), why your app produced an output, or whether the LLM judge that produced the score can be trusted.
causeval wraps DeepEval's metrics - they stay the measurement instrument - and adds the three things a score alone can't give you:
The one rule that drives everything: causeval never reports a bare score. Every public result carries an item count, a repeat count, a confidence interval, the method used, and full provenance (versions, seed, dataset hash, git SHA).
DeepEval is an excellent measurement library. causeval is not a competitor - it is the layer that turns DeepEval's measurements into decisions you can defend in a code review, a launch meeting, or a paper. Here is precisely what it adds:
| Gap when you use metrics alone | causeval's answer |
|---|---|
| Single-sample scores, threshold pass/fail, flaky results run to run | Repeated sampling, clustered-bootstrap CIs, three-valued gates (pass / regression / inconclusive) |
| No paired tests, no multiple-comparison control | Paired bootstrap of Δ, Wilcoxon, Holm-adjusted CIs across metrics |
| Faithfulness checks entailment, not dependence on the context | Context Reliance + Counterfactual Adherence: does the answer actually change when you change the evidence? |
| The judge is never itself evaluated | Position/verbosity/formatting bias probes, human calibration, PPI, conformal abstention |
| No robustness or fairness causal tests | Perturbation engine with metamorphic relations + counterfactual fairness gaps |
| Agent failures give you no step-level blame | Counterfactual replay: which step, if fixed, would have saved the run? |
| Wasteful test sets - every item costs the same | 2PL IRT difficulty/discrimination + information-based pruning |
| Shared metric objects corrupt scores under concurrency (issue #3356) | A fresh metric instance per measurement, enforced by a factory + a test |
| Telemetry on by default; dotenv loaded at import | An import guard sets the opt-outs before the first import deepeval |
causeval targets Python 3.10-3.12.
# from source (recommended while in alpha)
git clone https://github.com/routsom/causeval.git
cd causeval
uv sync # core install
uv sync --all-extras # + optional backends (see below)
Optional extras, each pulling in a heavier dependency only where you need it:
| Extra | Enables | Pulls in |
|---|---|---|
causeval[nli] | NLI-based equivalence checks for perturbation invariance | transformers, torch |
causeval[ppi] | numerical cross-check of the PPI estimator | ppi-python |
causeval[irt] | cross-check of the 2PL IRT fit | girth |
causeval[observational] | reference for observational causal estimation | dowhy, econml |
The core estimators (bootstrap CIs, PPI, IRT, AIPW) are implemented natively on numpy/scipy - the extras are for optional cross-checks and NLI, not for the math.
The core object is Experiment: run each metric over each item R times and get back a
RunResult whose estimates carry CIs and variance components.
from causeval import Experiment, MetricSpec
exp = Experiment(
dataset=goldens, # DeepEval Goldens or plain dicts
metrics=[MetricSpec("FaithfulnessMetric", {"threshold": 0.7})],
repeats=5, # resample the judge 5x per item
seed=0,
)
run = exp.run() # or: await exp.a_run()
print(run.summary())
Run 'run' (seed=0)
judge=gpt-4o-mini causeval=0.0.1
Faithfulness [base]: 0.812 [95% CI 0.771, 0.849] (n=50, R=5, icc=0.34, flaky=6)
method=cluster_bootstrap_studentized_B=2000
You immediately learn what a bare score hides: the CI, that ~34% of the variance is between items (the rest is judge noise), and that 6 items are flaky (pass rate between 0.2 and 0.8).
from causeval.stats import compare, gate
comparisons = compare( # Holm-adjusted across metrics
baseline_run.measurements,
candidate_run.measurements,
margin=0.02,
)
report = gate(comparisons)
print(report.overall) # "pass" | "regression" | "inconclusive"
raise SystemExit(report.exit_code) # 0 / 1 / 2 for CI
The gate is three-valued on purpose: "inconclusive" (the CI straddles the margin) is a
first-class outcome, so you never ship a regression that hid behind noise, and never block a
release on a difference you didn't have the power to detect. Plan that power up front with
causeval.stats.plan_power.
DeepEval's Faithfulness tells you the answer is entailed by the context. It cannot tell you the model would have said the same thing without it. causeval intervenes on the context and watches the answer move:
from causeval.interventions import ground, GroundingItem
result = ground(my_rag_app, dataset, repeats=8)
print(result.summary())
RAG grounding (tau=0.5, policy=follow_context)
context_reliance: +0.463 [95% CI +0.311, +0.621] (n=40, R=8)
counterfactual_adherence: +0.506 [95% CI +0.372, +0.643] (n=40, R=8)
classes: grounded=20, parametric=20
Counterfactual Adherence edits one supporting fact to a plausible-but-false value and checks whether the answer follows the edit. Items are labelled grounded / parametric / confabulating / mixed, and any answer that passes Faithfulness while being parametric is flagged false-faithful.
from causeval.judge_audit import calibrate, ppi_mean_ci
report = calibrate(judge_scores, human_labels) # Spearman, kappa, isotonic map, ECE + CI
effect = ppi_mean_ci(human_labels, judge_on_labeled, judge_on_unlabeled) # human mean, debiased
PPI gives you the human-quality mean with a CI that is unbiased (unlike averaging the judge) yet far tighter than using your few human labels alone.
| Module | What you get | Key estimand |
|---|---|---|
stats/ | Repeated-sampling estimates, cluster-bootstrap CIs (studentized / percentile / BCa), variance decomposition (ICC), paired Δ + Wilcoxon + McNemar, Holm, three-valued gates, power planning | μ = E_item[E_repeat[y]] with valid uncertainty |
interventions/rag.py | Context Reliance, Counterfactual Adherence, per-chunk leave-one-out effects, grounded/parametric classification, false-faithful flag | P(answer follows a counterfactual edit) |
interventions/perturb.py | Metamorphic perturbations (paraphrase, reorder, distractor, typo, attribute swap) with invariance rate, paired metric effect, Holm-adjusted fairness gaps | Δ metric under a meaning-preserving change |
interventions/cot.py | Chain-of-thought faithfulness: early-answering curve + area-over-curve, mistake-insertion sensitivity | does the stated reasoning drive the answer? |
judge_audit/ | Bias probes (position/verbosity/formatting/authorship), calibration (isotonic + ECE), sampling plans, PPI/PPI++, conformal abstention (Learn-then-Test), jury + Krippendorff's α | bias effect, E[human score], error rate ≤ α w.p. ≥ 1-δ |
attribution/ | Agent trace schema, record/replay cassettes, step-level counterfactual attribution, decisive-step detection | P(success | do(step_k = oracle)) - P(success | natural) |
stats/irt.py | 2PL item-response theory, Fisher-information pruning, system ranking | item difficulty / discrimination, ability |
stats/observational.py | Doubly-robust AIPW effect from production logs, with DoWhy-style refuters | ATE of a version on a metric |
checks/ | Deterministic pre-judge checks (JSON schema, regex, verifier callables, NLI) + the metamorphic-relation registry | fail fast before spending judge tokens |
core/ | Pydantic result schemas, content-hashed LLM cache (SQLite), cost planner, provenance | reproducibility & cost control |
Every statistical claim has a benchmark with a known ground truth, run in CI on synthetic
data (fakes, no network), with a live variant for real models. These are the committed offline
numbers (bench/results/):
| Benchmark | Claim | Target | Result |
|---|---|---|---|
| B2 RAG grounding | Counterfactual Adherence separates grounded from parametric answers | AUROC ≥ 0.9 | 1.000 |
| B3 Judge bias recovery | Probes recover an injected position (0.15) + verbosity (0.08) bias inside their CIs; non-injected ≈ 0 | ~95% coverage | recovered 0.156 / 0.070; formatting & authorship not flagged |
| B4 PPI | PPI covers the true human mean; the naive judge-mean does not | ~95% coverage, narrower than human-only | PPI 0.99 @ width 0.061; naive 0.00; human-only 0.98 @ width 0.112 |
| B5 Agent attribution | The decisive step equals the injected fault step | ≥ 90% of tasks | 100% |
| B6 IRT | 2PL parameter recovery; pruning to 50% preserves system ranking | corr ≥ 0.9; Kendall τ ≥ 0.9 | b=0.99, a=0.95, θ=0.98; τ=1.00 |
(B1 - CI coverage, false-regression rate, and power - is verified by simulation tests under
tests/.)
The numbers above are the offline benchmarks: synthetic fakes with a known ground truth,
run in CI on every commit so the statistics are provably correct without spending a token.
Each benchmark also has a live variant that swaps the fake for a real judge/model, so you
can publish the same claims against, say, gpt-4o-mini or claude-haiku.
🚧 Status: partial. The baseline row below is a real live run against Claude Haiku 4.5; B2-B6 live runs are still pending a larger budget (tracked in
PROGRESS.md). Run any of them yourself with the commands underneath.
| Benchmark | Live claim to verify | Judge model | Result |
|---|---|---|---|
| Baseline variance | DeepEval single-sample scores are flaky; repeats + CIs quantify it | claude-haiku-4-5 | clean items → 1.000, flip 0; borderline items → AnswerRelevancy 0.833 [0.633, 1.000], ContextualRelevancy 0.173 [0.000, 0.373] ¹ |
| B2 RAG grounding | Counterfactual Adherence separates grounded vs parametric answers (AUROC ≥ 0.9); Faithfulness cannot | claude-haiku-4-5 | CA AUROC 0.833 (n=12: 6 fictional vs 6 well-known) ³ |
| B3 Judge bias | Real judges show measurable position/verbosity bias with CIs | claude-haiku-4-5 | position −0.292 [−0.422, −0.126] (flagged), verbosity −0.061 (flagged), formatting/authorship n.s. ² |
| B4 PPI | PPI covers the human mean and beats the naive judge-mean at equal labels | claude-haiku-4-5 | PPI 0.510 [0.353, 0.667] covers truth (0.5); eff. n ≈ 41 from 12 labels, CI half of human-only ⁴ |
| B5 Agent attribution | Decisive step = injected fault step in ≥ 90% of tasks with a real model | claude-haiku-4-5 | decisive step = injected fault in 4/5 valid tasks (0.80) ⁶ |
| B6 IRT | On ≥ 5 real systems, pruning 30-50% of items keeps ranking Kendall τ ≥ 0.9 | 6 Claude models | pruned to 50% → Kendall τ 1.000 (systems near ceiling; narrow spread) ⁵ |
¹ Two real live runs against Claude Haiku 4.5. Clean run (n=6, R=5, temp 0,
baseline_live.md): every item a confident 1.000 with zero flip - the reassuring baseline, and an end-to-end pipeline confirmation. Borderline run (n=5, R=5, temp 1.0, deliberately partial/ambiguous items,baseline_live_borderline.md) surfaces what a bare score hides: AnswerRelevancy drops to 0.833 with a wide CI [0.63, 1.00] (a single item's score is genuinely uncertain across the set), ContextualRelevancy correctly falls to 0.17 on the off-topic contexts, and one item wobbled across repeats. Two findings you can only get by measuring: within-item flip is ≈ 0 even at temperature 1, so Haiku 4.5 is a stable judge; but it rated deliberately unsupported claims (invented patent counts, a made-up budget) as fully Faithful = 1.000 - a real judge-leniency signal that causeval's judge audit (bias probes, calibration, PPI) exists to catch.² Live judge-bias probes against Claude Haiku 4.5 (n=12 neutral answer pairs judged in both orders; n=12 pointwise items;
b3_judge_bias_live.md). The headline is real and significant: a position effect of −0.292 [−0.422, −0.126] means that, on answer pairs of equal quality, Haiku picks the second-presented option ~79% of the time - a textbook LLM-judge position bias, and exactly the kind of thing you must correct for before trusting a pairwise judge. It also mildly penalizes padded/verbose answers (−0.061), and shows no significant formatting or authorship-label effect. Unlike the offline B3 (which injects a known 0.15/0.08 bias to prove the probes recover it), the live run measures whatever bias the real judge actually has.³ Live RAG grounding against a real Claude Haiku 4.5 RAG app (
b2_rag_grounding_live.md), using the SPEC's fictional-vs-well-known design. Counterfactual Adherence (does the answer follow a false edit to a supporting fact?) separates the two groups with AUROC 0.833: all 6 fictional items score CA 1.0 with high Context Reliance (the model must use the context), while the well-known items mostly resist the false edit (CA ≈ 0, CR ≈ 0 - the model already knows the answer). It lands below the offline 1.000 / the 0.9 target for an honest reason on real data: on 2 of 12 items the model was swayed by the false context (e.g. it accepted "Romeo and Juliet was written by Dickens") - a real sycophancy signal the metric surfaces. DeepEval Faithfulness would rate every edited-context answer "faithful" and could not make this grounded-vs-parametric distinction at all.⁴ Live PPI (
b4_ppi_live.md) over 40 factual-QA items (20 correct, 20 with a plausible-but-wrong answer) where objective 0/1 correctness is the "human label" and Claude's pointwise score is the predictorf. From a random 12-item labeled subset, PPI estimates the true mean correctness as 0.510 [0.353, 0.667] (truth = 0.500) with an effective sample size ≈ 41 - i.e. 12 human labels bought the precision of ~41, and the CI is half the width of the human-only estimate (0.31 vs 0.58). Honest caveat: on these clear-cut items Claude was a well-calibrated grader (naive judge-mean 0.503, essentially unbiased), so there was little bias to correct here - unlike the offline synthetic judge. PPI's win on this run is label efficiency; its bias-correction matters most on the subtler tasks where judges drift (see the faithfulness leniency in the baseline footnote).⁵ Live IRT (
b6_irt_live.md) using 6 real Claude models as the systems (haiku-4-5, sonnet-4-5, sonnet-5, opus-4-5, opus-4-8, fable-5) on 30 hard short-answer items. Fitting 2PL and pruning to the most-informative 15 items preserved the system ranking exactly (Kendall τ = 1.000). Honest caveat: these are all frontier models, so they cluster near ceiling (accuracy 0.93-1.00) and the true ability spread is narrow - the ranking is close, so preserving it is a lighter test than the offline B6, which validates pruning across a wide simulated ability range. The live run is a real end-to-end confirmation that Fisher-information pruning doesn't scramble the ranking; the offline B6 is the rigorous one.⁶ Live agent attribution (
b5_agent_attribution_live.md) against a real Claude Haiku tool-agent (price → multiply → add-tax tasks) with a wrong tool argument injected at a known step. Counterfactual replay localized the injected fault as the decisive step in 4 of 5 valid tasks (0.80). It lands below the offline 100% for two honest, interesting reasons: (1) real Claude agents often self-correct an injected fault (one task was dropped as "no persistent failure" because the agent noticed and redid the step - a genuine robustness finding), and (2) this environment's Anthropic SDK build rejectstemperature=0, so the replay rollouts are noisy. The offline B5 (a deterministic scripted agent) validates the attribution engine at 100% over 40 tasks.
Reproduce a live run (needs an API key for the provider you name; nothing is hardcoded):
export OPENAI_API_KEY=... # or your provider's key
uv sync
# baseline flakiness with a real judge
uv run python -m causeval.bench.baseline_variance --model gpt-4o-mini
# the full live test suite (opt-in; excluded from the default offline run)
uv run pytest -m live
Live results are written under bench/results/ alongside the offline ones;
open a PR with your table and we'll add it here.
Every command writes a JSON result and prints a human-readable summary; none of them will ever print a bare number.
causeval run --config eval.yaml --out runs/ # repeated-sampling run
causeval compare --baseline a.json --candidate b.json --margin 0.02
causeval gate --baseline a.json --candidate b.json --margin 0.02 # exit 0/1/2
causeval plan --pilot-baseline a.json --pilot-candidate b.json --metric Faithfulness --detect 0.03
causeval ground --config rag.yaml --out runs/ # RAG causal grounding
causeval audit-judge --config judge.yaml --out runs/ # calibration + PPI
causeval attribute --config agent.yaml --out runs/ # agent step-level blame
causeval gate sets the process exit code (0 pass, 1 regression, 2 inconclusive), so it
drops straight into CI as a release gate.
These are enforced by tests, not just documented (see CLAUDE.md):
deepeval>=4.2,<5; never patch its internals.n_items, n_repeats, a CI, the
method, and provenance - including CLI output.Alpha. The full estimator suite (Phases 0-6) and the causal extensions of Phase 7 are
implemented, with 180+ offline tests and strict typing. causeval is a working name.
See SPEC.md for the full design and PROGRESS.md for the
decision log.
uv sync --all-extras
uv run pytest -m "not live" # fast offline suite (must always pass)
uv run pytest -m live # real LLM calls; needs API keys
uv run ruff check . && uv run ruff format --check .
uv run mypy src/causeval # strict
uv run python -m causeval.bench.rag_grounding # regenerate a benchmark report
Contributions are welcome. See CONTRIBUTING.md for setup, the checks your
PR must pass, and the non-negotiable rules (especially: never report a bare score, and every
statistical method needs a simulation test). Use the issue templates to
report a bug or
request a feature.
causeval is created, designed, and maintained by @routsom.
If this project is useful to you, please ⭐ star the repo and follow @routsom for more work on rigorous LLM evaluation. Issues, ideas, and pull requests are welcome.
Apache 2.0, matching DeepEval. Portions of DeepEval, where copied, retain their
Apache 2.0 headers and are recorded in NOTICE.
causeval is an independent project and is not affiliated with or endorsed by Confident AI, the maintainers of DeepEval.
Python
100.0%