Semantic loop detection for agent traces: a judged drop-in for deepeval's word-overlap stagnation check
Python
0
5 commits
updated Sep 27, 2026
Semantic loop detection for agent traces. A drop-in alternative to the
stagnation sub-signal in deepeval's AgentLoopDetectionMetric, which asks
"do these two steps share enough words?" This asks the question that was
actually meant: did the agent get anywhere?
from semloop import Trace, SemanticLoopDetection, JevBackend
trace = Trace.from_deepeval_trace_dict(test_case._trace_dict, goal=test_case.input)
metric = SemanticLoopDetection(JevBackend()) # needs TYPESAFE_API_KEY
result = metric.measure(trace)
print(result.score, result.reason)
JevBackend and LLMJudgeBackend both implement the same tiny interface —
ask(state, questions) -> JudgeResponse — so the same detector runs against
TypeSafe's Jev, any OpenAI-compatible chat model, or a scripted MockBackend
in tests.
deepeval/metrics/agent_loop_detection/agent_loop_detection.py (read
2026-09-27) scores reasoning stagnation with bigram Jaccard overlap and
difflib.SequenceMatcher on consecutive LLM outputs. Its own docstring names
the limitation:
Stagnation detection uses structural text similarity... It will miss semantically identical outputs that are worded very differently. An LLM-as-judge mode would solve this but would sacrifice the deterministic / zero-cost properties.
Jev (TypeSafe AI's "System One" model, $0.042 per million input tokens, output free, ~70–500ms per call — see the launch post) answers exactly this shape of question — "is this true, yes or no?" — for a fraction of a cent per trace, which makes the zero-cost tradeoff in that docstring no longer the right one to make.
bench/dataset.jsonl (60 hand-built traces, 10 families, 6 instances each —
generated by bench/make_dataset.py, seed fixed for reproducibility) isolates
five documented blind spots in the lexical check, plus four healthy runs it
should leave alone. Running the actual port of deepeval's scoring
(semloop/lexical.py) against it:
$ python bench/run.py
Stagnation signal, correct calls per family (like-for-like: reasoning check only)
family lexical
--------------------------------
alternating_loop 0/6 <- A-B-A-B cycle: only consecutive pairs are compared
boilerplate_progress 0/6 <- false positive: identical preamble, real work underneath
cosmetic_args_loop 0/6 <- not a reasoning defect; see the tool-repetition table below
identical_tool_loop 6/6 <- what the lexical check is already good at
paraphrased_loop 0/6 <- same intent, reworded every step
progress 6/6 <- healthy, correctly left alone
retry_then_recover 6/6 <- healthy, correctly left alone
template_progress 0/6 <- false positive: 5 different files, same sentence shape
terse_loop 0/6 <- steps under 20 meaningful words are skipped, not compared
verbatim_loop 6/6 <- what the lexical check is already good at
Totals: stagnation 24/60 (40%) · metric verdict 36/60 (60%)
Six of the ten families score 0/6. Two of those six (identical_tool_loop,
verbatim_loop) are cases the lexical check is supposed to get right —
they're here as a regression check on any replacement, not as a complaint.
An oracle run — SemanticLoopDetection wired to a backend that answers
each question the way a judge that actually read the trace should (see
"Verifying this without an API key" below) — gets 54/60 (90%):
$ PYTHONPATH=bench python bench/run.py --judge oracle
family lexical oracle
------------------------------------------------------
alternating_loop 0/6 6/6
boilerplate_progress 0/6 6/6
cosmetic_args_loop 0/6 0/6 <- see "A weighting quirk" below
identical_tool_loop 6/6 6/6
paraphrased_loop 0/6 6/6
progress 6/6 6/6
retry_then_recover 6/6 6/6
template_progress 0/6 6/6
terse_loop 0/6 6/6
verbatim_loop 6/6 6/6
Totals: lexical 24/60 (40%) · oracle 54/60 (90%)
The oracle is not a benchmark of Jev, or of any real judge. It scripts
answers from the dataset's own family labels — see
bench/oracle_backend.py's docstring. What it proves is that the mechanism
(question construction, answer parsing, score combination, one request per
trace) reaches the right verdict when its questions are answered correctly.
Whether a live Jev or LLM judge actually answers them correctly on your traces
is a separate, unavoidable question — see "What this doesn't prove" below.
The question above has an answer now, not just an oracle's stand-in for one.
bench/run.py --judge jev was run once against typesafe/jev-1.13 through a
host that serves the System One protocol at its own path rather than
TypeSafe's /v1/systemone — JevBackend now takes a path= override for
exactly this (--jev-path "" on the CLI, when --base-url is already the
full endpoint):
$ python bench/run.py --judge jev --model typesafe/jev-1.13 \
--base-url <endpoint> --jev-path ""
family lexical jev:typesafe/jev-1.13
-------------------------------------------------------------
alternating_loop 0/6 6/6
boilerplate_progress 0/6 6/6
cosmetic_args_loop 0/6 6/6
identical_tool_loop 6/6 6/6
paraphrased_loop 0/6 6/6
progress 6/6 6/6
retry_then_recover 6/6 6/6
template_progress 0/6 0/6 <- see below
terse_loop 0/6 6/6
verbatim_loop 6/6 6/6
Totals: stagnation — lexical 24/60 (40%) · jev 54/60 (90%)
verdict — lexical 36/60 (60%) · jev 60/60 (100%)
$0.0042 total (99,817 input tokens) · 779 ms mean latency · 7 undecided · 0 errors
This is not the oracle: every answer came from the model actually reading
each trace's content, with no knowledge of which family generated it. The
stagnation number lands almost exactly on the oracle's 90% — the mechanism
and a real judge agree, which is the result you want from this kind of check.
7 of the 60 traces fell in the 0.4–0.6 undecided band and were resolved by
fallback_to_lexical_when_uncertain; none errored.
One place the real judge disagrees with the oracle's script.
template_progress (five steps, each touching a different file, described in
the same sentence shape) fools the real judge on the fine-grained stagnation
question exactly the way it fools the lexical check — 0/6, where the oracle
had assumed a competent judge would score 6/6. It doesn't flip the overall
verdict here only because the 0.35 weight on reasoning stagnation isn't
enough on its own to fail the trace (see "A weighting quirk" below) — the
false-positive blind spot this package claims to close is, for this model, at
least partly still open, just hidden by the weighting rather than fixed. A
caller who raises the stagnation weight, or reads score_breakdown directly
instead of the combined score, would still see it.
This is one run, on one synthetic dataset, through one specific host — not proof this holds on production traffic, and not a claim about TypeSafe's own endpoint or any other model. It is a real data point where the section above could only offer a scripted one.
SequenceMatcher both stay below
0.47 on ten different ways to say "I'm going to search the codebase" (see
bench/make_dataset.py's PARAPHRASES). Fixed by asking directly: did
step j advance the goal beyond step i (progress_{i}_{j}, one noul
question per consecutive pair).if len(words) < 20: continue —
a loop of short retries is invisible to the lexical check by construction.
build_questions has no length filter; every consecutive pair gets a
question regardless of how short the steps are.revisit_{i}_{j} questions at lookahead distance ask
directly whether step j repeats step i, independent of what sits between
them.(name, sorted(args.items())); adding one flag resets the count. Closed
with a fourth question family, tool_repeat_{name}, for any tool called at
least repetition_threshold times: are these calls the same underlying
action despite different arguments? The more severe of the exact-match and
judged checks wins (min in SemanticLoopDetection.measure), so this can
only add detections, never hide one the lexical check already made.SequenceMatcher
past 0.85 (bench/make_dataset.py's template_progress and
boilerplate_progress — see the measured pairwise similarities in that
file's docstrings). Because the judged question is did this advance the
goal, not do these share words, genuine progress across different
files or underneath a fixed preamble reads as what it is.Every question about one trace — however many pairs, however many repeated
tool names — goes in a single request (test_one_request_per_trace_regardless_of_pair_count
in tests/test_semantic.py), which is what keeps the added cost near zero
even on a long trace.
deepeval's _combine_scores weights the three sub-signals 0.40 (tool
repetition) / 0.35 (reasoning stagnation) / 0.25 (call graph cycles), and the
default threshold is 0.5. That means no single sub-signal, however severe,
can flip the overall verdict alone: a fully-confident detection (sub-score
0.0) on the heaviest-weighted signal still leaves the composite at
1 - 0.40 = 0.60, above threshold. cosmetic_args_loop's oracle row shows
this directly — the tool-repetition sub-score hits 0.0 with a confidently
correct judge, and the overall verdict still doesn't flip until the caller
raises the threshold above 0.60 (tests/test_semantic.py's
test_a_lone_signal_needs_a_stricter_threshold_to_flip_the_verdict, and
tests/test_lexical.py's test_single_severe_signal_cannot_flip_the_default_verdict).
This applies identically whether the sub-signal comes from the lexical check
or a judge — it's a property of the combination step this package keeps
unchanged for compatibility, not a limitation of the semantic swap. A caller
who wants a single confirmed signal to be decisive should raise the metric's
threshold, or read score_breakdown directly instead of relying on the
combined verdict.
The dataset is hand-built from templates, not sampled from production traces.
It's enough to show each blind spot exists and is reproducible — it says
nothing about how often paraphrased loops, cosmetic argument changes, or false
positives on genuine multi-file progress actually occur in your traffic.
"A real judge run" above answers whether a live judge answers these
particular questions correctly on this dataset — mostly yes, with one
blind spot ( template_progress) still open on the model tested — but that
was one model, through one host, on one synthetic dataset. Run
bench/run.py --judge jev (needs TYPESAFE_API_KEY, or pass --base-url
and --jev-path "" for a host that serves the protocol somewhere other than
/v1/systemone) or --judge llm --model <your model> (needs an API key for
that model, plus --input-usd-per-mtok and --output-usd-per-mtok if you
want a real cost column) on your own traces before trusting this on
production traffic. The --judge columns report cost, latency, and an
undecided count (answers in the 0.4–0.6 band that fell back to the lexical
check — see fallback_to_lexical_when_uncertain) alongside accuracy,
specifically so a live run can be judged on more than a single number.
pip install -e ".[dev]"
python bench/make_dataset.py # writes bench/dataset.jsonl
pytest # 29 tests, all offline, MockBackend only
python bench/run.py # lexical baseline, no key needed
python bench/run.py --judge oracle # needs PYTHONPATH=bench (see below)
tests/test_semantic.py scripts a MockBackend explicitly for each blind
spot (never relying on a shared default answer across differently-worded
questions — an early draft of this benchmark got that wrong; see the git
history) and checks the detector reaches the right sub-score. tests/test_backends.py
checks the Jev and LLM-judge response parsers against fixed payloads, so
they're covered without a network call. tests/test_lexical.py pins the
lexical port's behaviour against the exact deepeval source it's copied from,
including the weighting quirk above, so a change to either can't silently
drift from what's documented here.
semloop/
trace.py Span/Trace model; adapts deepeval's nested trace dict
lexical.py Faithful port of deepeval's current scoring, for the baseline
semantic.py SemanticLoopDetection: the judged replacement
backends/
jev.py TypeSafe System One API client (stdlib only, no deps;
path= overridable for hosts that don't serve it at
/v1/systemone, e.g. an OpenRouter proxy)
llm.py Any OpenAI-compatible chat model as a judge, for comparison
__init__.py Answer/Backend/JudgeResponse types, plus MockBackend
bench/
make_dataset.py builds bench/dataset.jsonl from the 10 families
run.py runs lexical / oracle / jev / llm and reports accuracy, cost, latency
oracle_backend.py the scripted stand-in described above
tests/ 29 offline tests (pytest)
v0.1.0. Built from a read of deepeval's agent_loop_detection.py on
2026-09-27; the deepeval maintainers have their own reasons for keeping the
metric deterministic and zero-cost by default (their own docstring names the
tradeoff), so this is offered as a standalone package and a measured case for
an optional judged mode — not yet proposed upstream. Validated once against a
real judge the same day ("A real judge run" above); not yet run against
production traffic.
Python
100.0%
Semantic loop detection for agent traces: a judged drop-in for deepeval's word-overlap stagnation check
Python
0
5 commits
updated Sep 27, 2026
Semantic loop detection for agent traces. A drop-in alternative to the
stagnation sub-signal in deepeval's AgentLoopDetectionMetric, which asks
"do these two steps share enough words?" This asks the question that was
actually meant: did the agent get anywhere?
from semloop import Trace, SemanticLoopDetection, JevBackend
trace = Trace.from_deepeval_trace_dict(test_case._trace_dict, goal=test_case.input)
metric = SemanticLoopDetection(JevBackend()) # needs TYPESAFE_API_KEY
result = metric.measure(trace)
print(result.score, result.reason)
JevBackend and LLMJudgeBackend both implement the same tiny interface —
ask(state, questions) -> JudgeResponse — so the same detector runs against
TypeSafe's Jev, any OpenAI-compatible chat model, or a scripted MockBackend
in tests.
deepeval/metrics/agent_loop_detection/agent_loop_detection.py (read
2026-09-27) scores reasoning stagnation with bigram Jaccard overlap and
difflib.SequenceMatcher on consecutive LLM outputs. Its own docstring names
the limitation:
Stagnation detection uses structural text similarity... It will miss semantically identical outputs that are worded very differently. An LLM-as-judge mode would solve this but would sacrifice the deterministic / zero-cost properties.
Jev (TypeSafe AI's "System One" model, $0.042 per million input tokens, output free, ~70–500ms per call — see the launch post) answers exactly this shape of question — "is this true, yes or no?" — for a fraction of a cent per trace, which makes the zero-cost tradeoff in that docstring no longer the right one to make.
bench/dataset.jsonl (60 hand-built traces, 10 families, 6 instances each —
generated by bench/make_dataset.py, seed fixed for reproducibility) isolates
five documented blind spots in the lexical check, plus four healthy runs it
should leave alone. Running the actual port of deepeval's scoring
(semloop/lexical.py) against it:
$ python bench/run.py
Stagnation signal, correct calls per family (like-for-like: reasoning check only)
family lexical
--------------------------------
alternating_loop 0/6 <- A-B-A-B cycle: only consecutive pairs are compared
boilerplate_progress 0/6 <- false positive: identical preamble, real work underneath
cosmetic_args_loop 0/6 <- not a reasoning defect; see the tool-repetition table below
identical_tool_loop 6/6 <- what the lexical check is already good at
paraphrased_loop 0/6 <- same intent, reworded every step
progress 6/6 <- healthy, correctly left alone
retry_then_recover 6/6 <- healthy, correctly left alone
template_progress 0/6 <- false positive: 5 different files, same sentence shape
terse_loop 0/6 <- steps under 20 meaningful words are skipped, not compared
verbatim_loop 6/6 <- what the lexical check is already good at
Totals: stagnation 24/60 (40%) · metric verdict 36/60 (60%)
Six of the ten families score 0/6. Two of those six (identical_tool_loop,
verbatim_loop) are cases the lexical check is supposed to get right —
they're here as a regression check on any replacement, not as a complaint.
An oracle run — SemanticLoopDetection wired to a backend that answers
each question the way a judge that actually read the trace should (see
"Verifying this without an API key" below) — gets 54/60 (90%):
$ PYTHONPATH=bench python bench/run.py --judge oracle
family lexical oracle
------------------------------------------------------
alternating_loop 0/6 6/6
boilerplate_progress 0/6 6/6
cosmetic_args_loop 0/6 0/6 <- see "A weighting quirk" below
identical_tool_loop 6/6 6/6
paraphrased_loop 0/6 6/6
progress 6/6 6/6
retry_then_recover 6/6 6/6
template_progress 0/6 6/6
terse_loop 0/6 6/6
verbatim_loop 6/6 6/6
Totals: lexical 24/60 (40%) · oracle 54/60 (90%)
The oracle is not a benchmark of Jev, or of any real judge. It scripts
answers from the dataset's own family labels — see
bench/oracle_backend.py's docstring. What it proves is that the mechanism
(question construction, answer parsing, score combination, one request per
trace) reaches the right verdict when its questions are answered correctly.
Whether a live Jev or LLM judge actually answers them correctly on your traces
is a separate, unavoidable question — see "What this doesn't prove" below.
The question above has an answer now, not just an oracle's stand-in for one.
bench/run.py --judge jev was run once against typesafe/jev-1.13 through a
host that serves the System One protocol at its own path rather than
TypeSafe's /v1/systemone — JevBackend now takes a path= override for
exactly this (--jev-path "" on the CLI, when --base-url is already the
full endpoint):
$ python bench/run.py --judge jev --model typesafe/jev-1.13 \
--base-url <endpoint> --jev-path ""
family lexical jev:typesafe/jev-1.13
-------------------------------------------------------------
alternating_loop 0/6 6/6
boilerplate_progress 0/6 6/6
cosmetic_args_loop 0/6 6/6
identical_tool_loop 6/6 6/6
paraphrased_loop 0/6 6/6
progress 6/6 6/6
retry_then_recover 6/6 6/6
template_progress 0/6 0/6 <- see below
terse_loop 0/6 6/6
verbatim_loop 6/6 6/6
Totals: stagnation — lexical 24/60 (40%) · jev 54/60 (90%)
verdict — lexical 36/60 (60%) · jev 60/60 (100%)
$0.0042 total (99,817 input tokens) · 779 ms mean latency · 7 undecided · 0 errors
This is not the oracle: every answer came from the model actually reading
each trace's content, with no knowledge of which family generated it. The
stagnation number lands almost exactly on the oracle's 90% — the mechanism
and a real judge agree, which is the result you want from this kind of check.
7 of the 60 traces fell in the 0.4–0.6 undecided band and were resolved by
fallback_to_lexical_when_uncertain; none errored.
One place the real judge disagrees with the oracle's script.
template_progress (five steps, each touching a different file, described in
the same sentence shape) fools the real judge on the fine-grained stagnation
question exactly the way it fools the lexical check — 0/6, where the oracle
had assumed a competent judge would score 6/6. It doesn't flip the overall
verdict here only because the 0.35 weight on reasoning stagnation isn't
enough on its own to fail the trace (see "A weighting quirk" below) — the
false-positive blind spot this package claims to close is, for this model, at
least partly still open, just hidden by the weighting rather than fixed. A
caller who raises the stagnation weight, or reads score_breakdown directly
instead of the combined score, would still see it.
This is one run, on one synthetic dataset, through one specific host — not proof this holds on production traffic, and not a claim about TypeSafe's own endpoint or any other model. It is a real data point where the section above could only offer a scripted one.
SequenceMatcher both stay below
0.47 on ten different ways to say "I'm going to search the codebase" (see
bench/make_dataset.py's PARAPHRASES). Fixed by asking directly: did
step j advance the goal beyond step i (progress_{i}_{j}, one noul
question per consecutive pair).if len(words) < 20: continue —
a loop of short retries is invisible to the lexical check by construction.
build_questions has no length filter; every consecutive pair gets a
question regardless of how short the steps are.revisit_{i}_{j} questions at lookahead distance ask
directly whether step j repeats step i, independent of what sits between
them.(name, sorted(args.items())); adding one flag resets the count. Closed
with a fourth question family, tool_repeat_{name}, for any tool called at
least repetition_threshold times: are these calls the same underlying
action despite different arguments? The more severe of the exact-match and
judged checks wins (min in SemanticLoopDetection.measure), so this can
only add detections, never hide one the lexical check already made.SequenceMatcher
past 0.85 (bench/make_dataset.py's template_progress and
boilerplate_progress — see the measured pairwise similarities in that
file's docstrings). Because the judged question is did this advance the
goal, not do these share words, genuine progress across different
files or underneath a fixed preamble reads as what it is.Every question about one trace — however many pairs, however many repeated
tool names — goes in a single request (test_one_request_per_trace_regardless_of_pair_count
in tests/test_semantic.py), which is what keeps the added cost near zero
even on a long trace.
deepeval's _combine_scores weights the three sub-signals 0.40 (tool
repetition) / 0.35 (reasoning stagnation) / 0.25 (call graph cycles), and the
default threshold is 0.5. That means no single sub-signal, however severe,
can flip the overall verdict alone: a fully-confident detection (sub-score
0.0) on the heaviest-weighted signal still leaves the composite at
1 - 0.40 = 0.60, above threshold. cosmetic_args_loop's oracle row shows
this directly — the tool-repetition sub-score hits 0.0 with a confidently
correct judge, and the overall verdict still doesn't flip until the caller
raises the threshold above 0.60 (tests/test_semantic.py's
test_a_lone_signal_needs_a_stricter_threshold_to_flip_the_verdict, and
tests/test_lexical.py's test_single_severe_signal_cannot_flip_the_default_verdict).
This applies identically whether the sub-signal comes from the lexical check
or a judge — it's a property of the combination step this package keeps
unchanged for compatibility, not a limitation of the semantic swap. A caller
who wants a single confirmed signal to be decisive should raise the metric's
threshold, or read score_breakdown directly instead of relying on the
combined verdict.
The dataset is hand-built from templates, not sampled from production traces.
It's enough to show each blind spot exists and is reproducible — it says
nothing about how often paraphrased loops, cosmetic argument changes, or false
positives on genuine multi-file progress actually occur in your traffic.
"A real judge run" above answers whether a live judge answers these
particular questions correctly on this dataset — mostly yes, with one
blind spot ( template_progress) still open on the model tested — but that
was one model, through one host, on one synthetic dataset. Run
bench/run.py --judge jev (needs TYPESAFE_API_KEY, or pass --base-url
and --jev-path "" for a host that serves the protocol somewhere other than
/v1/systemone) or --judge llm --model <your model> (needs an API key for
that model, plus --input-usd-per-mtok and --output-usd-per-mtok if you
want a real cost column) on your own traces before trusting this on
production traffic. The --judge columns report cost, latency, and an
undecided count (answers in the 0.4–0.6 band that fell back to the lexical
check — see fallback_to_lexical_when_uncertain) alongside accuracy,
specifically so a live run can be judged on more than a single number.
pip install -e ".[dev]"
python bench/make_dataset.py # writes bench/dataset.jsonl
pytest # 29 tests, all offline, MockBackend only
python bench/run.py # lexical baseline, no key needed
python bench/run.py --judge oracle # needs PYTHONPATH=bench (see below)
tests/test_semantic.py scripts a MockBackend explicitly for each blind
spot (never relying on a shared default answer across differently-worded
questions — an early draft of this benchmark got that wrong; see the git
history) and checks the detector reaches the right sub-score. tests/test_backends.py
checks the Jev and LLM-judge response parsers against fixed payloads, so
they're covered without a network call. tests/test_lexical.py pins the
lexical port's behaviour against the exact deepeval source it's copied from,
including the weighting quirk above, so a change to either can't silently
drift from what's documented here.
semloop/
trace.py Span/Trace model; adapts deepeval's nested trace dict
lexical.py Faithful port of deepeval's current scoring, for the baseline
semantic.py SemanticLoopDetection: the judged replacement
backends/
jev.py TypeSafe System One API client (stdlib only, no deps;
path= overridable for hosts that don't serve it at
/v1/systemone, e.g. an OpenRouter proxy)
llm.py Any OpenAI-compatible chat model as a judge, for comparison
__init__.py Answer/Backend/JudgeResponse types, plus MockBackend
bench/
make_dataset.py builds bench/dataset.jsonl from the 10 families
run.py runs lexical / oracle / jev / llm and reports accuracy, cost, latency
oracle_backend.py the scripted stand-in described above
tests/ 29 offline tests (pytest)
v0.1.0. Built from a read of deepeval's agent_loop_detection.py on
2026-09-27; the deepeval maintainers have their own reasons for keeping the
metric deterministic and zero-cost by default (their own docstring names the
tradeoff), so this is offered as a standalone package and a measured case for
an optional judged mode — not yet proposed upstream. Validated once against a
real judge the same day ("A real judge run" above); not yet run against
production traffic.
Python
100.0%