VBS2004/semloop

Semantic loop detection for agent traces: a judged drop-in for deepeval's word-overlap stagnation check

Python

0

5 commits

updated Sep 27, 2026

See the code

See what people are saying

README

semloop

Semantic loop detection for agent traces. A drop-in alternative to the stagnation sub-signal in deepeval's AgentLoopDetectionMetric, which asks "do these two steps share enough words?" This asks the question that was actually meant: did the agent get anywhere?

from semloop import Trace, SemanticLoopDetection, JevBackend

trace = Trace.from_deepeval_trace_dict(test_case._trace_dict, goal=test_case.input)
metric = SemanticLoopDetection(JevBackend())  # needs TYPESAFE_API_KEY
result = metric.measure(trace)
print(result.score, result.reason)

JevBackend and LLMJudgeBackend both implement the same tiny interface — ask(state, questions) -> JudgeResponse — so the same detector runs against TypeSafe's Jev, any OpenAI-compatible chat model, or a scripted MockBackend in tests.

Why

deepeval/metrics/agent_loop_detection/agent_loop_detection.py (read 2026-09-27) scores reasoning stagnation with bigram Jaccard overlap and difflib.SequenceMatcher on consecutive LLM outputs. Its own docstring names the limitation:

Stagnation detection uses structural text similarity... It will miss semantically identical outputs that are worded very differently. An LLM-as-judge mode would solve this but would sacrifice the deterministic / zero-cost properties.

Jev (TypeSafe AI's "System One" model, $0.042 per million input tokens, output free, ~70–500ms per call — see the launch post) answers exactly this shape of question — "is this true, yes or no?" — for a fraction of a cent per trace, which makes the zero-cost tradeoff in that docstring no longer the right one to make.

What it fixes, with numbers

bench/dataset.jsonl (60 hand-built traces, 10 families, 6 instances each — generated by bench/make_dataset.py, seed fixed for reproducibility) isolates five documented blind spots in the lexical check, plus four healthy runs it should leave alone. Running the actual port of deepeval's scoring (semloop/lexical.py) against it:

$ python bench/run.py
Stagnation signal, correct calls per family (like-for-like: reasoning check only)
family                    lexical
--------------------------------
alternating_loop            0/6      <- A-B-A-B cycle: only consecutive pairs are compared
boilerplate_progress        0/6      <- false positive: identical preamble, real work underneath
cosmetic_args_loop          0/6      <- not a reasoning defect; see the tool-repetition table below
identical_tool_loop         6/6      <- what the lexical check is already good at
paraphrased_loop            0/6      <- same intent, reworded every step
progress                    6/6      <- healthy, correctly left alone
retry_then_recover          6/6      <- healthy, correctly left alone
template_progress           0/6      <- false positive: 5 different files, same sentence shape
terse_loop                  0/6      <- steps under 20 meaningful words are skipped, not compared
verbatim_loop                6/6      <- what the lexical check is already good at

Totals: stagnation 24/60 (40%) · metric verdict 36/60 (60%)

Six of the ten families score 0/6. Two of those six (identical_tool_loop, verbatim_loop) are cases the lexical check is supposed to get right — they're here as a regression check on any replacement, not as a complaint.

An oracle run — SemanticLoopDetection wired to a backend that answers each question the way a judge that actually read the trace should (see "Verifying this without an API key" below) — gets 54/60 (90%):

$ PYTHONPATH=bench python bench/run.py --judge oracle
family                    lexical          oracle
------------------------------------------------------
alternating_loop            0/6             6/6
boilerplate_progress        0/6             6/6
cosmetic_args_loop          0/6             0/6   <- see "A weighting quirk" below
identical_tool_loop         6/6             6/6
paraphrased_loop            0/6             6/6
progress                    6/6             6/6
retry_then_recover          6/6             6/6
template_progress           0/6             6/6
terse_loop                  0/6             6/6
verbatim_loop                6/6             6/6

Totals: lexical 24/60 (40%) · oracle 54/60 (90%)

The oracle is not a benchmark of Jev, or of any real judge. It scripts answers from the dataset's own family labels — see bench/oracle_backend.py's docstring. What it proves is that the mechanism (question construction, answer parsing, score combination, one request per trace) reaches the right verdict when its questions are answered correctly. Whether a live Jev or LLM judge actually answers them correctly on your traces is a separate, unavoidable question — see "What this doesn't prove" below.

A real judge run (2026-09-27)

The question above has an answer now, not just an oracle's stand-in for one. bench/run.py --judge jev was run once against typesafe/jev-1.13 through a host that serves the System One protocol at its own path rather than TypeSafe's /v1/systemone — JevBackend now takes a path= override for exactly this (--jev-path "" on the CLI, when --base-url is already the full endpoint):

$ python bench/run.py --judge jev --model typesafe/jev-1.13 \
    --base-url <endpoint> --jev-path ""
family                    lexical     jev:typesafe/jev-1.13
-------------------------------------------------------------
alternating_loop            0/6             6/6
boilerplate_progress        0/6             6/6
cosmetic_args_loop          0/6             6/6
identical_tool_loop         6/6             6/6
paraphrased_loop            0/6             6/6
progress                    6/6             6/6
retry_then_recover          6/6             6/6
template_progress           0/6             0/6   <- see below
terse_loop                  0/6             6/6
verbatim_loop                6/6             6/6

Totals: stagnation — lexical 24/60 (40%) · jev 54/60 (90%)
        verdict    — lexical 36/60 (60%) · jev 60/60 (100%)
        $0.0042 total (99,817 input tokens) · 779 ms mean latency · 7 undecided · 0 errors

This is not the oracle: every answer came from the model actually reading each trace's content, with no knowledge of which family generated it. The stagnation number lands almost exactly on the oracle's 90% — the mechanism and a real judge agree, which is the result you want from this kind of check. 7 of the 60 traces fell in the 0.4–0.6 undecided band and were resolved by fallback_to_lexical_when_uncertain; none errored.

One place the real judge disagrees with the oracle's script. template_progress (five steps, each touching a different file, described in the same sentence shape) fools the real judge on the fine-grained stagnation question exactly the way it fools the lexical check — 0/6, where the oracle had assumed a competent judge would score 6/6. It doesn't flip the overall verdict here only because the 0.35 weight on reasoning stagnation isn't enough on its own to fail the trace (see "A weighting quirk" below) — the false-positive blind spot this package claims to close is, for this model, at least partly still open, just hidden by the weighting rather than fixed. A caller who raises the stagnation weight, or reads score_breakdown directly instead of the combined score, would still see it.

This is one run, on one synthetic dataset, through one specific host — not proof this holds on production traffic, and not a claim about TypeSafe's own endpoint or any other model. It is a real data point where the section above could only offer a scripted one.

The five blind spots, and how each is closed

  1. Paraphrased loops. Bigram Jaccard and SequenceMatcher both stay below 0.47 on ten different ways to say "I'm going to search the codebase" (see bench/make_dataset.py's PARAPHRASES). Fixed by asking directly: did step j advance the goal beyond step i (progress_{i}_{j}, one noul question per consecutive pair).
  2. Terse steps are skipped, not compared. if len(words) < 20: continue — a loop of short retries is invisible to the lexical check by construction. build_questions has no length filter; every consecutive pair gets a question regardless of how short the steps are.
  3. A-B-A-B loops. Only consecutive pairs are compared, so a period-2 cycle never trips it. revisit_{i}_{j} questions at lookahead distance ask directly whether step j repeats step i, independent of what sits between them.
  4. Cosmetic argument changes reset tool repetition. The exact-match key is (name, sorted(args.items())); adding one flag resets the count. Closed with a fourth question family, tool_repeat_{name}, for any tool called at least repetition_threshold times: are these calls the same underlying action despite different arguments? The more severe of the exact-match and judged checks wins (min in SemanticLoopDetection.measure), so this can only add detections, never hide one the lexical check already made.
  5. False positives on real progress. Five steps that each process a different file, described in the same sentence shape, or a fixed corporate preamble in front of real one-line updates, both push SequenceMatcher past 0.85 (bench/make_dataset.py's template_progress and boilerplate_progress — see the measured pairwise similarities in that file's docstrings). Because the judged question is did this advance the goal, not do these share words, genuine progress across different files or underneath a fixed preamble reads as what it is.

Every question about one trace — however many pairs, however many repeated tool names — goes in a single request (test_one_request_per_trace_regardless_of_pair_count in tests/test_semantic.py), which is what keeps the added cost near zero even on a long trace.

A weighting quirk worth knowing about (not something this package changes)

deepeval's _combine_scores weights the three sub-signals 0.40 (tool repetition) / 0.35 (reasoning stagnation) / 0.25 (call graph cycles), and the default threshold is 0.5. That means no single sub-signal, however severe, can flip the overall verdict alone: a fully-confident detection (sub-score 0.0) on the heaviest-weighted signal still leaves the composite at 1 - 0.40 = 0.60, above threshold. cosmetic_args_loop's oracle row shows this directly — the tool-repetition sub-score hits 0.0 with a confidently correct judge, and the overall verdict still doesn't flip until the caller raises the threshold above 0.60 (tests/test_semantic.py's test_a_lone_signal_needs_a_stricter_threshold_to_flip_the_verdict, and tests/test_lexical.py's test_single_severe_signal_cannot_flip_the_default_verdict). This applies identically whether the sub-signal comes from the lexical check or a judge — it's a property of the combination step this package keeps unchanged for compatibility, not a limitation of the semantic swap. A caller who wants a single confirmed signal to be decisive should raise the metric's threshold, or read score_breakdown directly instead of relying on the combined verdict.

What this doesn't prove

The dataset is hand-built from templates, not sampled from production traces. It's enough to show each blind spot exists and is reproducible — it says nothing about how often paraphrased loops, cosmetic argument changes, or false positives on genuine multi-file progress actually occur in your traffic. "A real judge run" above answers whether a live judge answers these particular questions correctly on this dataset — mostly yes, with one blind spot ( template_progress) still open on the model tested — but that was one model, through one host, on one synthetic dataset. Run bench/run.py --judge jev (needs TYPESAFE_API_KEY, or pass --base-url and --jev-path "" for a host that serves the protocol somewhere other than /v1/systemone) or --judge llm --model <your model> (needs an API key for that model, plus --input-usd-per-mtok and --output-usd-per-mtok if you want a real cost column) on your own traces before trusting this on production traffic. The --judge columns report cost, latency, and an undecided count (answers in the 0.4–0.6 band that fell back to the lexical check — see fallback_to_lexical_when_uncertain) alongside accuracy, specifically so a live run can be judged on more than a single number.

Verifying this without an API key

pip install -e ".[dev]"
python bench/make_dataset.py      # writes bench/dataset.jsonl
pytest                             # 29 tests, all offline, MockBackend only
python bench/run.py                # lexical baseline, no key needed
python bench/run.py --judge oracle # needs PYTHONPATH=bench (see below)

tests/test_semantic.py scripts a MockBackend explicitly for each blind spot (never relying on a shared default answer across differently-worded questions — an early draft of this benchmark got that wrong; see the git history) and checks the detector reaches the right sub-score. tests/test_backends.py checks the Jev and LLM-judge response parsers against fixed payloads, so they're covered without a network call. tests/test_lexical.py pins the lexical port's behaviour against the exact deepeval source it's copied from, including the weighting quirk above, so a change to either can't silently drift from what's documented here.

Layout

semloop/
  trace.py       Span/Trace model; adapts deepeval's nested trace dict
  lexical.py     Faithful port of deepeval's current scoring, for the baseline
  semantic.py    SemanticLoopDetection: the judged replacement
  backends/
    jev.py       TypeSafe System One API client (stdlib only, no deps;
                 path= overridable for hosts that don't serve it at
                 /v1/systemone, e.g. an OpenRouter proxy)
    llm.py       Any OpenAI-compatible chat model as a judge, for comparison
    __init__.py  Answer/Backend/JudgeResponse types, plus MockBackend
bench/
  make_dataset.py    builds bench/dataset.jsonl from the 10 families
  run.py             runs lexical / oracle / jev / llm and reports accuracy, cost, latency
  oracle_backend.py  the scripted stand-in described above
tests/               29 offline tests (pytest)

Status

v0.1.0. Built from a read of deepeval's agent_loop_detection.py on 2026-09-27; the deepeval maintainers have their own reasons for keeping the metric deterministic and zero-cost by default (their own docstring names the tradeoff), so this is offered as a standalone package and a measured case for an optional judged mode — not yet proposed upstream. Validated once against a real judge the same day ("A real judge run" above); not yet run against production traffic.

VBS2004/semloop

Semantic loop detection for agent traces: a judged drop-in for deepeval's word-overlap stagnation check

Python

0

5 commits

updated Sep 27, 2026

See the code

See what people are saying

README

semloop

Semantic loop detection for agent traces. A drop-in alternative to the stagnation sub-signal in deepeval's AgentLoopDetectionMetric, which asks "do these two steps share enough words?" This asks the question that was actually meant: did the agent get anywhere?

from semloop import Trace, SemanticLoopDetection, JevBackend

trace = Trace.from_deepeval_trace_dict(test_case._trace_dict, goal=test_case.input)
metric = SemanticLoopDetection(JevBackend())  # needs TYPESAFE_API_KEY
result = metric.measure(trace)
print(result.score, result.reason)

JevBackend and LLMJudgeBackend both implement the same tiny interface — ask(state, questions) -> JudgeResponse — so the same detector runs against TypeSafe's Jev, any OpenAI-compatible chat model, or a scripted MockBackend in tests.

Why

deepeval/metrics/agent_loop_detection/agent_loop_detection.py (read 2026-09-27) scores reasoning stagnation with bigram Jaccard overlap and difflib.SequenceMatcher on consecutive LLM outputs. Its own docstring names the limitation:

Stagnation detection uses structural text similarity... It will miss semantically identical outputs that are worded very differently. An LLM-as-judge mode would solve this but would sacrifice the deterministic / zero-cost properties.

Jev (TypeSafe AI's "System One" model, $0.042 per million input tokens, output free, ~70–500ms per call — see the launch post) answers exactly this shape of question — "is this true, yes or no?" — for a fraction of a cent per trace, which makes the zero-cost tradeoff in that docstring no longer the right one to make.

What it fixes, with numbers

bench/dataset.jsonl (60 hand-built traces, 10 families, 6 instances each — generated by bench/make_dataset.py, seed fixed for reproducibility) isolates five documented blind spots in the lexical check, plus four healthy runs it should leave alone. Running the actual port of deepeval's scoring (semloop/lexical.py) against it:

$ python bench/run.py
Stagnation signal, correct calls per family (like-for-like: reasoning check only)
family                    lexical
--------------------------------
alternating_loop            0/6      <- A-B-A-B cycle: only consecutive pairs are compared
boilerplate_progress        0/6      <- false positive: identical preamble, real work underneath
cosmetic_args_loop          0/6      <- not a reasoning defect; see the tool-repetition table below
identical_tool_loop         6/6      <- what the lexical check is already good at
paraphrased_loop            0/6      <- same intent, reworded every step
progress                    6/6      <- healthy, correctly left alone
retry_then_recover          6/6      <- healthy, correctly left alone
template_progress           0/6      <- false positive: 5 different files, same sentence shape
terse_loop                  0/6      <- steps under 20 meaningful words are skipped, not compared
verbatim_loop                6/6      <- what the lexical check is already good at

Totals: stagnation 24/60 (40%) · metric verdict 36/60 (60%)

Six of the ten families score 0/6. Two of those six (identical_tool_loop, verbatim_loop) are cases the lexical check is supposed to get right — they're here as a regression check on any replacement, not as a complaint.

An oracle run — SemanticLoopDetection wired to a backend that answers each question the way a judge that actually read the trace should (see "Verifying this without an API key" below) — gets 54/60 (90%):

$ PYTHONPATH=bench python bench/run.py --judge oracle
family                    lexical          oracle
------------------------------------------------------
alternating_loop            0/6             6/6
boilerplate_progress        0/6             6/6
cosmetic_args_loop          0/6             0/6   <- see "A weighting quirk" below
identical_tool_loop         6/6             6/6
paraphrased_loop            0/6             6/6
progress                    6/6             6/6
retry_then_recover          6/6             6/6
template_progress           0/6             6/6
terse_loop                  0/6             6/6
verbatim_loop                6/6             6/6

Totals: lexical 24/60 (40%) · oracle 54/60 (90%)

The oracle is not a benchmark of Jev, or of any real judge. It scripts answers from the dataset's own family labels — see bench/oracle_backend.py's docstring. What it proves is that the mechanism (question construction, answer parsing, score combination, one request per trace) reaches the right verdict when its questions are answered correctly. Whether a live Jev or LLM judge actually answers them correctly on your traces is a separate, unavoidable question — see "What this doesn't prove" below.

A real judge run (2026-09-27)

The question above has an answer now, not just an oracle's stand-in for one. bench/run.py --judge jev was run once against typesafe/jev-1.13 through a host that serves the System One protocol at its own path rather than TypeSafe's /v1/systemone — JevBackend now takes a path= override for exactly this (--jev-path "" on the CLI, when --base-url is already the full endpoint):

$ python bench/run.py --judge jev --model typesafe/jev-1.13 \
    --base-url <endpoint> --jev-path ""
family                    lexical     jev:typesafe/jev-1.13
-------------------------------------------------------------
alternating_loop            0/6             6/6
boilerplate_progress        0/6             6/6
cosmetic_args_loop          0/6             6/6
identical_tool_loop         6/6             6/6
paraphrased_loop            0/6             6/6
progress                    6/6             6/6
retry_then_recover          6/6             6/6
template_progress           0/6             0/6   <- see below
terse_loop                  0/6             6/6
verbatim_loop                6/6             6/6

Totals: stagnation — lexical 24/60 (40%) · jev 54/60 (90%)
        verdict    — lexical 36/60 (60%) · jev 60/60 (100%)
        $0.0042 total (99,817 input tokens) · 779 ms mean latency · 7 undecided · 0 errors

This is not the oracle: every answer came from the model actually reading each trace's content, with no knowledge of which family generated it. The stagnation number lands almost exactly on the oracle's 90% — the mechanism and a real judge agree, which is the result you want from this kind of check. 7 of the 60 traces fell in the 0.4–0.6 undecided band and were resolved by fallback_to_lexical_when_uncertain; none errored.

One place the real judge disagrees with the oracle's script. template_progress (five steps, each touching a different file, described in the same sentence shape) fools the real judge on the fine-grained stagnation question exactly the way it fools the lexical check — 0/6, where the oracle had assumed a competent judge would score 6/6. It doesn't flip the overall verdict here only because the 0.35 weight on reasoning stagnation isn't enough on its own to fail the trace (see "A weighting quirk" below) — the false-positive blind spot this package claims to close is, for this model, at least partly still open, just hidden by the weighting rather than fixed. A caller who raises the stagnation weight, or reads score_breakdown directly instead of the combined score, would still see it.

This is one run, on one synthetic dataset, through one specific host — not proof this holds on production traffic, and not a claim about TypeSafe's own endpoint or any other model. It is a real data point where the section above could only offer a scripted one.

The five blind spots, and how each is closed

  1. Paraphrased loops. Bigram Jaccard and SequenceMatcher both stay below 0.47 on ten different ways to say "I'm going to search the codebase" (see bench/make_dataset.py's PARAPHRASES). Fixed by asking directly: did step j advance the goal beyond step i (progress_{i}_{j}, one noul question per consecutive pair).
  2. Terse steps are skipped, not compared. if len(words) < 20: continue — a loop of short retries is invisible to the lexical check by construction. build_questions has no length filter; every consecutive pair gets a question regardless of how short the steps are.
  3. A-B-A-B loops. Only consecutive pairs are compared, so a period-2 cycle never trips it. revisit_{i}_{j} questions at lookahead distance ask directly whether step j repeats step i, independent of what sits between them.
  4. Cosmetic argument changes reset tool repetition. The exact-match key is (name, sorted(args.items())); adding one flag resets the count. Closed with a fourth question family, tool_repeat_{name}, for any tool called at least repetition_threshold times: are these calls the same underlying action despite different arguments? The more severe of the exact-match and judged checks wins (min in SemanticLoopDetection.measure), so this can only add detections, never hide one the lexical check already made.
  5. False positives on real progress. Five steps that each process a different file, described in the same sentence shape, or a fixed corporate preamble in front of real one-line updates, both push SequenceMatcher past 0.85 (bench/make_dataset.py's template_progress and boilerplate_progress — see the measured pairwise similarities in that file's docstrings). Because the judged question is did this advance the goal, not do these share words, genuine progress across different files or underneath a fixed preamble reads as what it is.

Every question about one trace — however many pairs, however many repeated tool names — goes in a single request (test_one_request_per_trace_regardless_of_pair_count in tests/test_semantic.py), which is what keeps the added cost near zero even on a long trace.

A weighting quirk worth knowing about (not something this package changes)

deepeval's _combine_scores weights the three sub-signals 0.40 (tool repetition) / 0.35 (reasoning stagnation) / 0.25 (call graph cycles), and the default threshold is 0.5. That means no single sub-signal, however severe, can flip the overall verdict alone: a fully-confident detection (sub-score 0.0) on the heaviest-weighted signal still leaves the composite at 1 - 0.40 = 0.60, above threshold. cosmetic_args_loop's oracle row shows this directly — the tool-repetition sub-score hits 0.0 with a confidently correct judge, and the overall verdict still doesn't flip until the caller raises the threshold above 0.60 (tests/test_semantic.py's test_a_lone_signal_needs_a_stricter_threshold_to_flip_the_verdict, and tests/test_lexical.py's test_single_severe_signal_cannot_flip_the_default_verdict). This applies identically whether the sub-signal comes from the lexical check or a judge — it's a property of the combination step this package keeps unchanged for compatibility, not a limitation of the semantic swap. A caller who wants a single confirmed signal to be decisive should raise the metric's threshold, or read score_breakdown directly instead of relying on the combined verdict.

What this doesn't prove

The dataset is hand-built from templates, not sampled from production traces. It's enough to show each blind spot exists and is reproducible — it says nothing about how often paraphrased loops, cosmetic argument changes, or false positives on genuine multi-file progress actually occur in your traffic. "A real judge run" above answers whether a live judge answers these particular questions correctly on this dataset — mostly yes, with one blind spot ( template_progress) still open on the model tested — but that was one model, through one host, on one synthetic dataset. Run bench/run.py --judge jev (needs TYPESAFE_API_KEY, or pass --base-url and --jev-path "" for a host that serves the protocol somewhere other than /v1/systemone) or --judge llm --model <your model> (needs an API key for that model, plus --input-usd-per-mtok and --output-usd-per-mtok if you want a real cost column) on your own traces before trusting this on production traffic. The --judge columns report cost, latency, and an undecided count (answers in the 0.4–0.6 band that fell back to the lexical check — see fallback_to_lexical_when_uncertain) alongside accuracy, specifically so a live run can be judged on more than a single number.

Verifying this without an API key

pip install -e ".[dev]"
python bench/make_dataset.py      # writes bench/dataset.jsonl
pytest                             # 29 tests, all offline, MockBackend only
python bench/run.py                # lexical baseline, no key needed
python bench/run.py --judge oracle # needs PYTHONPATH=bench (see below)

tests/test_semantic.py scripts a MockBackend explicitly for each blind spot (never relying on a shared default answer across differently-worded questions — an early draft of this benchmark got that wrong; see the git history) and checks the detector reaches the right sub-score. tests/test_backends.py checks the Jev and LLM-judge response parsers against fixed payloads, so they're covered without a network call. tests/test_lexical.py pins the lexical port's behaviour against the exact deepeval source it's copied from, including the weighting quirk above, so a change to either can't silently drift from what's documented here.

Layout

semloop/
  trace.py       Span/Trace model; adapts deepeval's nested trace dict
  lexical.py     Faithful port of deepeval's current scoring, for the baseline
  semantic.py    SemanticLoopDetection: the judged replacement
  backends/
    jev.py       TypeSafe System One API client (stdlib only, no deps;
                 path= overridable for hosts that don't serve it at
                 /v1/systemone, e.g. an OpenRouter proxy)
    llm.py       Any OpenAI-compatible chat model as a judge, for comparison
    __init__.py  Answer/Backend/JudgeResponse types, plus MockBackend
bench/
  make_dataset.py    builds bench/dataset.jsonl from the 10 families
  run.py             runs lexical / oracle / jev / llm and reports accuracy, cost, latency
  oracle_backend.py  the scripted stand-in described above
tests/               29 offline tests (pytest)

Status

v0.1.0. Built from a read of deepeval's agent_loop_detection.py on 2026-09-27; the deepeval maintainers have their own reasons for keeping the metric deterministic and zero-cost by default (their own docstring names the tradeoff), so this is offered as a standalone package and a measured case for an optional judged mode — not yet proposed upstream. Validated once against a real judge the same day ("A real judge run" above); not yet run against production traffic.

Languages

Python

100.0%