chunxiaox/nautilus-compass-test-data

Dataset

Nautilus Compass · Drift Detection Anchors

0

3 commits

1 linked in READMEs

updated May 13, 2026

See the code

README

Nautilus Compass · Drift Detection Anchors

Behavioral anchor packs used by Nautilus Compass for black-box agent persona drift detection.

Each anchor is a short natural-language sentence representing either aligned behavior (positive anchor) or drifted/anti-pattern behavior (negative anchor). Drift detection works by computing the weighted top-k mean cosine similarity (BGE-m3 embeddings) of an agent's latest message against the positive vs. negative anchor cloud:

drift_score = weighted_topk_mean(cos(msg, positive)) - weighted_topk_mean(cos(msg, negative))

Implementation: daemon.py:344-365.

Configs

marketing-anchors-v1.2 (this release)

  • Domain: compass project self-marketing / public-facing copy
  • Language: English
  • Size: 31 positive + 37 negative = 68 anchors
  • Purpose: prevent compass agents (V7 partnership-loop, engagement-cron, blog drafts) from over-claiming, dismissing competitors, or hype-spamming
  • Calibration: round 2 (2026-05-11), 5 topical-FP anchors rewritten to extreme-literal phrasing to reduce topical-similarity false positives (10/13 alert → target <5/13 on v14 13-paragraph dogfood)

Anchor schema:

fieldtypeexample
anchor_idstrp01, n05
polaritystrpositive / negative
textstr"compass is a black-box memory layer · no LLM extraction at index time"
anchor_packstrcompass_marketing_v1.2

Usage

from datasets import load_dataset

ds = load_dataset("chunxiaox/nautilus-compass-test-data", "marketing-anchors-v1.2", split="anchors")
print(ds[0])
# {'anchor_id': 'p01', 'polarity': 'positive', 'text': '...', 'anchor_pack': 'compass_marketing_v1.2'}

positive = ds.filter(lambda r: r["polarity"] == "positive")
negative = ds.filter(lambda r: r["polarity"] == "negative")

To use with the runtime drift checker, embed each anchor with BGE-m3 and apply the formula above.

Coming next

This is the first config in a series. Planned additions:

  • roc-eval-v1 — labeled drift-detection ROC evaluation data (cosine scores + binary labels, ~500 utterances) — needs PII sanitization, ETA 2-3 days
  • claude-code-sessions-v1 — redacted multi-turn agent traces for downstream benchmarking — ETA 2-3 weeks
  • anchors-domain-finance / anchors-domain-legal — per-domain anchor packs

Citation

@misc{wang2026compass,
  title = {Nautilus Compass: Black-Box Persona Drift Detection for Production LLM Agents},
  author = {Wang, Chunxiao},
  year = {2026},
  eprint = {2605.09863},
  archivePrefix = {arXiv},
  primaryClass = {cs.CL},
  url = {https://arxiv.org/abs/2605.09863}
}

Paper page: https://huggingface.co/papers/2605.09863

License

MIT — same as the compass codebase. Free to use for research and commercial purposes; attribution appreciated.

Acknowledgments

Thanks to Niels Rogge at Hugging Face for the onboarding push (issue #8).

agent-memory
anchors
behavioral-drift
compass
llm-safety

chunxiaox/nautilus-compass-test-data

Dataset

Nautilus Compass · Drift Detection Anchors

0

3 commits

1 linked in READMEs

updated May 13, 2026

See the code

README

Nautilus Compass · Drift Detection Anchors

Behavioral anchor packs used by Nautilus Compass for black-box agent persona drift detection.

Each anchor is a short natural-language sentence representing either aligned behavior (positive anchor) or drifted/anti-pattern behavior (negative anchor). Drift detection works by computing the weighted top-k mean cosine similarity (BGE-m3 embeddings) of an agent's latest message against the positive vs. negative anchor cloud:

drift_score = weighted_topk_mean(cos(msg, positive)) - weighted_topk_mean(cos(msg, negative))

Implementation: daemon.py:344-365.

Configs

marketing-anchors-v1.2 (this release)

  • Domain: compass project self-marketing / public-facing copy
  • Language: English
  • Size: 31 positive + 37 negative = 68 anchors
  • Purpose: prevent compass agents (V7 partnership-loop, engagement-cron, blog drafts) from over-claiming, dismissing competitors, or hype-spamming
  • Calibration: round 2 (2026-05-11), 5 topical-FP anchors rewritten to extreme-literal phrasing to reduce topical-similarity false positives (10/13 alert → target <5/13 on v14 13-paragraph dogfood)

Anchor schema:

fieldtypeexample
anchor_idstrp01, n05
polaritystrpositive / negative
textstr"compass is a black-box memory layer · no LLM extraction at index time"
anchor_packstrcompass_marketing_v1.2

Usage

from datasets import load_dataset

ds = load_dataset("chunxiaox/nautilus-compass-test-data", "marketing-anchors-v1.2", split="anchors")
print(ds[0])
# {'anchor_id': 'p01', 'polarity': 'positive', 'text': '...', 'anchor_pack': 'compass_marketing_v1.2'}

positive = ds.filter(lambda r: r["polarity"] == "positive")
negative = ds.filter(lambda r: r["polarity"] == "negative")

To use with the runtime drift checker, embed each anchor with BGE-m3 and apply the formula above.

Coming next

This is the first config in a series. Planned additions:

  • roc-eval-v1 — labeled drift-detection ROC evaluation data (cosine scores + binary labels, ~500 utterances) — needs PII sanitization, ETA 2-3 days
  • claude-code-sessions-v1 — redacted multi-turn agent traces for downstream benchmarking — ETA 2-3 weeks
  • anchors-domain-finance / anchors-domain-legal — per-domain anchor packs

Citation

@misc{wang2026compass,
  title = {Nautilus Compass: Black-Box Persona Drift Detection for Production LLM Agents},
  author = {Wang, Chunxiao},
  year = {2026},
  eprint = {2605.09863},
  archivePrefix = {arXiv},
  primaryClass = {cs.CL},
  url = {https://arxiv.org/abs/2605.09863}
}

Paper page: https://huggingface.co/papers/2605.09863

License

MIT — same as the compass codebase. Free to use for research and commercial purposes; attribution appreciated.

Acknowledgments

Thanks to Niels Rogge at Hugging Face for the onboarding push (issue #8).

agent-memory
anchors
behavioral-drift
compass
llm-safety