The NLA paper claims NLAs surface safety-relevant cognition that models represent internally but do not verbalize, with evaluation awareness as the headline case. They could not validate this against ground truth because a real model's internal beliefs are unobservable. This project replaces unobservable ground truth with constructed ground truth, using model organisms where we control exactly what cognition is present and whether it is verbalized, and then measures whether the NLA actually detects it.
Work In Progress :D
Python
97.8%
Shell
2.2%
The NLA paper claims NLAs surface safety-relevant cognition that models represent internally but do not verbalize, with evaluation awareness as the headline case. They could not validate this against ground truth because a real model's internal beliefs are unobservable. This project replaces unobservable ground truth with constructed ground truth, using model organisms where we control exactly what cognition is present and whether it is verbalized, and then measures whether the NLA actually detects it.
Work In Progress :D
Python
97.8%
Shell
2.2%