senku14x/Validating-NLAs

1

stars

123

commits

Python

primary language

Aug 22, 2026

updated

README

Validating-NLAs

The NLA paper claims NLAs surface safety-relevant cognition that models represent internally but do not verbalize, with evaluation awareness as the headline case. They could not validate this against ground truth because a real model's internal beliefs are unobservable. This project replaces unobservable ground truth with constructed ground truth, using model organisms where we control exactly what cognition is present and whether it is verbalized, and then measures whether the NLA actually detects it.

Work In Progress :D

Contributors

claude

70 commits

senku14x

51 commits

Veeraraju-E

2 commits

senku14x/Validating-NLAs

1

stars

123

commits

Python

primary language

Aug 22, 2026

updated

README

Validating-NLAs

The NLA paper claims NLAs surface safety-relevant cognition that models represent internally but do not verbalize, with evaluation awareness as the headline case. They could not validate this against ground truth because a real model's internal beliefs are unobservable. This project replaces unobservable ground truth with constructed ground truth, using model organisms where we control exactly what cognition is present and whether it is verbalized, and then measures whether the NLA actually detects it.

Work In Progress :D

Contributors

claude

70 commits

senku14x

51 commits

Veeraraju-E

2 commits

Languages

Python

97.8%

Shell

2.2%