Empirical framework for testing whether RL training instills genuine behavioral dispositions or surface compliance in language models, using compute frugality as a controllable proxy value.
-
Updated
Jul 3, 2026 - Python
Empirical framework for testing whether RL training instills genuine behavioral dispositions or surface compliance in language models, using compute frugality as a controllable proxy value.
Tested whether Anthropic's sleeper agent detection technique generalizes to naturally-arising deception. Ran Llama 70B inference across 90 prompt-conditions on rented A100s, manually scored every output. Honest null result. BlueDot Impact AI Safety Sprint.
Add a description, image, and links to the alignment-faking topic page so that developers can more easily learn about it.
To associate your repository with the alignment-faking topic, visit your repo's landing page and select "manage topics."