Eval-gaming detector
A defensive eval-gaming / sandbagging detector. Flags when a model-evaluation score is decoupled from true capability. Abstract scores only; no techniques.
No separate design note is published for this file. The description above is taken from the toolkit README and architecture index. See the source on GitHub for the implementation and self-test.