Models are most confident when they are wrong, an eval harness finds
An eval harness surfaced a pattern qualitative review had missed: AI models display their highest confidence precisely when their answers are wrong.
Research
An eval harness surfaced a pattern qualitative review had missed: AI models display their highest confidence precisely when their answers are wrong.