Case Study: When 0% Hallucination Isn’t What It Seems — Rethinking Benchmark Contradictions After Claude 4.1 Opus
https://spiral-yamamomo-ae7.notion.site/Debate-Mode-Oxford-Style-for-Strategy-Validation-Using-Structured-Argument-AI-3b3825cb4779804db9eced87f73db21c
How a leading hallucination benchmark was flipped by a refusal-first model Last year a public benchmark report showed Claude 4.1 Opus with a startling outcome: 0% hallucination