tested, not assumed

The benchmark, every answer included.

10 prompts across five scenarios, 3 samples each, blind-judged against the same model with no instructions. Required facts per prompt were fixed before the run. Pick a comparison:

raw data, fixtures & rubric
Don't trust us.
Rerun it.

how we measured

WordsCounted by splitting on whitespace.
TokensEstimated with OpenAI's tiktoken (o200k_base) on the answer text only, so both models are counted the same way.
JudgesClaude answers were judged by Claude Opus, GPT answers by GPT-6.1 Sol. The Anthropic-vs-OpenAI head-to-head is the only comparison where the judge (Claude Opus) didn't write either answer.