10 prompts across five scenarios, 3 samples each, blind-judged against the same model with no instructions. Required facts per prompt were fixed before the run. Pick a comparison:
TokensEstimated with OpenAI's tiktoken (o200k_base) on the answer text only, so both models are counted the same way.
JudgesClaude answers were judged by Claude Opus, GPT answers by GPT-6.1 Sol. The Anthropic-vs-OpenAI head-to-head is the only comparison where the judge (Claude Opus) didn't write either answer.