the.ai

Evaluation / Capabilities

verified

Chain of Thought

Asked for the answer to a multi-step problem, a model that must produce it immediately often gets it wrong. Asked to work through the steps first, it does much better. The reasoning is written into the output because that is the only place the model has to put it — there is no scratchpad, so the tokens are the scratchpad.

Viz primitive · budget-splitreasoning-tokens = 80

reasoning-tokens holds 67% of the budget; rest holds the remaining 33%.

Share of the generated tokens spent reasoning before the answer, against the answer itself. Drag the chain length to watch the answer become a small part of what is produced.

80

Reviewed by opendroid · 2026-08-04

  • arXiv:2201.11903 — Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
  • arXiv:2305.04388 — Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting