Test-Time Compute
Spending more compute during inference to improve answer quality without retraining the model.
Plain English
The model gets more time or attempts to think when the task is hard.
Example
A coding benchmark run gives an agent multiple attempts, tool calls, or verifier passes before selecting a final patch.
Why it matters
The frontier is moving from only bigger training runs to smarter inference strategies.