SWE-bench
A benchmark family that evaluates AI systems on real software engineering issues.
Plain English
It tests whether an AI can fix real code, not just answer coding trivia.
Example
A model is scored by whether its patch passes tests for a real GitHub issue.
Why it matters
SWE-bench became a shorthand for practical coding-agent capability.