Specific Labs benchmark puts AI coding agents to work on private enterprise code
Real-SWE recorded a 28.6% pass rate for runs under 10 minutes and 26.6% for longer runs across 640 scored rollouts.
Specific Labs has released Real-SWE, a benchmark designed to assess AI coding agents on tasks drawn from licensed private production codebases. The benchmark covers 10 enterprise software-engineering tasks and eight model-and-harness configurations.
The 640 scored rollouts produced a 28.6% pass rate among runs completed in under 10 minutes, compared with 26.6% for runs taking 10 minutes or longer. The results do not indicate that longer runtimes improved performance.
The tasks represent work involving billing, tax, datastores, identity and analytics. Specific Labs says the benchmark is intended to test whether agents can navigate existing architecture, business logic, coding conventions and operational constraints rather than solve isolated or synthetic coding problems.
Reported resolution rates ranged from 16.2% for GPT-5.6 Sol to 38.8% for Fable 5.1. Those figures compare combined model-and-harness configurations, not models in isolation.
Real-SWE uses native harnesses, isolated sandboxes and verifiers injected during grading. Specific Labs says a typical instruction is about 1,742 characters and spans roughly 11 files, while the tasks are inspired by or lifted from private company codebases.
The supplied material does not independently verify the licensing arrangements, participating companies or methodology. Some model names and cost figures also require verification, meaning the results should be treated as findings from Specific Labs’ benchmark rather than a definitive measure of general software-engineering ability.
