AI Model Report

Benchmarks · 11 pieces on file

Benchmarks

Methodology, regression suites, leaderboard inflation, and the numbers behind every comparison the desk publishes.


Feature · SEPTEMBER 14, 2026

Real-SWE puts frontier coding agents at 38.8% on real enterprise tickets — and one task went 0-for-64

Specific Labs' new benchmark runs eight frontier models against ten licensed private-codebase tickets. Fable 5.1 leads at 38.8% for $6.96 a rollout; Gemini 3.8 Flash trails by 7.6 points at a third the cost.

By Linnea Halberg · Benchmarks desk

Read the full piece →


More in Benchmarks