
Specific Labs
Can coding agents can do real engineering work on codebases they've never seen?
launched 21h ago ยท by Janak & Sid
Roasty review of Specific Labs
Nobody has swiped on Specific Labs yet.
Specific Labs launched on 10 September 2026 on Y Combinator, in the week of 7 September 2026. Every swipe on the card above counts.
What they say it does
TL;DR\ Specific Labs (YC F25) is launching Real-SWE, a benchmark that evaluates coding agents on real software engineering tasks that engineers performed, using private, out-of-distribution codebases. Public benchmarks are built on open source repos models have already seen. Real-SWE tests whether an agent can do the work inside a company it has never encountered. Leaderboard and task examples at \ \ Hi, we're Janak and Sid\ Over the past year we've acquired and licensed operational data and codebases from real companies, and one question kept coming up with labs: how do coding agents actually perform on private code? Nobody had a clean way to measure it. So we built one. The problem\ Every major coding benchmark is built on public repos. Models have trained on that code or on code that looks a lot like it. Scores keep climbing, but the number companies care about is different - can this agent land a fix in our billing system, our permissions layer, our customer data pipeline, without
Alternatives to Specific Labs
All Specific Labs alternatives โFounder and want the vote updates? Email hello@roasty.tech from your @withspecific.com address and we'll set it up.



