Roasty
Y Combinator๐Ÿ”ฅ live ยท 7d left

Specific Labs

Can coding agents can do real engineering work on codebases they've never seen?

launched 21h ago ยท by Janak & Sid

โ€”
toast rating ยท 0 votes
๐Ÿซฅ
Untouched
0 toasts0 roasts
Visit site โ†—Launch page โ†—Share on X

Roasty review of Specific Labs

Nobody has swiped on Specific Labs yet.

Specific Labs launched on 10 September 2026 on Y Combinator, in the week of 7 September 2026. Every swipe on the card above counts.

B2BArtificial IntelligenceSaaS๐Ÿƒ B2B deck โ†’

What they say it does

TL;DR\ Specific Labs (YC F25) is launching Real-SWE, a benchmark that evaluates coding agents on real software engineering tasks that engineers performed, using private, out-of-distribution codebases. Public benchmarks are built on open source repos models have already seen. Real-SWE tests whether an agent can do the work inside a company it has never encountered. Leaderboard and task examples at \ \ Hi, we're Janak and Sid\ Over the past year we've acquired and licensed operational data and codebases from real companies, and one question kept coming up with labs: how do coding agents actually perform on private code? Nobody had a clean way to measure it. So we built one. The problem\ Every major coding benchmark is built on public repos. Models have trained on that code or on code that looks a lot like it. Scores keep climbing, but the number companies care about is different - can this agent land a fix in our billing system, our permissions layer, our customer data pipeline, without

Founder and want the vote updates? Email hello@roasty.tech from your @withspecific.com address and we'll set it up.