Originally published on LinkedIn.
Specialized models are now viable for real engineering work. Almost all of our development runs through an autonomous loop, and we'd been on a trajectory where each new frontier model made that loop markedly better. With the latest generation, that's no longer the pattern.
We run that loop on Cursor Cloud Agents. We started there because the harness was easy to stand up, and it's the most ergonomic remote agent environment we've used. We assumed the coding in that loop needed the best general model available.
We tried Composer on a lark. We were blown away: for most of the coding we run through the loop, it was a drop-in replacement. It's faster too. We haven't benchmarked it precisely, but the gap is obvious, and public numbers put it near 50% quicker on coding tasks. It's also 90% cheaper than Opus.
It isn't good at everything. Behind our QA agent it was noticeably worse. That's the point. It's specialized for coding.
For us the tradeoff flipped. A cheaper, faster, good-enough coding model beats a general model that's stronger at things the task never touches. Composer is the good-enough one today. I'd expect that to flip: specialized coding models pulling ahead of general ones at coding outright, and getting cheaper and faster as they do.
QA is the gap for us now, so if you know a model specialized for it, tell me.
I'm surprised I don't hear more about Composer. And it raises a real question for model specialization: who wins there? Maybe it's the major foundation labs that already hold all the leverage. Maybe specialization is the crack where someone else gets in.
