Lovable now builds with Fable 5.1.
Early testing shows it's especially good at fixing and improving existing apps without breaking what already works, outperforming Fable 5 by up to 17% on the most difficult tasks.
This matters because many Lovable builders spend a lot of their time improving live apps, not starting from scratch. Working in existing, complex codebases has historically been more difficult for models, but Fable 5.1 handles this significantly better, at a 31% lower cost.
It's also markedly better at checking its own work and has sharper visual taste, with a 3.5% lift in visual design quality.
For people building with Lovable, this means more reliable updates, higher design quality, and less time spent fixing unintended issues.
This post breaks down how we benchmarked it, where it wins, and what it means for your builds.
How we evaluate new models
We run every candidate model against Lovable's internal benchmark suite, which measures end-to-end app-building performance on:
- 0-to-1 building: building from scratch with Lovable and achieving the desired result for our users
- Fixing and evolving existing codebases: going back into a live app to add a feature, fix a bug, or rework a screen without breaking what's already there (the largest category of work builders do in Lovable)
- Design quality: how closely the output matches what a builder actually pictured
We evaluated Fable 5.1 against Fable 5 at equivalent reasoning effort levels: low, medium, and high. For every benchmark, each task was run five times.
The scores below show Fable 5.1’s weighted composite performance relative to Fable 5. Our evaluation measures token efficiency, self-verification, including whether the model tests its own claims, and how closely the final result matches the user’s request.
Score deltas for Fable 5.1 vs. Fable 5
| Benchmark / Effort Level | Low | Medium | High |
|---|---|---|---|
| 0-to-1 building | -1.4% | -0.9% | -2.7%* |
| Iterative code fixing | +11.9%* | +12.1%* | +17.2%* |
| UI and visual design | +1.6% | +2.7%* | +3.5%* |
Cost-per task deltas for Fable 5.1 vs. Fable 5
| Benchmark / Effort Level | Low | Medium | High |
|---|---|---|---|
| 0-to-1 Building | +6.3% | +2.9% | +7.1% |
| Iterative code fixing | -36.3% | -39.2% | -31.4% |
| UI and visual design | +15.8% | +5.4% | +7.8% |
* Indicates a statistically significant difference between Fable 5.1 and Fable 5 at the 95% confidence level. Green indicates that Fable 5.1 scored higher, while red indicates that it scored lower.
What changed under the hood
Four things stand out in how Fable 5.1 handles Lovable workloads:
- More disciplined self-verification. Fable 5.1 opens the running app in the browser before calling a task finished. This improved discipline improves request-fidelity, meaning the output matches what the user asked for more often.
- A real edge on existing codebases. Taking something live and making it do more is where Fable 5.1 pulls furthest ahead of Fable 5, and the lead grows as reasoning effort increases, all at a consistently lower cost per task across varying effort levels.
- Sharper taste at higher effort. Design quality improves at medium and high reasoning effort, so the more Fable 5.1 is asked to think, the closer the result lands to what you pictured.
What this means for builders
Fable 5.1 is the model Lovable reaches for when you're building something complex on top of apps that are already live.
Where Fable 5.1 pulls furthest ahead of Fable 5 is the work our users do most: taking an app that's already up and running and making it do more, design included. It's just more careful, checking its own work and opening the app to make sure it actually runs before calling anything done. The result looks and works the way you pictured it.
— Fabian Hedin, CTO & Co-founder, Lovable
As our A/B tests continue, we’re routing sessions that get stuck, along with longer and more complex work, to Fable 5.1.



