I had Fable and Astra build the same iPhone app
One intentionally vague prompt, two frontier models. Astra finished in 18 minutes for $24; Fable took three hours and $375 and shipped the better app. Codex is the better executor, Claude is the better manager — which you prefer says as much about how you like to work as it does about the model.
I had Fable and Astra build the same iPhone app, but that's not the interesting part.
Pretty much all the impressive one-shot demos I've seen are either:
- copies of something else or
- a visually impressive demo
But that's not how real products get built. In the real world you start with a vague idea that you refine over time while you figure out the requirements. So that's the test I put both through.
One prompt, intentionally vague.
Astra was the first to finish. It didn't ask me a single question and "completed" the task in 18 minutes. But... the app opened to a blank screen. After a round of review and iteration, at 41 minutes I had a working app. Aesthetically nice but no onboarding and where you set which apps to restrict is buried. Otherwise, solid for quite literally one prompt.
The cost... 17m tokens and ~$24 at API rates.
Fable on the other hand took 80 minutes to produce a working app. It then did a final adversarial review loop which took another hour and used 42 parallel agents. Same as Astra, no questions just one shot.
The upside, no fixes required. Aesthetically slightly simpler design, but it included an onboarding and also had a simulator mode so I could test in the sim which doesn't support screen time. It was also 4,400 lines to Astra's 790.
The cost... 205m tokens and ~$375 at API rates.
Which is better very much depends on what you're trying to do. Astra won hands down on speed and cost. If you're working interactively it's the better model for it. Getting a first draft in 18 minutes leaves a lot of room for iteration.
Fable did a better job accomplishing the goal at a cost of time and money. It's the model I'd choose if you want to hand off complex work with more limited oversight. It wins when you want to walk away for a couple of hours and come back to something finished
At its core, to me this highlights the increasing divergence between the two. Codex is the better executor, Claude is the better manager.
Which you prefer comes down not just to benchmarks or API rates, but how you like to work.
This divergence has been growing since 4.7 when Anthropic increasingly optimized for long-horizon work. When I see someone say Opus 4.6 was the last good model, that tells me as much about how they like to work as it tells me about the model.









