The early Gemini 4 and Fable 5.2 reports are worth taking seriously as signs of better visual coding—but they point to different strengths and tradeoffs. A controller SVG attributed to Gemini 4 looks competitive with GPT-6 Astra in a close comparison. A tester of the suspected next Fable reports better output than Astra at high reasoning, alongside slower, more expensive runs. That makes this an interesting quality-versus-time comparison, even before the model names and access routes are fully settled. Those observations come from Lumina's controller comparison and JAZII's early Fable report, checked on September 20, 2026.
Our provisional reading is that alleged Gemini 4 is showing promising visual precision, the suspected next Fable is attracting attention for richer results, and Astra is the documented baseline against which those improvements can be judged. None of the demonstrations establishes a general coding winner. This article examines the actual reported outputs and their implications; we did not run these models or reproduce the tests. GPT-6 here means Astra, while names such as “Gemini 4 Pro” and “Fable 5.2” retain the qualifications attached by their sources.
What the early comparisons actually show
| Report | Observed or reported result | Reasonable takeaway | What remains unresolved |
|---|---|---|---|
| Lumina: alleged Gemini 4 versus Astra controller SVG | Author calls both nearly perfect, with a slightly wrong logo on Astra's version | In this visual coding example, the alleged Gemini output appears competitive with Astra, with a possible advantage in a small detail | Backend identity and repeatability; the author did not have Fable 5.2 access |
| JAZII: suspected next Fable versus Astra, same prompt and high reasoning | Author prefers the Fable result, calls it a major improvement, but says it is slower and expensive | A possible quality gain may come with a higher time and cost budget | “5.2” is the author's assumed name; no measured latency or price table |
| Chetaslua: three pelican-on-bicycle graphics at claimed maximum settings | Fable/Opus examples show fuller coastal scenes; Astra uses a simpler vignette | More scene detail can explain why viewers prefer one output, depending on the task | Exact prompt, backend attribution, and whether extra detail was required |
| Bhavy: Rocket demonstration | Author reports more realistic output from alleged newer Fable routing | Another early signal of improved visual presentation | Main comparison is Fable versus Opus; an Astra preference in a reply is not an equivalent Astra test |
These reports are useful because they show or describe specific changes. They are also uneven: some include comparable images, some offer a tester's judgment, and some attach a next-generation name to an existing product label. Keeping those distinctions visible lets us analyze the improvements without turning every circulating attribution into a confirmed release.
Gemini 4: a close visual result, with a repeatability question
The Lumina controller montage is more informative than a blanket claim that Gemini has overtaken everything. Both the alleged Gemini 4 and Astra renderings depict detailed controller bodies, controls, and shading. Lumina describes them as nearly perfect, while pointing out a slightly incorrect logo in Astra's version.
For an SVG or frontend prototype, that is a meaningful comparison. Getting the overall silhouette right is one level of performance; placing controls, shaping surfaces, and reproducing small identifying details is another. The example supports a narrow judgment that the alleged Gemini output can match Astra's visual polish and may handle a particular detail better. It does not show a large general lead, because the two examples are already close.
The surrounding account makes waiting time and repeatability part of the story. The output was associated with an Arena entry labeled gemini 3.7 flash, reportedly took about 20 minutes, and came from the previous night. The author said they could not obtain similar results the next day. That could make the difference between an impressive one-off artifact and a useful everyday tool. The post does not establish why the behavior changed, so a model switch, different conditions, and ordinary output variation remain possibilities rather than settled explanations. Lumina also said they lacked Fable 5.2 access, which rules out reading this as a three-way test.
A separate Antigravity user report describes better output, roughly half the token-generation speed, and fewer tokens used under the Gemini 3.8 Flash label. It includes a graphic comparison the author says used the same prompt. The practical insight is to separate generation speed, total tokens, and time to a useful result. A model can generate tokens more slowly yet finish with fewer of them; it can also deliver a better first attempt and avoid a second run. The report supplies a reason to investigate those dimensions, not enough measurements to calculate a net productivity gain or identify the backend.
Secondary reporting widens the range of examples. TMTPost describes a voxel pagoda, helicopter, and monochrome website attributed by community testers to Gemini 4, with reported runs of roughly 8, 10, and 14 minutes. Those examples suggest that interest extends beyond one controller drawing into more elaborate visual construction. They still leave open whether the same system can reliably modify an existing application, preserve behavior, and pass tests.
Fable 5.2: the stronger-output claim comes with a price
JAZII's September 18 report is the clearest account of a direct quality tradeoff. The author says they used the same prompt and high reasoning, preferred the suspected next Fable over Astra, and saw a substantial improvement over current Fable. The same report calls the new behavior slower and expensive. In a follow-up, the author explicitly says the name is not confirmed: “Fable 5.2” is their assumption.
The useful conclusion is conditional but substantive: if that quality improvement repeats on your work, the next Fable could be attractive for jobs where a better first result is worth a longer wait. An intricate visual prototype or a deliverable requiring careful composition might fit. An interactive loop with many small edits could be less forgiving of the delay. “Expensive” remains the tester's assessment; the post supplies neither an official tariff nor an itemized cost per completed task.
Other examples help explain what people may be responding to. Chetaslua's comparison presents three pelicans riding bicycles, labeled as updated Fable routing to 5.2, Opus next, and Astra at maximum settings. The first two pictures use fuller coastal scenes; the Astra example has a more restrained, circular composition. The richer framing is a visible difference, but preference depends on the assignment. A scenic illustration rewards elaboration; a compact icon or reusable interface element may reward restraint. Without the exact prompt, extra scenery cannot automatically count as better instruction following.
Bhavy's Rocket post adds an enthusiastic report of more realistic output after claimed routing changes. Its main comparison is Fable against Opus. A reply favoring it over Astra is a separate judgment, and other viewers disagreed about realism. It adds to the early visual-quality signal without supplying another controlled Astra head-to-head.
These graphic and code-output demonstrations also do not establish native image generation by the underlying models. The relevant question is how well a system produces the code or artifact requested, not whether an attractive screenshot proves an additional modality.

Astra is a useful baseline, not an automatic winner
Astra matters here because the comparison does not have to wait for every rumor to become official. You can ask whether the alleged gains exceed what a documented model already delivers—and whether they would change a real workflow.
OpenAI's Astra model page identifies gpt-6-astra, lists text and image input with text output, and documents reasoning settings from low through max. Standard API pricing starts at $10 per million ordinary input tokens and $50 per million output tokens for requests with up to 272,000 input tokens; longer inputs use higher rates. Those are direct API prices, not subscription allowances. Our Astra pricing guide explains the remaining billing conditions.
That creates a concrete reference for testing the “better but slower and expensive” claim. Log total elapsed time, the cost of all attempts, whether the output meets the requirement, and the effort needed to repair it. A more expensive model can be cheaper per accepted result if it avoids retries. Equally, a visually richer answer can be a poor bargain if the extra material is unwanted or breaks the intended design.
The circulating benchmark sheet is weaker evidence for an overall ranking. TMTPost reports the sheet alongside discrepancies between it and the associated backend screenshot, an Astra comparator that does not match official numbers, and no endorsement from Google, Arena, or the benchmark maintainers. We therefore do not turn those circulated scores into a factual leaderboard. This problem concerns that numerical comparison; it does not erase the visual demonstrations or the testers' firsthand observations. Published evaluator results, such as Artificial Analysis, also identify particular models and settings rather than establish this exact rumored three-way matchup.
What this means for your next model choice
For SVG, visual prototypes, and frontend presentation, the alleged Gemini and Fable improvements are worth following now. Gemini's close controller comparison suggests strong execution of an existing visual target. The suspected Fable examples and reports suggest more ambitious detail and composition, with a possible time and cost penalty. These are provisional readings of the cited cases, not measured win rates.
For maintaining a real codebase, the missing evidence is different: can the system fix a bug, preserve interfaces, avoid regressions, and finish a usable change? A controller or pelican graphic answers little of that. Keep your current model as the baseline and compare on a representative repository task before migrating. Check Astra access for your actual project if it is a candidate; seeing a name in a picker is not the same as validating the route you intend to use.

For a subscription or API purchase, the reports justify interest rather than a premium paid for a presumed hidden model. Google's July statement confirms that Gemini 4 pre-training had started then, but does not identify today's alleged Arena backend. The Gemini catalog and Claude catalog checked on September 20 did not establish public Gemini 4 or Fable 5.2 offerings. Use documented account benefits when spending money, while treating changes in output as observations worth testing.
The early matchup is therefore already interesting: Astra faces credible-looking visual competition in the shared examples, and the suspected next Fable may trade more time and money for better results. The next evidence that would change a decision is a repeatable improvement on the work you actually need done. A launch announcement would clarify naming and access; it would not replace that comparison.



