For months this project told agents one thing about text in generated images: never trust it. We tested whether Grok Imagine Image 2.0 changes that — at three difficulty levels, comparing every rendered string character by character against the brief.
Required: NIGHT MARKET · SATURDAY 8PM · PIER 32 — exact, middot included
Same brief, different render — exact again. All three variants were correct.
Five separate strings, in perspective, down to 6pt-equivalent small print: HARBOUR ROAST, Single Origin Ethiopia, Medium Roast, 340g, Roasted in Lisbon · Best before 03/2027 — all exact.
reading exactly: … and it comes back character for character — on a flat poster and on a photorealistic packshot with perspective and small print.
Every label is right — title, four region names twice over, four percentages, the axis label. Every bar is not.
| Region | Label says | Bar drawn at | |
|---|---|---|---|
| North America | 42% | ~41.5% | ✓ |
| Europe | 27% | ~25% | ✗ |
| Asia Pacific | 23% | ~20.5% | ✗ |
| Latin America | 8% | ~8% | ✓ |
sensenova-u1-infographic or render the chart deterministically and composite it.flux-schnell is ~2s. This is a deliberate pick for text-critical work, never a default.num_images bills per image.