What to write, in what order, and which of it the model will actually act on.
On-screen text in generated video fails in a specific, predictable way. Naming the exact string is the difference between a sign that reads and a sign that smears.
Every few weeks someone posts a generated clip where a shop sign reads ANTIQEUS or a book cover says a word that is almost a word, and the comments fill up with "the model can't spell." It is a fair description of the symptom and a bad description of the cause, and the difference matters because it points at a completely different fix.
The model is not spelling badly. It is not spelling at all.
Generated video is produced in a compressed latent space and decoded back to pixels. That compression is extremely good at things which are statistically smooth — skin, fabric, foliage, motion. Text is the opposite: a small number of high-contrast, high-frequency shapes where a one-pixel difference changes the meaning entirely. There is no gradient of "almost the letter Q." You are either exactly on the glyph or you are on a shape that reads as wrong.
Add temporal compression and it gets worse, because now the glyph has to be exactly the same shape across every frame. Letters that shimmer are not the model changing its mind; they are the same near-miss re-rolled per frame.
Two predictions fall out of this, and both hold up:
Bigger text works better than small text. More pixels per glyph means the near-miss is a smaller fraction of the letterform. A single word filling a third of the frame is often clean. A paragraph on a page is never clean.
Familiar text works better than novel text. Common words, common logos, and common signage phrases appear often enough in training to be effectively memorised as shapes. Your client's brand name, invented last year, has no such support.
Once you accept that text is a rendering problem, the production decision becomes simple. Put every piece of text in your shot into one of three tiers.
Tier 1 — must be legible and correct. Brand names, prices, legal copy, URLs, anything a viewer will read. Never generate this. Composite it in post. This is not a workaround; it is how the rest of the industry has always done motion graphics, and it gives you font control, kerning, animation timing and a spell-check.
Tier 2 — should read as text but the content does not matter. Background signage, a newspaper in someone's hand, distant shopfronts. Generate it, and write the brief so the text is soft: out of focus, at an angle, partly occluded, or small. Actively ask for it to be indistinct. "Distant illuminated signage, out of focus" gets you the impression of a city without a legibility problem.
Tier 3 — text that should not be there at all. Watermarks, subtitles the model adds unprompted, captions burned in because half the training data had them. This is what a negative or exclusion clause is for, and it is one of the highest-value exclusions you can write: no subtitles, no watermark, no captions.
Most bad text in generated video comes from Tier 1 content being left in the generation, or from Tier 3 content never being excluded.
The objection to "do it in post" is always time. It is worth doing the arithmetic once.
Re-rolling a shot until the sign spells the word correctly takes an unbounded number of attempts, each one billed, each one producing a slightly different shot that has to be re-reviewed for everything else. Compositing a title over a generated plate takes a few minutes in any motion tool and is deterministic.
The only real requirement is that the generation gives you somewhere to put it. So brief for that: a clean wall, a plain surface, a deliberate empty third of the frame. Ask for the space, not for the words. A shot generated with a blank sign is a better asset than a shot generated with a nearly-right sign, because the blank one can be finished and the nearly-right one can only be redone.
The general picture above is architectural and applies broadly. The details — how large is large enough, whether Latin script does better than others, whether the model tends to invent subtitles — vary per model and change with each release, so they are worth checking against something current rather than against a blog post from last year. A page that collects what one model actually does with on-screen text is the kind of reference to look for: specific behaviours, dated, with the failure cases shown rather than described.
If you want a sense of where the boundary sits before you plan a job around it, generate one shot with a large single word and one with a small block of text and compare — most hosted tools will show you the difference in two generations, and the MiniMax H3 AI video generator at minimax-h3ai.video runs the first one without an account.