What to write, in what order, and which of it the model will actually act on.
A prompt field with room for seven thousand characters is not an invitation to write seven thousand characters. These notes are about what belongs in a brief for a generative video model, what belongs somewhere else, and the order that makes the difference.
Most advice about prompting generative video models is a list of adjectives. The three notes here are not that. They are about structure: which information a model can act on, where it has to appear, and what to do with the parts of a brief that no amount of description will fix.
The first takes a specification number seriously. A model that accepts seven thousand characters is not asking for seven thousand characters — it is telling you it expects structured input, and structure has a shape. Knowing the shape is worth more than knowing the limit.
The second is about text inside the picture, which is the single most reliable way to make a generated clip look generated. The failure is not random and it is not a rendering weakness. It has a specific cause and a specific fix, and the fix is one sentence long.
The third is about order. When a model produces picture and sound in one pass rather than two, writing the sound last means writing it into a frame that has already been decided. Writing it first changes the picture. That is not a stylistic preference; it is a consequence of how the generation works.
The reference behind these — field names, caps, the order the model expects them in — is collected at https://minimax-h3ai.video.
Long prompt limits tell you the model expects structure, not volume. Where the useful length actually sits and what to cut first.
On-screen text in generated video fails in a specific, predictable way. Naming the exact string is the difference between a sign that reads and a sign that smears.
When picture and audio come out of one forward pass, the sound is not a layer you add afterwards. What changes in the brief when you write it first.