Two ways to start: a topic or a script
Give a topic and the maker writes the script for you. It researches nothing and relies on general knowledge, so it stays strongest on evergreen subjects: how things work, why things happen, lists, comparisons, reframes and stories. It writes a hook first, keeps sentences short so captions fit, and ends with a payoff instead of a call to action.
Give a script and the maker becomes an assembler rather than a writer. Your wording is preserved and only split into beats of a sentence or two, each paired with a picture brief that illustrates that moment. Light grammar fixes are the only edits. This is the right route for anything you have already written carefully, or anything that has to be accurate word for word.
What makes text read well as a short
Shorts are heard before they are read, so sentences that work on a page can trip over themselves out loud. The beat writer favours concrete nouns, one idea per sentence and numbers spelled the way people say them. If you paste your own text, the same rules apply: short sentences caption cleanly in groups of three words with the spoken word highlighted, while long clauses produce captions that lag behind the voice.
Length matters too. A thirty-second short holds roughly seventy words; forty-five seconds about a hundred and ten; sixty seconds about a hundred and fifty. Paste more than that and the maker will trim to fit the length you chose, so if you want every word kept, choose sixty seconds or cut the text first.
From words to picture briefs
Every beat gets a picture brief written alongside the narration: a fifteen to thirty word description of one scene with a subject, a setting, lighting and mood. The briefs never ask for text, logos or real people, and they are run through the same moderation gate as your input before any image is rendered. Your chosen look is then appended so all the images share a palette and medium.
This matters because a short is really a slideshow with intent. A beat about pressure at ocean depth should show a crushed submersible in blue dark, not a stock photo of a beach. Writing the brief from the narration, beat by beat, is how the pictures stay on-message without you art-directing anything.
Timing, captions and the final file
After the voice is generated the full narration is transcribed with word-level timestamps, and captions are rendered from those times rather than from guesses about reading speed. Each caption shows a few words at a time with the current word in the style's accent colour, placed in the lower third where platform interfaces leave room. The output is a 1080×1920 MP4 at thirty frames per second with AAC audio and fast-start headers, so it previews quickly wherever you upload it.