What a faceless short is, and why people make them
Plenty of the most-watched vertical channels never show a person. A voice tells a story or explains an idea while the screen cycles through pictures, text and motion. The format works because attention on a phone is won in the first second by a strong line and a strong image, not by a presenter. It also means you can publish without a studio, a ring light, or a good hair day.
The catch has always been the labour. A single sixty-second faceless short normally means writing a script, recording a voiceover, hunting for visuals, timing captions by hand, and assembling it all in an editor. This maker collapses that into one request. You describe the short; the service produces the beats, the audio, the images and the edit in a few minutes, server-side, and hands you the file.
How the maker builds your short
Your topic is turned into five to eight beats. A beat is one spoken sentence or two paired with a picture brief that describes exactly what should be on screen while those words are said. The first beat is written as a hook, the last one as a payoff, and the middle beats carry one idea each so the pacing never sags.
Each beat is voiced separately in the voice style you chose, then the clips are joined with a short breath between them. The combined narration is transcribed word by word so every caption lands on the exact moment it is spoken. For the pictures, one vertical AI image is rendered per beat in your chosen visual style, and a slow push-in or drift is applied so the stills feel filmed rather than pasted. Finally the captions are burned in with the current word lit in the style's accent colour, and the whole thing is encoded as an MP4 ready to upload.
If you already have a script, paste it instead of a topic. The service keeps your wording and only splits it into beats, so a script you have polished stays yours line for line.
Looks and voices that suit faceless formats
Faceless channels live or die on consistency, so the ten looks are designed to stay recognisable from short to short: flat neon graphics, cinematic photographic stills, paper-cut dioramas, ink sketches, claymation, retro screen prints, blueprints, dark fantasy paintings, watercolour and pixel art. Pick one and keep using it; viewers will learn to spot your channel in the feed.
There are thirteen voice styles ranging from a deep documentary read to a bright, quick presenter. Each look also carries its own delivery notes, so a dark fantasy short is read slower and lower than a quiz short even with the same voice. You are choosing a voice style, not cloning anyone; nothing here is a real person's voice.
Still images or motion clips
The standard tier animates still images with gentle camera moves, which is what most faceless channels use and costs fifteen credits per short. The motion tier swaps each still for a short AI video clip generated from that image, so smoke drifts, water moves and light flickers. It costs forty credits and takes a few minutes longer. Both tiers produce the same 1080×1920 H.264 file with AAC audio.