How a video gets made
Twelve production stages, four specialized AI models, and a check on every image. This is exactly what happens between your idea and the upload.
The twelve stages
The same stages you watch tick off live on a video's page while it's being made.
It learns the topic before it writes a word
Most AI video tools write the script straight from a prompt, from whatever the language model happens to remember. YourFacelessFactory researches first.
- Live web search. Before any writing starts, the topic is researched with live web searches: up to 25 for a long video, fewer for a Short. The goal is real facts, real numbers and fresh examples, not the same five talking points every other video on the topic repeats.
- Reworking a YouTube video. The video's transcript is read to understand its subject and angle. It is never used as the script. About half of its points are kept in completely new words, the other half are replaced with newly researched facts, and the two are woven together so the result stands on its own.
- Starting from an idea. A one-line idea is researched from scratch. Paste a longer brief (notes, an outline, even a draft) and it's treated as your content: the facts in it are checked against search, anything wrong is corrected, and gaps are filled.
Research makes a script far more accurate than writing from memory. No AI is perfect, though, so skim the important claims before you publish.
Written like a real channel, and written for the ear
The script is written with the same structure successful explainer channels use, as an entirely original piece of writing.
- A hook that earns the click. The first seconds lead with a consequence, not a topic, and state what happened while holding back how.
- Loops that keep people watching. The body is built from three to six setup, tension and payoff loops, threaded into one story instead of a list of facts.
- Comparisons you can picture. Big numbers are turned into something physical and relatable, and earlier moments are called back later.
- Made to be spoken. Numbers, ranges and units are written the way a narrator says them, units are converted to metric, and abbreviations are spelled out, so the voice never stumbles over "92x" or "10-15".
- Every line is also a shot. The script marks which moments are story scenes with the narrator, which cut away to a map, an object or a comparison, and which become a bold text card for a twist.
- List videos done properly. "Top 10" style topics are recognized automatically and get a numbered roadmap with a transition into every item. You can also force the list format on or off.
- Pacing you choose. Five pacing levels, from quick cuts to slow, deliberate scenes. The script is written to that pace, then sentences that pack in several ideas are split into separate beats (or short beats are joined, at the calmer levels). Every split is verified word for word, so no narration is ever lost or changed.
A real voice, recorded in natural takes
- Your choice of voice. Narration is recorded with ElevenLabs, in the voice you pick and preview in the template builder.
- Long takes, not clip after clip. Scenes are recorded together in long continuous takes and then cut at the exact word. The delivery flows naturally across scene changes instead of restarting its tone every few seconds.
- Everything is timed to the words. The recording's word-level timing decides how long every scene lasts, when a pop-up appears, when each part of a reveal lands, and where every caption starts and ends.
- The length you asked for. After recording, the real duration is measured. If it falls outside your target, the script is rewritten at a corrected pace, up to four times.
Every line gets a planned shot, not just a picture
- Staged with the whole story in view. A planning pass reads the entire script at once and stages every scene: where the character stands and which way they face, what the background shows, and which object deserves the emphasis. It first writes down what each image is actually meant to show, so a figure of speech gets drawn as what it means, not literally.
- A second pass for continuity. A separate review re-reads the finished plans and fixes drift, such as a scene contradicting an earlier one or a room that resets after something dramatic just happened in it.
- Call-outs that point at things. The first time a named person, place or object appears on screen, a labeled arrow points right at it, timed to the moment it's mentioned. Each one is checked against what the picture actually shows, and against text already on screen, so labels land on the right thing. A pop sound is optional, and it can be your own.
- A camera with intent. The camera moves only when the narration calls for it: a slow drift across a place being described, or a push-in on the exact object being named, found in the image by a vision model. Now and then it snaps in on the thing being named, right as it's said, for emphasis (a switch you can turn off). Otherwise it holds still, and two pans in a row never drift the same way.
- Cut-aways that explain. Places get a map, with the year when it matters. Numbers and comparisons build up piece by piece. "Wrong vs. right" moments get a split screen with a red X and a green check, and two related things are joined by an arrow that names the relationship.
- No near-duplicate drawings. When a moment simply continues the previous shot, the same picture can carry on with a slow pan instead of a near-copy being drawn. It's optional, and capped at a quarter of the scenes.
One hand-drawn style. One cast. Every scene.
- Drawn from scratch, line by line. Every scene is illustrated for the exact line being spoken. There's no stock footage and no clip-art library.
- A strict style rulebook. Flat solid colors, bold black outlines, clean white heads, simple two-to-three-shape backgrounds and consistent sky colors, with the focal subject drawn in extra detail and color so it pops. Every single image request carries the same rulebook.
- Your narrator, always on model. The main character is drawn from reference images matched to the angle each shot needs, or from your own uploaded character.
- Every important character gets a locked look. Before their first scene, each important character gets their own reference portrait. Real people from history are researched for their clothing, headwear and signature props. Recurring groups get a group design sheet, and recurring objects and animals are locked after their first appearance.
- Build-ups that grow. When the narration adds to something, like armor piece by piece or a fortress wall by wall, the drawing is edited one step at a time instead of redrawn, so it stays the same drawing as it grows.
- List videos with a real roadmap. Each item is drawn as its own icon and placed on an exact grid. The camera zooms to each item as it's introduced, and the backdrop grows darker item by item as the list builds toward its most extreme entry.
It checks its own work
Generating an image is the easy part. Knowing it's the right image is the hard part, so every image is checked by a different AI model than the one that drew it.
- The right subject. Does the picture show the specific thing that was asked for, or a vague, generic stand-in?
- The right character. Does a recurring character match their locked reference: hair, head shape, build and clothing color?
- The right style. Does it follow the style rules, and did it borrow only the reference's technique rather than copying its scenery?
- Fixed, not re-rolled. A real mistake is corrected with a targeted edit of that same image, keeping everything else and changing only what's wrong. That's far more reliable than starting over.
- Stops before it wastes your tokens. If five scenes in a row fail, something is systematically wrong, so the run stops instead of spending your tokens on a video that won't come out right. And a video that fails on our end is never charged.
Finished, not almost finished
- Rendered at full size. Pictures, camera moves, pop-ups, reveals and narration are rendered into one video at the platform's own recommended size: 1920×1080 landscape, 1080×1920 vertical, 1080×1080 square or 1080×1350 portrait.
- Captions timed to the voice. Optional burned-in captions follow the narration word by word: nine animated styles to start from (a karaoke box, a highlighted spoken word, pop-ins and more), and every detail adjustable.
- Thumbnails built to be different. Optionally, several thumbnail concepts are planned together so they don't all say the same thing, drawn at a higher resolution, and checked for correctly spelled title text. Upload a reference and they'll follow its look.
- Ready for YouTube. A title, a description, five hashtags and chapter timestamps are written from the video's own script, with chapters taken from the real scene timings.
- Download or autopilot. Download one zip with the video, title, description, thumbnails and narration audio, or link the template to a schedule. Scheduled videos are made up to a day ahead and posted to your channel right on time.
Who does what
Each job goes to whatever is best at it, instead of one model trying to do everything.
| Job | Done by |
|---|---|
| Research, script writing, scene planning, continuity review, pop-up and camera planning, titles and descriptions | Anthropic's Claude, with live web search for research |
| Drawing every scene, character portrait and thumbnail | Google's Gemini image model |
| Checking every image | A separate Google Gemini vision model |
| Narration and word-level timing | ElevenLabs |
| Finding where sentences can be split | A grammar parser: rules, not guesses |
| Rendering, layouts, arrows, maps, captions | Our own renderer |
| Posting (only if you connect a channel) | YouTube's official upload API |
What you control
Set it once in a template, then reuse it for every video.
- SourceA YouTube link, a one-line idea, or your own detailed brief
- Format16:9, 9:16, 1:1 or 4:5, with thumbnails matching or their own shape
- Length & pacingYour target length and five pacing levels
- Voice & languagePick and preview the narrator's voice. English today, more languages coming soon
- Art styleHand-drawn today, with more styles on the way
- CharactersOur narrator or your own uploaded character, plus side characters
- StructureIntro, outro and list format
- FramesVisual cut-aways, comparison layouts and text cards, each on or off
- Motion & call-outsCamera motion, shot reuse, pop-ups and their sound
- CaptionsNine animated styles; font, size, colors, highlight, animation and position
- ThumbnailsHow many, which shape, and a reference image for them to follow
- PublishingTitle and description, your YouTube channel and a schedule
See it for yourself
Build a template and look around the dashboard for free. You'll only need an account when you generate a video.