How to make a cartoon video in 2026, step by step
A practical walkthrough for making a 30-second cartoon video from a written idea: writing the prompt, choosing a style, fixing common problems, and exporting for social.
Ten years ago, a thirty-second cartoon was a two-week job. You needed a script, a storyboard, character designs, a rig for every character, a pass for backgrounds, an animator to key the motion, a compositor to assemble the layers, and a sound designer to make any of it feel alive. Studios charged four figures a minute and they were not overcharging — that is genuinely how much labour was in the frame.
That pipeline still exists, and for a feature film it still matters. But for the thing most people actually need — a short, clear, charming animated clip to explain a product, open a video, tell a joke, or hold a child's attention for half a minute — the work has collapsed into a text box and a wait of a few minutes. This guide is the practical version of how that works now, written for someone who has never opened animation software and does not intend to.
What you need before you start
Almost nothing, which is the point. You need an idea you can describe in a couple of sentences, and a browser. You do not need a drawing tablet, a subscription to a compositing suite, a microphone, or a stock music licence. A modern cartoon video creator handles the animation, the voices, the sound effects, and the music in a single generation.
What genuinely helps, and what most people skip, is spending sixty seconds deciding three things before you type anything:
- Who is in the shot. One character is dramatically easier to get right than three. If your idea has a crowd, pick the one character the viewer should watch and let the rest be background.
- What single thing happens. Thirty seconds is roughly one beat. A character wants something, tries, and either gets it or comically does not. That is the whole story. Two beats in thirty seconds reads as rushed.
- Where it happens. A specific location does more work than a specific adjective. "A cluttered garage workshop at night" gives the model lighting, props, colour and mood in five words.
Step 1: Write the prompt like a shot description
The single biggest quality jump available to a beginner is switching from describing a topic to describing a shot. These two prompts produce wildly different results:
Weak: "A cartoon about recycling."
Strong: "A round, cheerful robot with mismatched eyes sorts glass bottles into a blue bin in a sunlit suburban kitchen. It picks up a bottle, squints at it, tosses it in, and gives a proud thumbs-up to the camera. Warm 2D cartoon style, thick outlines, pastel palette, gentle bouncing motion."
The second one names a subject, an action, a setting, a camera relationship, and a visual style. That is the full checklist. If your prompt is missing one of those five, the model will invent it, and it will invent something generic.
A reliable template to start from:
[character with 2–3 physical details] [does a specific action] in [specific location with lighting]. [Camera behaviour]. [Style words]. [Mood or pacing].
If you want a much longer set of these, we keep a library of tested cartoon video prompts organised by use case.
Say the dialogue out loud, in quotes
Seedance 2.5 generates audio alongside the picture, and it treats quoted text as spoken dialogue. Writing The fox leans in and whispers: "You did not see anything." produces a lip-synced whisper. Writing The fox says something secretive produces mumbling. Quotation marks are the switch.
Keep spoken lines short. Roughly two to three seconds of speech per line, and no more than about three lines in a thirty-second clip, or the pacing turns into a monologue with no room for animation.
Step 2: Choose a style deliberately
"Cartoon" is not one look. The model will happily give you a different one every run unless you pin it down. The five style families that behave most predictably:
| Style | Prompt words that trigger it | Best for |
|---|---|---|
| Modern 2D flat | flat vector cartoon, bold outlines, limited palette | Explainers, product demos, ads |
| Storybook / watercolour | soft watercolour, paper texture, hand-painted | Kids' content, gentle narration |
| 3D toon render | 3D toon shading, rounded shapes, soft studio light | Mascots, app characters, tech |
| Anime | cel-shaded anime, expressive eyes, speed lines | Story clips, gaming, drama |
| Retro rubber-hose | 1930s rubber hose cartoon, black and white, film grain | Comedy, nostalgia, music |
Pick one family and use its words consistently. Mixing "anime" and "watercolour" in the same prompt usually produces something that commits to neither. There is a longer breakdown in our guide to cartoon animation styles.
Step 3: Decide how the clip starts
You have three genuinely different starting points, and choosing the right one is most of the battle.
Start from text
Type a description, get a cartoon. Best when you have no existing assets and want the model to have full freedom over composition. This is the fastest route and the one most people should use first.
Start from an image
Upload a picture and the model animates it, treating your image as the opening frame. This is the right choice when you already have a character design, a logo, a product photo, or a drawing you want to keep. Because the first frame is fixed, the character stays recognisably itself instead of drifting.
You can also supply a last frame as well as a first frame. The model then animates the journey between the two images, which is the cleanest way to control exactly where a shot ends up — useful for transformations, before-and-after reveals, and logo stings.
Start from a reference
Supply reference images of a character, a prop, and a location, and the model carries those identities into a brand-new scene. This is how you make episode two look like episode one. If you are building a series or a mascot-led brand, this is the mode that matters, and it is worth reading the comparison of text, image and reference inputs before you commit to a workflow.
Step 4: Set duration, aspect ratio and sound
Three settings, each with an obvious right answer most of the time.
- Duration. Thirty seconds is the maximum in a single generation and it is a genuinely different creative unit from the five-second clips earlier models produced — there is room for a setup and a payoff. If your idea is a single visual gag, eight to twelve seconds will feel tighter and cost less.
- Aspect ratio. 9:16 for TikTok, Reels and Shorts. 16:9 for YouTube, websites and presentations. 1:1 for feed posts. Choose before you generate; cropping afterwards throws away the composition the model built.
- Audio. Leave it on unless you are cutting the clip into an existing edit with its own soundtrack. Generated audio includes ambience and effects, not just voice, and it does more for perceived quality than an extra resolution tier.
Step 5: Watch it once, then fix one thing
The temptation after a first generation is to rewrite the whole prompt. Resist it. Change one variable and regenerate, or you will never learn which word did what. The most common fixes, in the order they usually come up:
| Problem | What actually causes it | Fix |
|---|---|---|
| Character changes appearance mid-clip | Description too vague to hold identity | Add 2–3 fixed physical details, or supply a first-frame image |
| Everything drifts and floats | No anchored camera | Add "static camera" or "locked-off shot" |
| Feels empty and slow | Thirty seconds asked to carry one static idea | Shorten to 10–15s, or add a second action beat |
| Mouth movement does not match speech | Dialogue not in quotes | Put spoken lines in double quotes |
| Style looks generic | Only the word "cartoon" given | Name a style family and a palette |
| Text in the scene is garbled | Rendered lettering is unreliable in any model | Add captions in your editor afterwards |
Step 6: Export and finish
Download the MP4 and it is ready to post. Two optional finishing touches are worth the extra two minutes:
Add burned-in captions if the clip has dialogue. The large majority of social video is watched muted, and captions reliably lift completion rate more than any visual change you could make to the animation itself.
Trim the first quarter-second if the clip opens on a held frame. Generated video sometimes settles for a few frames at the start; cutting them makes the opening feel snappier, which matters enormously in a feed.
A realistic sense of what this does and does not do
Being straight about the limits saves you a frustrating afternoon. AI cartoon generation is excellent at short, self-contained, character-led shots with clear motion. It is not yet a replacement for a studio pipeline when you need frame-exact timing to a pre-recorded voice track, a character who must be pixel-identical across fifty shots, or on-screen text rendered correctly.
What it is genuinely, decisively better at is the thing that used to be impossible: making twelve versions of an idea on a Tuesday afternoon to find out which one works. The iteration speed is the feature. Most people get their best result on the third or fourth attempt, not the first, and the whole exercise still takes less than half an hour.
Try it with one sentence
The fastest way to understand any of this is to generate something. Open the cartoon video creator, paste one of the strong prompts above, set thirty seconds, and watch what comes back. Then change exactly one word and go again — that second run is where the intuition starts.
Try it while it is fresh
Free credits on sign-up. Up to 30 seconds with sound.
Related guides
40 Cartoon Video Prompts That Actually Work (Copy & Paste)
Tested prompt templates for AI cartoon videos, grouped by use case: explainers, mascots, kids' stories, ads, comedy and title cards — plus the prompt structure behind them.
12 Cartoon Animation Styles Explained (With Prompt Words)
A visual vocabulary for cartoon video: rubber hose, flat vector, cel-shaded anime, storybook watercolour, 3D toon and more — what each style signals, and the exact words that trigger it.
Text, Image or Reference: Which Cartoon Input Should You Use?
Text-to-video, image-to-video, first-and-last-frame and multimodal reference each solve a different problem. Here is when to use each one, and what each gives up.