Skip to content
Cartoon Video Creator
Getting started

How to make a cartoon video in 2026, step by step

A practical walkthrough for making a 30-second cartoon video from a written idea: writing the prompt, choosing a style, fixing common problems, and exporting for social.

9 min read1,578 words

Ten years ago, a thirty-second cartoon was a two-week job. You needed a script, a storyboard, character designs, a rig for every character, a pass for backgrounds, an animator to key the motion, a compositor to assemble the layers, and a sound designer to make any of it feel alive. Studios charged four figures a minute and they were not overcharging — that is genuinely how much labour was in the frame.

That pipeline still exists, and for a feature film it still matters. But for the thing most people actually need — a short, clear, charming animated clip to explain a product, open a video, tell a joke, or hold a child's attention for half a minute — the work has collapsed into a text box and a wait of a few minutes. This guide is the practical version of how that works now, written for someone who has never opened animation software and does not intend to.

What you need before you start

Almost nothing, which is the point. You need an idea you can describe in a couple of sentences, and a browser. You do not need a drawing tablet, a subscription to a compositing suite, a microphone, or a stock music licence. A modern cartoon video creator handles the animation, the voices, the sound effects, and the music in a single generation.

What genuinely helps, and what most people skip, is spending sixty seconds deciding three things before you type anything:

  • Who is in the shot. One character is dramatically easier to get right than three. If your idea has a crowd, pick the one character the viewer should watch and let the rest be background.
  • What single thing happens. Thirty seconds is roughly one beat. A character wants something, tries, and either gets it or comically does not. That is the whole story. Two beats in thirty seconds reads as rushed.
  • Where it happens. A specific location does more work than a specific adjective. "A cluttered garage workshop at night" gives the model lighting, props, colour and mood in five words.

Step 1: Write the prompt like a shot description

The single biggest quality jump available to a beginner is switching from describing a topic to describing a shot. These two prompts produce wildly different results:

Weak: "A cartoon about recycling."
Strong: "A round, cheerful robot with mismatched eyes sorts glass bottles into a blue bin in a sunlit suburban kitchen. It picks up a bottle, squints at it, tosses it in, and gives a proud thumbs-up to the camera. Warm 2D cartoon style, thick outlines, pastel palette, gentle bouncing motion."

The second one names a subject, an action, a setting, a camera relationship, and a visual style. That is the full checklist. If your prompt is missing one of those five, the model will invent it, and it will invent something generic.

A reliable template to start from:

[character with 2–3 physical details] [does a specific action] in [specific location with lighting]. [Camera behaviour]. [Style words]. [Mood or pacing].

If you want a much longer set of these, we keep a library of tested cartoon video prompts organised by use case.

Say the dialogue out loud, in quotes

Seedance 2.5 generates audio alongside the picture, and it treats quoted text as spoken dialogue. Writing The fox leans in and whispers: "You did not see anything." produces a lip-synced whisper. Writing The fox says something secretive produces mumbling. Quotation marks are the switch.

Keep spoken lines short. Roughly two to three seconds of speech per line, and no more than about three lines in a thirty-second clip, or the pacing turns into a monologue with no room for animation.

Step 2: Choose a style deliberately

"Cartoon" is not one look. The model will happily give you a different one every run unless you pin it down. The five style families that behave most predictably:

StylePrompt words that trigger itBest for
Modern 2D flatflat vector cartoon, bold outlines, limited paletteExplainers, product demos, ads
Storybook / watercoloursoft watercolour, paper texture, hand-paintedKids' content, gentle narration
3D toon render3D toon shading, rounded shapes, soft studio lightMascots, app characters, tech
Animecel-shaded anime, expressive eyes, speed linesStory clips, gaming, drama
Retro rubber-hose1930s rubber hose cartoon, black and white, film grainComedy, nostalgia, music

Pick one family and use its words consistently. Mixing "anime" and "watercolour" in the same prompt usually produces something that commits to neither. There is a longer breakdown in our guide to cartoon animation styles.

Step 3: Decide how the clip starts

You have three genuinely different starting points, and choosing the right one is most of the battle.

Start from text

Type a description, get a cartoon. Best when you have no existing assets and want the model to have full freedom over composition. This is the fastest route and the one most people should use first.

Start from an image

Upload a picture and the model animates it, treating your image as the opening frame. This is the right choice when you already have a character design, a logo, a product photo, or a drawing you want to keep. Because the first frame is fixed, the character stays recognisably itself instead of drifting.

You can also supply a last frame as well as a first frame. The model then animates the journey between the two images, which is the cleanest way to control exactly where a shot ends up — useful for transformations, before-and-after reveals, and logo stings.

Start from a reference

Supply reference images of a character, a prop, and a location, and the model carries those identities into a brand-new scene. This is how you make episode two look like episode one. If you are building a series or a mascot-led brand, this is the mode that matters, and it is worth reading the comparison of text, image and reference inputs before you commit to a workflow.

Step 4: Set duration, aspect ratio and sound

Three settings, each with an obvious right answer most of the time.

  • Duration. Thirty seconds is the maximum in a single generation and it is a genuinely different creative unit from the five-second clips earlier models produced — there is room for a setup and a payoff. If your idea is a single visual gag, eight to twelve seconds will feel tighter and cost less.
  • Aspect ratio. 9:16 for TikTok, Reels and Shorts. 16:9 for YouTube, websites and presentations. 1:1 for feed posts. Choose before you generate; cropping afterwards throws away the composition the model built.
  • Audio. Leave it on unless you are cutting the clip into an existing edit with its own soundtrack. Generated audio includes ambience and effects, not just voice, and it does more for perceived quality than an extra resolution tier.

Step 5: Watch it once, then fix one thing

The temptation after a first generation is to rewrite the whole prompt. Resist it. Change one variable and regenerate, or you will never learn which word did what. The most common fixes, in the order they usually come up:

ProblemWhat actually causes itFix
Character changes appearance mid-clipDescription too vague to hold identityAdd 2–3 fixed physical details, or supply a first-frame image
Everything drifts and floatsNo anchored cameraAdd "static camera" or "locked-off shot"
Feels empty and slowThirty seconds asked to carry one static ideaShorten to 10–15s, or add a second action beat
Mouth movement does not match speechDialogue not in quotesPut spoken lines in double quotes
Style looks genericOnly the word "cartoon" givenName a style family and a palette
Text in the scene is garbledRendered lettering is unreliable in any modelAdd captions in your editor afterwards

Step 6: Export and finish

Download the MP4 and it is ready to post. Two optional finishing touches are worth the extra two minutes:

Add burned-in captions if the clip has dialogue. The large majority of social video is watched muted, and captions reliably lift completion rate more than any visual change you could make to the animation itself.

Trim the first quarter-second if the clip opens on a held frame. Generated video sometimes settles for a few frames at the start; cutting them makes the opening feel snappier, which matters enormously in a feed.

A realistic sense of what this does and does not do

Being straight about the limits saves you a frustrating afternoon. AI cartoon generation is excellent at short, self-contained, character-led shots with clear motion. It is not yet a replacement for a studio pipeline when you need frame-exact timing to a pre-recorded voice track, a character who must be pixel-identical across fifty shots, or on-screen text rendered correctly.

What it is genuinely, decisively better at is the thing that used to be impossible: making twelve versions of an idea on a Tuesday afternoon to find out which one works. The iteration speed is the feature. Most people get their best result on the third or fourth attempt, not the first, and the whole exercise still takes less than half an hour.

Try it with one sentence

The fastest way to understand any of this is to generate something. Open the cartoon video creator, paste one of the strong prompts above, set thirty seconds, and watch what comes back. Then change exactly one word and go again — that second run is where the intuition starts.

Try it while it is fresh

Free credits on sign-up. Up to 30 seconds with sound.

Make a cartoon

Related guides

Make the cartoon you just read about

Describe a scene, pick a length up to 30 seconds, and watch it animate with sound.

Make a cartoon free