Text vs image vs reference: choosing the right input for your cartoon
Text-to-video, image-to-video, first-and-last-frame and multimodal reference each solve a different problem. Here is when to use each one, and what each gives up.
There is a moment, usually around your fifth or sixth generation, where you stop asking "why does this look wrong" and start asking "why does this look different every time". That is the moment to stop typing longer prompts and change your input mode instead.
A modern cartoon video creator accepts several kinds of starting material, and each one trades a different amount of creative freedom for a different amount of control. Picking the right one is usually a bigger quality lever than any amount of prompt refinement.
The four modes at a glance
| Mode | You supply | You control | You give up |
|---|---|---|---|
| Text to video | A description | Concept, style, action | Exact appearance, framing |
| Image to video | One opening image + prompt | Exactly how it starts | Freedom of composition |
| First & last frame | Two images + prompt | Start and end state | Aspect ratio, surprise |
| Reference to video | Character/prop images + prompt | Identity across scenes | Setup time |
Text to video: maximum freedom, minimum control
You describe a scene; the model invents everything. This is the mode to use when you are exploring — when you do not yet know what the video should look like and you want to see options.
Use it when: you have no existing artwork, you are testing an idea, you want variety across runs, or the character does not need to reappear anywhere else.
Its real weakness is identity drift. Nothing anchors the character's appearance except your words, so two runs of the same prompt produce two different characters — and within a single long clip, a character described loosely can subtly change as the shot progresses. Fixed physical details help ("one blue eye and one green eye, a dented left shoulder") but they are guidance, not a guarantee.
Practical tip: text-to-video is also the only mode where you get free choice of aspect ratio. Once you attach an image or a source video, the model inherits that asset's shape.
Image to video: the opening frame is yours
You upload one picture and the model treats it as the first frame, then animates forward from it. Everything visible in that image — the character design, the palette, the composition, the lighting — is locked at the start and tends to persist.
Use it when: you already have a character design, a logo, a product shot, a piece of concept art, or a child's drawing you want to bring to life. Also use it when text-to-video keeps producing the right idea in the wrong style: generate a still you like, then animate that.
What the prompt does here changes completely. It no longer describes what things look like — the image already did that. It describes what happens. "The fox turns its head toward the window and its ears twitch" is a good image-to-video prompt. "A cartoon fox in a forest" is a wasted one, because you have just described the picture you already supplied.
Constraint to know: the output inherits the aspect ratio of your image. If you need a 9:16 vertical video, crop your source image to 9:16 before uploading. You cannot override the ratio in this mode.
First and last frame: controlling the destination
Supply two images — one tagged as the first frame, one as the last — and the model animates a plausible journey between them. This is the most underused mode and the one that solves the most annoying problem in generated video: the ending.
Ordinary generation gives you a beginning you chose and an ending you did not. For anything that has to land on something specific, that is a real limitation. First-and-last-frame flips it.
Use it when:
- Logo stings. Start on chaos, end on your locked-up logo, exactly as your brand guidelines draw it.
- Transformations. Caterpillar to butterfly, messy desk to tidy desk, sketch to finished art. The model invents the middle, which is the hard part, while you control both ends.
- Before and after. Product demos where the payoff has to be accurate rather than imagined.
- Loops. Use the same image as both first and last frame and you get a clip that returns to its start — seamless for a background loop or a repeating social post.
- Stitching longer sequences. Generate clip one, take its final frame, and use it as the first frame of clip two. Do that repeatedly and you can build a continuous piece far longer than any single generation, with no visible jump at the joins.
Constraints: aspect ratio is inherited from the first-frame image, and the two images need to be plausibly connected. Asking the model to travel from a watercolour forest to a pixel-art spaceship in thirty seconds will produce something incoherent — it interpolates, it does not teleport.
Reference to video: identity across many scenes
Instead of one frame, you supply a set of reference images — a character from a few angles, a prop, a location — and the model carries those identities into a brand-new shot that none of the references depict.
This is the difference between making a video and making a series. With references, episode four can star recognisably the same character as episode one, in a place you never photographed, doing something you never drew.
Use it when: you are building a mascot-led brand, a recurring cast, a set of ads that must feel related, or any body of work where consistency is the point.
How to prepare good references:
- Show the character clearly and large in frame, against a plain background where possible.
- Two or three angles beat one. A front view plus a three-quarter view resolves most ambiguity about a face.
- Keep lighting neutral. A strongly coloured light in the reference gets read as part of the character's colour scheme.
- Reference the character in your prompt by a short consistent name — "the orange cat mascot" — so the text and the images agree about who is who.
Seedance 2.5 accepts a large number of reference assets in a single generation, so you can supply a character, a costume, a prop and an environment together rather than choosing between them.
A decision path
In practice, most people should follow this order:
- Start with text. Run three or four quick generations to find the concept and style that work. This is exploration; do not polish yet.
- When one run is nearly right, capture a frame. Pull a still you like from that generation and switch to image-to-video, using it as the first frame. You have now locked the look.
- If the ending matters, add a last frame. Especially for anything ending on a logo, a product or a specific pose.
- If you need more than one video, build references. Once you have a character you want to keep, collect two or three clean stills of it and move to reference mode for everything after.
The mistake almost everyone makes is trying to do step four's job with step one's tools — writing ever more elaborate text prompts hoping to pin down a character that only an image can pin down. Text describes; images specify. Reach for the right one.
What none of the modes fix
Two limitations survive every input mode, and it is worth knowing them before you plan a project around them. Rendered text inside the frame — signage, labels, captions — is unreliable in every current model; add it afterwards in an editor. And precise frame-level timing to a pre-recorded voiceover is not yet something you can specify; if your clip must hit a beat at 00:14, plan to edit the audio to the video rather than the other way round.
Everything else is now a question of choosing the right starting material. Try the same idea through two different modes in the cartoon video creator and the difference in control will be obvious within one generation each.
Try it while it is fresh
Free credits on sign-up. Up to 30 seconds with sound.
Related guides
How to Make a Cartoon Video in 2026 (No Drawing Needed)
A practical walkthrough for making a 30-second cartoon video from a written idea: writing the prompt, choosing a style, fixing common problems, and exporting for social.
40 Cartoon Video Prompts That Actually Work (Copy & Paste)
Tested prompt templates for AI cartoon videos, grouped by use case: explainers, mascots, kids' stories, ads, comedy and title cards — plus the prompt structure behind them.
Seedance 2.5 for Cartoons: What 30 Seconds Actually Changes
Seedance 2.5 generates up to 30 seconds with synchronised audio in a single pass. Here is what that unlocks for cartoon makers, and how to plan a clip that uses the length well.