Skip to content
Cartoon Video Creator
Craft

Text vs image vs reference: choosing the right input for your cartoon

Text-to-video, image-to-video, first-and-last-frame and multimodal reference each solve a different problem. Here is when to use each one, and what each gives up.

10 min read1,309 words

There is a moment, usually around your fifth or sixth generation, where you stop asking "why does this look wrong" and start asking "why does this look different every time". That is the moment to stop typing longer prompts and change your input mode instead.

A modern cartoon video creator accepts several kinds of starting material, and each one trades a different amount of creative freedom for a different amount of control. Picking the right one is usually a bigger quality lever than any amount of prompt refinement.

The four modes at a glance

ModeYou supplyYou controlYou give up
Text to videoA descriptionConcept, style, actionExact appearance, framing
Image to videoOne opening image + promptExactly how it startsFreedom of composition
First & last frameTwo images + promptStart and end stateAspect ratio, surprise
Reference to videoCharacter/prop images + promptIdentity across scenesSetup time

Text to video: maximum freedom, minimum control

You describe a scene; the model invents everything. This is the mode to use when you are exploring — when you do not yet know what the video should look like and you want to see options.

Use it when: you have no existing artwork, you are testing an idea, you want variety across runs, or the character does not need to reappear anywhere else.

Its real weakness is identity drift. Nothing anchors the character's appearance except your words, so two runs of the same prompt produce two different characters — and within a single long clip, a character described loosely can subtly change as the shot progresses. Fixed physical details help ("one blue eye and one green eye, a dented left shoulder") but they are guidance, not a guarantee.

Practical tip: text-to-video is also the only mode where you get free choice of aspect ratio. Once you attach an image or a source video, the model inherits that asset's shape.

Image to video: the opening frame is yours

You upload one picture and the model treats it as the first frame, then animates forward from it. Everything visible in that image — the character design, the palette, the composition, the lighting — is locked at the start and tends to persist.

Use it when: you already have a character design, a logo, a product shot, a piece of concept art, or a child's drawing you want to bring to life. Also use it when text-to-video keeps producing the right idea in the wrong style: generate a still you like, then animate that.

What the prompt does here changes completely. It no longer describes what things look like — the image already did that. It describes what happens. "The fox turns its head toward the window and its ears twitch" is a good image-to-video prompt. "A cartoon fox in a forest" is a wasted one, because you have just described the picture you already supplied.

Constraint to know: the output inherits the aspect ratio of your image. If you need a 9:16 vertical video, crop your source image to 9:16 before uploading. You cannot override the ratio in this mode.

First and last frame: controlling the destination

Supply two images — one tagged as the first frame, one as the last — and the model animates a plausible journey between them. This is the most underused mode and the one that solves the most annoying problem in generated video: the ending.

Ordinary generation gives you a beginning you chose and an ending you did not. For anything that has to land on something specific, that is a real limitation. First-and-last-frame flips it.

Use it when:

  • Logo stings. Start on chaos, end on your locked-up logo, exactly as your brand guidelines draw it.
  • Transformations. Caterpillar to butterfly, messy desk to tidy desk, sketch to finished art. The model invents the middle, which is the hard part, while you control both ends.
  • Before and after. Product demos where the payoff has to be accurate rather than imagined.
  • Loops. Use the same image as both first and last frame and you get a clip that returns to its start — seamless for a background loop or a repeating social post.
  • Stitching longer sequences. Generate clip one, take its final frame, and use it as the first frame of clip two. Do that repeatedly and you can build a continuous piece far longer than any single generation, with no visible jump at the joins.

Constraints: aspect ratio is inherited from the first-frame image, and the two images need to be plausibly connected. Asking the model to travel from a watercolour forest to a pixel-art spaceship in thirty seconds will produce something incoherent — it interpolates, it does not teleport.

Reference to video: identity across many scenes

Instead of one frame, you supply a set of reference images — a character from a few angles, a prop, a location — and the model carries those identities into a brand-new shot that none of the references depict.

This is the difference between making a video and making a series. With references, episode four can star recognisably the same character as episode one, in a place you never photographed, doing something you never drew.

Use it when: you are building a mascot-led brand, a recurring cast, a set of ads that must feel related, or any body of work where consistency is the point.

How to prepare good references:

  • Show the character clearly and large in frame, against a plain background where possible.
  • Two or three angles beat one. A front view plus a three-quarter view resolves most ambiguity about a face.
  • Keep lighting neutral. A strongly coloured light in the reference gets read as part of the character's colour scheme.
  • Reference the character in your prompt by a short consistent name — "the orange cat mascot" — so the text and the images agree about who is who.

Seedance 2.5 accepts a large number of reference assets in a single generation, so you can supply a character, a costume, a prop and an environment together rather than choosing between them.

A decision path

In practice, most people should follow this order:

  1. Start with text. Run three or four quick generations to find the concept and style that work. This is exploration; do not polish yet.
  2. When one run is nearly right, capture a frame. Pull a still you like from that generation and switch to image-to-video, using it as the first frame. You have now locked the look.
  3. If the ending matters, add a last frame. Especially for anything ending on a logo, a product or a specific pose.
  4. If you need more than one video, build references. Once you have a character you want to keep, collect two or three clean stills of it and move to reference mode for everything after.

The mistake almost everyone makes is trying to do step four's job with step one's tools — writing ever more elaborate text prompts hoping to pin down a character that only an image can pin down. Text describes; images specify. Reach for the right one.

What none of the modes fix

Two limitations survive every input mode, and it is worth knowing them before you plan a project around them. Rendered text inside the frame — signage, labels, captions — is unreliable in every current model; add it afterwards in an editor. And precise frame-level timing to a pre-recorded voiceover is not yet something you can specify; if your clip must hit a beat at 00:14, plan to edit the audio to the video rather than the other way round.

Everything else is now a question of choosing the right starting material. Try the same idea through two different modes in the cartoon video creator and the difference in control will be obvious within one generation each.

Try it while it is fresh

Free credits on sign-up. Up to 30 seconds with sound.

Make a cartoon

Related guides

Make the cartoon you just read about

Describe a scene, pick a length up to 30 seconds, and watch it animate with sound.

Make a cartoon free