Seedance 2.0 Prompting

Write Seedance prompts with a four-axis camera codec — distance, height, bearing, lens — so movement becomes coordinates instead of adjectives. Use when prompting a video model and the camera has to do something specific.

We wrote the same shot two ways and sent both to Seedance. Same model, same run, 76 characters apart.

Both generated on seedance-2.0-text-to-video. Left: written by feel. Right: the same shot through the four-axis codec. One generation each, not a repeated trial.

Watch what each one assembles into. The left-hand generation invents a photorealistic mountain landscape and mounts it on four curved panels — there is nothing about storyboard pages in the result, because nothing in the prompt told the camera or the subject what to do. The right-hand one assembles the pages into a continuous strip with the hand-drawn panels still legible, in the dawn light the prompt asked for.

The shot was chosen to be impossible without camera movement: scattered storyboard pages lift off a studio floor, arrange themselves into a timeline, and assemble into one image. If the movement is described badly, the result collapses.

  • A · written by feel — 241 characters, zero camera verbs, and three adjectives — “cinematic”, “high-end”, “beautiful light” — of exactly the kind the skill opens by criticising. The model is left to guess how the camera moves.
  • B · four-axis codec — 1,009 characters, three camera verbs with explicit start and end coordinates: dolly out, crane up, rack focus, resolving into one crane move.
  • C · minimal — 317 characters, the same three camera verbs. Barely longer than A, and still unambiguous.

The gap between A and C is the argument: C is seventy-six characters longer than A, and every one of those characters buys a camera instruction instead of an adjective.

The three prompts

Read A and C back to back — that is the whole case in seventy-six characters. One note: the skill is written for a Chinese-language workflow and its rules say prompts should be Chinese, with the 2000-character ceiling counted in Chinese characters. These are English; the derivation method is what transfers, and English prompts sit well inside the same ceiling.

A · written by feel — 241 characters

An empty studio with storyboard pages scattered on the floor. The pages slowly float up, line themselves into a row, and finally assemble into one complete picture. Make it cinematic, high-end, with beautiful light. 6 seconds, vertical 9:16.

B · four-axis codec — 1,009 characters

An empty studio at dawn, dozens of hand-drawn storyboard pages scattered across a wooden floor.

The camera starts at floor height in a worm's-eye angle, 24mm wide, three-quarter side-on and close enough to read the paper fibre of a single page, everything beyond it thrown out by a shallow depth of field.

The edges of the pages tremble, then the whole stack lifts off the ground. The camera performs a crane move in sync: dollying out from close-up to wide while craning up from the worm's-eye angle to eye level, the bearing rotating from three-quarter side to straight on.

Part-way through the rise it racks focus, the focal point travelling from the nearest page to the queue of pages forming in the distance, that queue arranged into a leading line receding into depth.

In the final second all the pages align dead ahead and assemble into one complete picture, edges seamless, dust motes suspended in the dawn light.

No text, no watermark, no logo. No people on camera.
6 seconds, vertical 9:16, 2K.

C · minimal — 317 characters

Empty studio, scattered storyboard pages float up into a timeline and assemble into one picture.
Crane move: worm's-eye close-up dollies out and cranes up to an eye-level wide, three-quarter side rotating to straight on, racking focus part-way.
24mm wide, dawn light, dust motes. No people, no text. 6s vertical 9:16.

The skill publishes measured limits for the minimal strategy: roughly a quarter pass, half acceptable, a quarter fail — and it says outright that reaction shots (surprise, panic, tension) need sub-beats and must not be minimised. Our shot is object motion, not a reaction, which is why C works here.

What it is

A prompting system, not a generator. It decomposes camera movement into four look-up-table axes — Z distance, Y height, X bearing, F lens capability — so a move becomes a pair of coordinates instead of an adjective. Each axis position maps to a named move with an English equivalent, which is what makes the prompt legible to the model.

It carries 31 worked examples, concentrated on live-action social video: concerts, karaoke, kitchen shorts, vlogs, stadium crowds.

What to know

It writes prompts, it does not make video. You still need Seedance access to render anything. The skill itself costs nothing and requires no keys.

There is an overlapping skill. Seedance Storyboard targets the same model. We compared them directly: the platform specifications are identical and its camera guidance is an order of magnitude thinner. Install this one. That page explains the one case where the other is still worth raiding.

Related

For deciding what the shots should be before you prompt them, see Video Spec Builder.

Installing

Not ours — install from the source repository. MIT.

相关技能

探索更多 →