Turn a story, novel excerpt, or script into a producible short-drama workflow: beats, characters, scenes, storyboard prep, and video generation handoff. Use when orchestrating script-to-video pipelines, micro-drama production, or adapting narrative source material into short-form video.
Organizes story/novel/script input into a repeatable short-drama production line: story beats, character and scene setup, storyboard prep, and video generation handoff.
Four tasks, each run with the skill installed and again with skills disabled, identical task text. Both arms produced 22 files in total, so volume tells you nothing here. What the two arms actually built is different in kind.
The skill arm assumes the thing will be generated; the control arm assumes it will be shot. Every skill run produced an asset registry with stable IDs — 03_asset_registry.md (1,631 B) for the café scene, 03_人物与场景资产清单.md (2,323 B) for the woodworker, 03_场景道具清单.md (3,186 B) registering five ENV and eight PROP identifiers with a cross-shot consistency table. That registry is the mechanism the skill uses to keep a character's face and a prop's shape stable from beat to beat, and it is followed in every run by a generation handoff (05_generation_handoff.md, 03_generation_handoff_checklist.md, 05_视频生成交接清单.md). The control arm never produced a registry in any of the four. It produced a casting brief, a shooting schedule, an equipment-and-budget list, and a risk-and-contingency sheet instead — a call sheet for a crew.
Neither framing is wrong; they answer different questions from the same brief. But if you are feeding this into a video model, the control arm's output leaves the consistency problem entirely unaddressed.
Two smaller things the skill arm did that the control arm did not: it diagnosed the source material in task 1, noting the excerpt does not carry enough dialogue to fill 60 seconds before proposing two invented beats to cover the gap; and it landed task 3's beat table on exactly 100 seconds and added 05_采访提纲.md (1,821 B), an interview outline — correct for a documentary-style piece where the subject has to actually say the lines.
Cost was close and not consistently in one direction: $0.39 / $0.29 / $0.51 / $0.44 with the skill against $0.49 / $0.23 / $0.50 / $0.30 without.
What we did not test: nothing was shot and nothing was generated. Every artifact in both arms is a document. The asset-ID scheme is the skill's central claim and we have no evidence it actually holds a character consistent, because no frames were ever produced from these manifests. All four tasks are Chinese-language vertical short form between 8 and 100 seconds.
13 天前 · v1.0.0
1 个月前 · v1.0.0
1 个月前 · v1.0.0