Seedance 2.5 · Blog
ByteDance Seed Team Visual 2026-07-31

One-take Creation, Flexible Referencing:
Introducing Seedance 2.5

Today, we are officially launching Seedance 2.5, the new-generation video creation model. Since the release of Seedance 2.0, we have noticed a shift in what users expect from video creation models: from merely generating a clip to completing a creative work. Building on the unified multimodal audio-video joint-generation architecture of Seedance 2.0, Seedance 2.5 centers on foundational generation and reference-based generation, delivering major breakthroughs in long-form storytelling, multimodal reference, and editing.

Key highlights

  • Up to 30 seconds per generation, with multi-round extensions: High-quality, 30-second audio-video clips in a single pass. Improved shot transitions and scene changes for stronger continuity. Users can produce high-quality multi-minute content with a consistent audiovisual language.
  • Fully upgraded multimodal referencing: Up to 30 images, 10 video clips, and 10 audio clips as reference materials in a single pass. Strengthened clay render, motion, and creative references.
  • More precise and stable editing capabilities: Timestamp-level control for targeted editing of audio and video content. Enhanced green screen, camera perspective, and reference-based editing.

Access Seedance 2.5

Rolling out on Jimeng AI, Doubao Pro, and other platforms. API access via BytePlus ModelArk.

30-second long-form storytelling with multi-round extensions

Seedance 2.5 extends single-pass video generation from 15 to 30 seconds and further strengthens its storytelling in longer videos. Within 30 seconds, the model can organize multiple logically connected shots so that a story unfolds through setup, development, turning points, and resolution, rather than simply extending a single moment.

For example, in a one-take clip of a singer's stage performance, the model portrays the full story of the singer interacting with staff in the dressing room, then walking through the backstage corridor, meeting the dancers, and stepping onto the stage with them for the performance, instead of only the moment of walking on stage.

T2V Prompt — Singer performance (30s)
One-take handheld gimbal tracking shot. The camera slowly pushes in through a gap in a heavy red curtain and enters a warm-toned backstage dressing room. A young female singer, with her back to the camera, is adjusting her earpiece as a staff member reminds her it's time to go on. She turns toward the camera and starts singing citypop. The camera pulls back and tracks her as she passes through the curtain into a dim backstage corridor, interacting naturally with her dancers along the way; one staff member hands her a microphone. She and the dancers then step onto the stage, and the camera arcs around to the back, gradually revealing the red-and-black stage design, LED screens, spotlights, haze, and reflective floor. The camera finally pulls out to a wide shot of the arena, showing the packed audience, light boards, glow sticks, and cheering crowd, capturing the youthful, free-spirited climax of the concert.
Singer performance — 30s one-take (output)
Multi-round extension example

Fully upgraded multimodal referencing

Users can now input up to 30 images, 10 video clips, and 10 audio clips as reference materials in a single pass. The model strengthens a range of reference capabilities, including clay render, motion, and creative references, enabling it to better grasp the creator's intent and realize complex ideas that span multiple subjects, scenes, and shot changes.

Clay render reference

The clay render capability allows creators to use 3D blockout renders as visual references for video generation. The model preserves composition, camera, and scene structure from the clay render while generating photorealistic or stylized output.

Clay render reference — input
Clay render — output

Motion reference

The model extracts motion patterns from reference videos and applies them to new subjects or scenes, preserving timing, rhythm, and spatial logic.

Motion reference — input
Motion reference — output

More precise and stable editing capabilities

Seedance 2.5 offers timestamp-level control for targeted editing of audio and video content, notably improving efficiency and controllability. The model also enhances advanced editing features, such as green screen, camera perspective, and reference-based editing, to meet the rigorous demands of professional, complex fields like film and advertising.

Green screen / compositing

Green screen edit — example
Camera perspective edit
Reference-based editing
Audio editing (timestamp control)

With advancements in long-form storytelling, multimodal reference, and editing, Seedance 2.5 goes beyond longer single-pass video generation. The model better understands creative intent and delivers the journey from idea to finished video with greater control.