01
Quad-Modal Input: Text, Image, Video, and Audio in One Prompt
One of the most distinct features of Seedance 2.0 is its support for quad-modal input. Users can combine up to 12 different assets — including text, images, video clips, and audio files — into a single generation request. Upload up to 9 images, 3 videos (15s total), and 3 audio files. Combine them freely to express your creative vision with unprecedented flexibility. Reference motion, effects, camera movements, characters, scenes, and sounds from any uploaded content. This quad-modal approach means you no longer rely on lengthy text descriptions alone. You can show the model what you want — a dance style, a camera angle, a musical rhythm — and it follows.
One of the most distinct features of Seedance 2.0 is its support for quad-modal input. Users can combine up to 12 different assets — including text, images, video clips, and audio files — into a single generation request.Upload up to 9 images, 3 videos (15s total), and 3 audio files. Combine them freely to express your creative vision with unprecedented flexibility. Reference motion, effects, camera movements, characters, scenes, and sounds from any uploaded content.This quad-modal approach means you no longer rely on lengthy text descriptions alone. You can show the model what you want — a dance style, a camera angle, a musical rhythm — and it follows.