Dance video prompts in Seedance require three simultaneous specifications that most prompt writers treat as one: what the body is doing step by step, what the camera is doing in relation to that body, and how the prompt handles the parts of choreography that are hardest for video models to sustain — prop continuity, cloth physics, body connection in couple dances, and environmental coherence across multiple seconds of motion. Getting any one of these wrong produces a video where the dancer appears technically capable but the performance feels ungrounded: the fabric clips through itself, the selendang disappears mid-sequence, the partner loses contact, or the background resets rather than transforms.
The five Seedance dance prompts below cover five structurally different problems in AI dance direction. The traditional selendang prompt demonstrates how to use a reference image as a choreography instruction sheet and how to specify four simultaneous tracking constraints (hands, feet, expression, prop fabric) without slowing the camera down. The street dance prompt shows how "one continuous shot, no cuts" plus cloth physics creates a fixed-camera stage that prioritizes footwork legibility over cinematic movement. The tango prompt shows how a couples dance prompt is more a social ritual script — invitation, frame, walk, ocho, dip — than a movement sequence. The hip-hop morphing prompt demonstrates how to lock a dancer in position while making the entire world transform around them using motion-driven transition triggers. And the decade-tracking prompt shows how to use camera direction as a time axis — spatial left-to-right movement becomes temporal progression through dance history.
1. The traditional selendang — reference sheet as instruction + four simultaneous tracking constraints
See the full prompt on scenic.sh →
"the selendang floating naturally as the dancer glides forward and backward. End on a still closing pose with hands at the heart, peaceful smile, and elegant silence."
Why this works: At 108 likes, this prompt's highest-leverage instruction is the reference image designation: "Use the first reference image as the exact choreography and motion-process guide." This converts the image from a style reference into an instruction sheet — the model is told to follow all 16 illustrated steps in order, not to capture the general aesthetic. The difference matters: a style-reference image produces approximate visual similarity, while an instruction-sheet image produces step-by-step sequencing of specific named moves (opening pose, cross steps, wrist waves, shoulder accents, hip sways, flowing turns, closing pose).
The prompt stacks four simultaneous visibility constraints in one sentence: "The dancer's hands, feet, facial expression, and selendang fabric must remain visible throughout." Each element has its own tracking logic — hands for wrist wave precision, feet for footwork clarity, facial expression for emotional registration, selendang for prop continuity. Rather than writing four separate instructions, the prompt collapses them into a single constraint sentence, which signals to the model that these four elements form a single non-negotiable unit.
The camera instruction inverts the usual relationship: "Keep the camera in a clean full-body cinematic frame, mostly front-facing, with slow controlled movement that supports the dance instead of distracting from it." The camera subordinates itself to the choreography rather than directing it. The only exception — "During the left and right turns, allow a subtle circular camera drift, then return to a centered frontal composition" — gives the camera one permitted move and explicitly instructs it to return, establishing front-center as the default state.
The closing instruction defines the sequence's terminus as a physical-emotional composite: "still closing pose with hands at the heart, peaceful smile, and elegant silence." All three elements (stillness, gesture, expression) are specified simultaneously at the end point, giving the model a concrete landing target for the final frame rather than an open-ended "end naturally" instruction.
The takeaway: use a reference image as a choreography instruction sheet by specifying it as "the exact choreography and motion-process guide" rather than as a style reference. Stack multiple tracking constraints in a single sentence ("hands, feet, expression, selendang fabric must remain visible") to signal they form one non-negotiable unit. Specify the camera's default state and its one permitted deviation: "mostly front-facing... allow a subtle circular drift during turns, then return to centered frontal." Define the closing pose as a physical-emotional composite (stillness + gesture + expression) rather than an ambient description.
2. The fixed-camera street dance — footwork legibility and cloth physics on a continuous stage
See the full prompt on scenic.sh →
"One continuous shot (no cuts) Full-body framing, camera fixed and centered Smooth, steady perspective (no zooms or cinematic transitions) Duration follows a 16-step choreography sequence"
Why this works: At 20 likes, this prompt establishes its structural logic in four consecutive constraints that together create a fixed-camera stage: one continuous shot, no cuts, full-body framing, camera fixed and centered. Each constraint does different work. "One continuous shot" prevents editorial shortcuts. "No cuts" eliminates the possibility of covering choreography gaps with a scene change. "Full-body framing" means the dancer's feet and head must remain in frame simultaneously — the camera cannot tighten on the face to hide footwork issues. "Camera fixed and centered" removes the possibility of repositioning to compensate for positioning drift.
The environment specification — "Ground texture (asphalt or concrete) visible for footwork clarity" — reveals the prompt's priority: footwork legibility over environmental atmosphere. The street setting is not there for mood; it is there because its surface texture makes foot placement readable. "No crowd, no vehicles, no distractions" removes visual noise that would compete with the dancer's feet for attention at ground level.
The cloth physics instruction — "Cloth physics for skirt movement (gentle sway and rotation response)" and "Natural body mechanics and timing / Lighting interacts realistically with the character and ground" — addresses the most common failure mode in AI dance video: clothing that clips, freezes, or moves at wrong scale relative to the body. Specifying cloth physics at the prompt level, not as a style tag, ensures the model allocates generation capacity to fabric dynamics throughout the full 16-step sequence rather than treating cloth as a secondary detail.
The motion detail instruction distinguishes choreography from cinematography: "Realistic weight shifts and foot placement / Smooth transitions between poses (no stiffness) / Optional very subtle motion trails to enhance clarity (kept minimal and natural)." The motion trail caveat — "kept minimal and natural" — prevents the visual effect from overwhelming the footwork it is meant to clarify.
The takeaway: use the fixed-camera + no-cuts + full-body combination to create a stage where choreography has nowhere to hide — the model must generate complete footwork because the camera cannot reframe or cut away. Make the environment serve legibility: "ground texture visible for footwork clarity" is a purpose statement, not an atmosphere note. Specify cloth physics explicitly ("gentle sway and rotation response") rather than relying on a style tag. Separate choreography detail from cinematography — motion trails are a cinematography tool that must be explicitly restrained when the goal is choreography legibility.
3. The couple tango — invitation-to-dip as social ritual arc
See the full prompt on scenic.sh →
"they form a classic tango frame, glide... and finish with a dramatic elegant dip pose, maintaining romantic tension and continuous body connection throughout."
Why this works: At 9 likes, this couples dance prompt is structured as a social ritual arc rather than a movement sequence: invitation (man extends hand), acceptance (she accepts and steps closer), union (form a classic tango frame), execution (forward walk, side step, pivot turn, ocho setup, ocho), and resolution (dramatic dip). The sequence moves from social gesture to technical execution to physical terminus — the same structure as a real tango with a stranger.
"Romantic tension and continuous body connection throughout" is a single compound constraint that governs the entire 15 seconds. "Romantic tension" is a psychological instruction that specifies the performers' relationship, not just their positions. "Continuous body connection" is a physical instruction that prevents the AI from separating the couple to show off individual technique. These two terms anchor every other movement in the sequence: the forward walk happens within the tango frame, the pivot happens with sustained connection, the dip maintains the embrace.
The reference image strategy here differs from the selendang prompt: "Use the tango choreography sheet only for movement order, dance structure, and body-action logic. Do not recreate the sheet, text, arrows, numbers, borders, or panel layout." The negative instruction — do not recreate the visual content of the sheet, only its structural logic — prevents the model from treating the reference image as a visual target and instead uses it purely as a sequencing guide.
Character specification at the couple level requires both visual consistency and relational consistency. The detailed visual descriptions (her red dress with thigh slit, his fitted black shirt) lock both performers' appearances for the full 15 seconds. The dip is specified as dramatic and elegant — both the energy level (dramatic) and the aesthetic register (elegant) are named, preventing an abrupt or clumsy terminus.
The takeaway: structure couples dance as a social ritual arc — invitation, acceptance, frame, execution, resolution — rather than as a movement list. Use a compound constraint ("romantic tension and continuous body connection") to govern the entire sequence's emotional and physical relationship simultaneously. Separate the reference image's visual content from its structural logic: specify "use for movement order and body-action logic only, do not recreate the visual content" to prevent the model from treating the sheet as the output target. Specify the terminus at two registers (dramatic + elegant) to anchor both energy and aesthetic for the final pose.
4. The morphing hip-hop — locked position with motion-driven world transformation
See the full prompt on scenic.sh →
"dancer locked in position (no teleporting) / environment changes around her only / motion must stay perfectly continuous"
Why this works: At 7 likes, this prompt's central technique is the split instruction: the dancer is locked in one place while the world transforms continuously around her. "CRITICAL: dancer locked in position (no teleporting) / no body or face distortion / environment changes around her only" — three CRITICAL constraints that together define the video's entire visual logic. The dancer is a fixed reference point; everything else is variable. This division prevents the common failure mode where morphing videos lose the subject's position or consistency as the world changes.
The motion-driven transition system is the prompt's most original structural feature: "camera whip → environment shifts / spin → 360° background morph / moonwalk → ground transforms under feet / crash zoom → new environment revealed / light flicker → environment rebuilds." Each camera or body action is explicitly mapped to a specific world-transformation trigger. This mapping removes ambiguity about when transitions occur — they are not random or time-based but triggered by specific movements. The moonwalk-to-ground transformation is the most precise example: when the dancer performs the moonwalk, the ground surface underneath her feet specifically transforms, not the entire background.
The three-layer transformation (clothing, style, environment) is specified separately: clothing changes "every frame (no repetition)" through streetwear to performance wear; style switches "between realistic, anime, line art, watercolor, comic, pastel, digital paint, sketch, collage, clay, low-poly, neon outline, glitch graphic, storyboard, chalk, airbrush, silhouette"; and environment "constantly transforms in real-time around her — NOT hard cuts" through twelve named locations. The "NOT hard cuts" instruction for the environment layer specifies a continuous transformation rather than a scene change, while the clothing and style layers imply frame-by-frame switching.
The depth and lighting specifications respond to the transformation's visual complexity: "Real parallax (foreground / midground / background)" prevents the transformed environments from reading as flat backdrops, while "Lighting adapts to each environment... and stays consistent on her body" ensures the dancer is consistently lit despite the background changing from neon to studio to digital grid. Consistent body lighting is the technical requirement that makes the locked position actually read as stable.
The takeaway: split the video's transformation logic explicitly — "dancer locked in position / environment changes around her only" — so the model has an unambiguous anchor. Map each motion to a specific transformation trigger (moonwalk → ground transforms, spin → 360° morph, camera whip → environment shifts) rather than using time-based or random transitions. Separate the transformation layers (clothing, style, environment) and specify each layer's change rate differently. Maintain consistent body lighting through the transformations with "lighting adapts to each environment and stays consistent on her body."
5. The decade-tracking shot — spatial movement as time axis through dance history
See the full prompt on scenic.sh →
"She spins one final time. Every era's dance floor appears as transparent layers behind her — speakeasy, ballroom, disco, street, studio — all dancing simultaneously. She stops. Bows."
Why this works: At 3 likes but with one of the most precise spatial-temporal structures in the gallery, this prompt establishes the camera's left-to-right tracking movement as the timeline axis. Moving right is moving forward in time. Each segment is named with year and location, not time code and shot description: "1920s speakeasy," "1940s ballroom," "1960s TV studio," "1977," "1984 street corner," "2000s music video set," "present-day bedroom." The progression is geographic movement that is simultaneously temporal — the camera does not cut to a new era, it tracks into it.
Each temporal segment is built from three elements that anchor the AI in a specific cultural period: costume ("fringed flapper dress," "flowing gown," "go-go boots," "sequins and platforms"), environment detail ("smoke and amber light," "Big band horns blare," "checkerboard floor," "Mirror ball"), and a specific named dance move associated with that era ("Charleston," "swing dancing... lands in a split," "doing the Twist," "pointing to the sky and striking the classic pose," "windmill, headspin, freeze"). All three elements resolve simultaneously — if only the costume changes but the dance move is wrong, the era is ambiguous.
The exit mechanism between eras uses a physical action that carries the camera forward: "She kicks right and the room stretches" (transition from 1920s to 1940s). "The camera pushes through a bead curtain" (transition from 1960s to 1977). Each era ends with a specific physical action that propels the camera into the next segment without a cut. These transition mechanisms are embedded in the physical space of each segment — the room stretches, the camera pushes through an object — rather than being camera-only instructions.
The closing overlay — "Every era's dance floor appears as transparent layers behind her — speakeasy, ballroom, disco, street, studio — all dancing simultaneously" — resolves the sequence by collapsing the time axis back onto a single point. The transparent-layer instruction converts the decade-progression from a sequence into a simultaneous superposition: all eras visible at once, all applauding. The sound design mirrors this: "all eras merging into one final beat → applause → silence" — the audio arc is a compression of the spatial arc.
The takeaway: use camera tracking direction as a time axis — left-to-right movement becomes temporal progression through history. Anchor each temporal segment with three cultural elements simultaneously: costume, environment detail, and era-specific named dance move. Build exit transitions into the segment's physical space (the room stretches, the camera pushes through an object) rather than using pure camera cuts. Close with a superposition image where all eras become simultaneously visible — this inverts the tracking logic (sequential → simultaneous) and provides a visual resolution that a single final frame cannot.
Seedance dance prompt cheat sheet
Across all five, the structural principles that make AI dance video Seedance prompts work:
- Reference sheet methodology — specify a choreography image as "the exact choreography and motion-process guide," not as a style reference. Stack four tracking constraints in one sentence (hands, feet, expression, prop fabric). Define the closing pose as a physical-emotional composite (stillness + gesture + expression) rather than an ambient ending.
- Fixed-camera stage — "one continuous shot, no cuts, full-body framing, camera fixed and centered" creates a stage where choreography has nowhere to hide. Make the environment serve legibility (ground texture for footwork clarity). Specify cloth physics explicitly ("gentle sway and rotation response") rather than relying on style tags.
- Social ritual arc for couples — structure partner dance as invitation-to-terminus arc (extend hand → accept → frame → execute → dip), not as a movement list. Use a compound constraint ("romantic tension and continuous body connection") to govern the entire sequence's emotional and physical relationship simultaneously.
- Locked position + motion-driven world transformation — "dancer locked in position / environment changes around her only" splits the transformation logic so the subject is always the visual anchor. Map each camera or body action to a specific transformation trigger (moonwalk → ground transforms, spin → 360° morph). Maintain consistent body lighting through world changes.
- Camera direction as time axis — left-to-right tracking becomes temporal progression through history. Anchor each temporal segment with costume + environment detail + era-specific dance move simultaneously. Build exit transitions into the segment's physical space, not the camera alone. Close with a superposition image where all eras are simultaneously visible.
Browse the Scenic dance gallery for more AI dance video examples, or see Seedance cinematic prompts and Seedance surreal prompts for related techniques. Read how to write Seedance 2 prompts for the complete prompting guide.