AI Video Prompts for Music Video

Seedance 2.0 prompts for music video — K-pop spectacle, hip-hop urban performance, indie golden-hour, EDM abstract VFX, and the structural techniques that give AI-generated music video its visual logic.

Music video is the genre where visual language and audio structure are most explicitly linked. Every shot decision — the width of the frame, the pace of the edit, the camera's distance from the performer — carries sonic meaning. A wide-angle slow tracking shot says "this is a slow build"; a rapid snap-cut close-up says "this is the beat drop." Most music video prompts fail because they describe appearance without rhythm: they name the artist and the song style but don't encode the visual grammar that makes the footage feel like it belongs to music. The fix is to think in structural terms first — what is the audio architecture of this section, and what camera and editing language matches it? The most fundamental structural decision in music video is the performance-to-atmosphere ratio: how much of the frame is occupied by the performer's body, and how much is given to the environment. This ratio changes by section. In verses, a narrower ratio (performer in a large environment, wide or medium framing) gives the visual weight to setting and mood, matching the lower audio energy of verse melody. In choruses, the ratio flips: tight framing on the performer's face, upper body, or hands concentrates visual energy on the human, matching the chorus's harmonic and lyrical payoff. Encoding this ratio shift explicitly — "verse: wide establishing pull-back showing performer alone in vast urban nightscape; chorus: cut to tight medium-close-up on performer, city lights bokeh behind" — gives Seedance the structural brief that makes the edit feel like a real music video rather than a series of unrelated performance shots. Genre-specific visual grammar is the single highest-leverage specification in music video prompting. Each genre carries a complete visual vocabulary: K-pop shoots are characterized by synchronized group choreography in high-contrast artificial light (LED wall back-projection, colored fluorescent banks, mirror floor reflections), fast cut rhythm matching the drop, and close-up insert shots on hands, faces, and footwork. Encoding K-pop grammar means naming the stage infrastructure (LED wall, mirror floor, color-temperature breakdown per fixture), not just the mood. Hip-hop music video grammar is different: handheld camera that places the camera operator in the performance space rather than observing from a fixed angle, urban practical light (streetlights, car headlights, neon storefronts), and a relationship between camera movement and the rapper's physical movement that suggests the operator is following rather than staging. Indie-folk grammar is naturalistic light (golden-hour outdoor, window light indoors, firelight at night), static or very slow dolly camera that separates indie from the constant motion of pop video, and a performer relationship to camera that suggests observation rather than performance — looking off-frame, unaware of the camera, intimate-documentary aesthetic. Naming the genre grammar rather than just the genre produces footage that reads as that genre. Beat synchronization is a core music video technique that requires explicit temporal instruction when prompting AI video. "Camera crash-zoom on the kick drum hit" is synchronization language — it names the audio event (kick drum hit) and the visual event (crash zoom) and their simultaneity. Without this synchronization specification, AI video defaults to smooth continuous camera movement uncoupled from audio structure, which produces footage that looks like a fashion commercial rather than a music video. The most effective beat-sync prompts name: (1) the audio event (kick, snare, hi-hat, bass hit, lyric line end, verse-to-chorus transition); (2) the visual event (cut to close-up, flash frame, camera punch-in, subject's physical movement peak); (3) their relationship (simultaneous, visual leads audio by half a beat, visual trails audio by one beat). A prompt that says "cut to ECU of performer's eyes simultaneous with the final word of each lyric line" gives Seedance a precise audio-visual synchronization model that produces footage with the timing feel of an edited music video. Lip sync technique in performance-focused music video requires understanding that the performer's relationship to the lyric is not about accuracy — it's about physicality and commitment. The best lip sync in music video looks like the performer is feeling the lyric rather than reciting it, which means the physical performance leads the vowels: the jaw opens before the vowel peaks, the expression changes a fraction before the lyric lands. Prompting lip sync means encoding the physical commitment: "performer sings directly to camera with high physical commitment — jaw movement emphatic, head position shifts with lyric emphasis, eyes holding camera contact through the line." This is different from "performer lip-syncs to music," which produces a flat neutral performance. The camera position for lip sync also carries meaning: front-facing tight close-up is confession or intimacy; three-quarter profile is performance-outward; looking-off-frame is contemplative or narrative. Specifying the camera-to-performer angle as part of the performance direction produces footage where the performance style and camera language are coherent. Narrative arc structure for a complete music video requires thinking in sections: intro, verse, pre-chorus, chorus, verse, chorus, bridge, final chorus, outro. Each section has a visual assignment — a set, a lighting scheme, a camera grammar, a performer relationship to frame. The most efficient way to encode this in a prompt is to use a three-column table (section, visual assignment, camera grammar) or a numbered shot list with section labels. A prompt that says "Intro: exterior establishing shot, performer walking toward camera in rain, natural streetlight, handheld follow. Verse 1: interior bedroom, single window light, static camera, performer seated, contemplative. Chorus: rooftop exterior, golden hour backlight, dolly circular track around performer, upward energy" gives Seedance a complete structural brief. The bridge is the section with the most compositional latitude — the visual equivalent of a harmonic pivot — and should be given the most distinct visual grammar from the verse and chorus.

No prompts here yet — browse the full gallery.

More use cases

Frequently asked questions

What are the best AI video prompts for music video?

The best music video prompts specify the structural grammar of the section — not just the look. For a chorus, name the tight framing on the performer (performer-to-environment ratio shifts to performer-dominant in choruses), the camera's physical energy (push-in, crash-zoom, circular track), and the beat-sync instruction ('cut simultaneous with the kick drum hit'). For a verse, name the wide framing that gives environmental weight, the slower camera movement that matches verse energy, and the performer's relationship to camera (observational, not performing-to-lens). Genre grammar is the second layer: K-pop needs LED wall back-projection and synchronized group choreography named explicitly; hip-hop needs handheld follow and urban practical light; indie-folk needs static or slow dolly with golden-hour or window light. Every strong music video prompt in this gallery combines section structure + genre grammar + beat-sync instruction.

How do I write a Seedance 2.0 prompt for a music video performance shoot?

A performance shoot prompt has four layers. First, name the performer's physical commitment: 'high physical commitment, emphatic jaw movement on vowels, head position shifts with lyric emphasis, eyes holding direct camera contact through each line' — this is different from 'performer lip-syncs to music,' which produces flat neutral performance. Second, specify the camera-to-performer relationship: front-facing tight close-up is intimacy/confession; three-quarter profile is outward performance; looking-off-frame is contemplative. Third, name the light source and its position relative to the performer (key light direction, fill ratio, practical lights in background). Fourth, add a beat-sync instruction for at least one camera event: 'camera punches in to ECU simultaneous with the word [lyric].' Those four layers together produce a performance shoot that reads as directed music video rather than candid footage.

Can Seedance 2.0 generate different styles of music video — K-pop, hip-hop, indie?

Yes, but each genre requires its specific visual grammar named explicitly. For K-pop: name the stage infrastructure (LED wall with dynamic color animation behind the group, mirror floor reflecting colored light, synchronized choreography with formation changes, close-up insert shots on hands and footwork). For hip-hop: name the camera operator's relationship to the scene (handheld operator inside the performance space following the artist rather than observing from fixed position, urban practical light from streetlights and car headlights, low-angle looking up at artist). For indie-folk: name the naturalistic light (golden-hour window, single practical lamp, overcast exterior) and the static or barely-moving camera that observes rather than stages. The genre name alone is not enough — the visual infrastructure that defines each genre is what tells Seedance which visual language to apply.