Close-up shots in AI video fail in a specific way: the model receives "close-up" as a vague size instruction, not as a shot-design decision. The result is a frame that is technically tight but structurally wrong — no depth of field motivated by a lens spec, no transition logic between the tight frame and the wider world it belongs to, no distinction between a probe-lens push-in through steam and a 100mm macro arc pulling back to reveal a transformed subject. A close-up is not a framing choice; it is a complete camera-optics-and-motion decision that encodes why the lens is this close, what it expects to find there, and how it plans to leave.
The five prompts below cover five entirely different close-up disciplines: a 100mm macro prime that starts on a pencil drawing and arcs back to reveal a fully realized 3D Formula 1 car; a live sports broadcast structure that uses ECU on the athlete's face to carry emotional stakes before the wide shot shows physics; a probe lens that enters the cooking environment through backlit steam at 120fps; a fashion morning routine that encodes skin pores, water beads, and fabric texture as a three-layer material grammar; and a sci-fi visor-slam where a single ECU moment at the climax carries the weight of an entire scene. Together they map from object-reveal to material-texture to dramatic-punctuation — the full range of what a close-up can structurally accomplish in a short video.
1. The 100mm macro arc reveal — object transform via ECU pull-back
See the full prompt on scenic.sh →
"A 100mm Macro Prime, aperture f/2.8 for shallow depth of field. Starts as an extreme close-up on the pencil drawing, then slowly pulls back and arcs around the car as it transitions to a 3D object. The transformation is complete. The car is now fully 3D, sitting on the paper as if it's a real, miniature vehicle."
Why this works: At 323 likes, this prompt's defining technique is the "reveal via pull-back" — the camera begins at ECU of a pencil sketch, so close that only graphite lines and paper texture are visible, and the pull-back reveals context rather than entering it. The choice of 100mm macro is structural, not decorative: a 100mm prime at f/2.8 produces a specific background compression ratio. When the camera is at ECU on the sketch, the desk lamp and scattered tools are invisible in the blur; as the camera pulls back and arcs, they resolve gradually into the frame, giving the transformation a spatial anchor and credibility. The desk objects — pencil, compass, eraser — function as scale witnesses. The viewer never sees the full desk until the car has already become three-dimensional, which means the impossible event is framed by mundane physical context rather than preceded by it.
The transformation technique is equally specific: the wave of change "begins at the front wing and sweeps backward," a directional animation that the camera arc reinforces by orbiting in the same axis as the wave's travel. The camera movement and the transformation movement are synchronized; they resolve in the same spatial direction. The audio architecture mirrors this: graphite scratching on paper transitions to "low magical humming" that grows, then full "swelling orchestral score" as the transformation completes. The prompt provides three temporal audio layers mapped to three visual phases (sketch, transformation, complete 3D car), giving the model a sound design structure that prevents audio from being generic ambient music.
The takeaway: use a macro lens spec (100mm, f/2.8) to control depth-of-field behavior across a pull-back reveal. At ECU, the tight DOF isolates the subject and hides context. As the camera pulls back, resolving DOF re-introduces the environment in stages — this is the reveal happening through optics, not through a cut. Synchronize camera arc direction with the transformation's directional logic (front-to-back wave → camera arcs around the same axis). Provide audio phases mapped to visual phases so the model can build the soundscape in temporal layers rather than defaulting to background ambience.
2. The sports ECU structure — concentration face → wide action → ECU impact
See the full prompt on scenic.sh →
"0-3s: [Close-up shot] addressing the ball, focused expression. He swings with perfect physics. 8-12s: [Extreme close-up cut] The golf ball lands softly on the green, rolls with realistic friction, and drops perfectly into the hole. [Sound: The 'plop' of the ball hitting the bottom of the cup followed by a sudden, deafening crowd explosion]."
Why this works: At 164 likes, this prompt demonstrates the standard grammar of live sports television close-up coverage — a grammar that works in AI video for the same reason it works in broadcast: it sequences close-up cuts to carry distinct information at each stage. The opening close-up is on the athlete's face because the face carries the psychological stake (concentration, focus, the weight of the shot). The wide shot (3–6 seconds) tracks the ball because the ball's physics is the event — trajectory, distance, angle of descent. The ECU at 8–12 seconds is on the result — the ball dropping into the cup — because the result is a tiny physical event (a ball dropping into a four-inch hole) that is only legible at macro scale.
Each close-up in this structure encodes different information: psychological state (face), event physics (wide), and physical resolution (ECU impact). This sequencing is what makes the structure feel like real sports coverage rather than a random close-up. The audio architecture reinforces the same sequence: "hushed commentator voice-over" during the face close-up (which cues the audience to hold attention), "wind whistle" during the wide ball-tracking (which gives the trajectory acoustic motion), and the "deafening crowd explosion" at the cup drop ECU (which provides the emotional payoff that the face close-up set up). The crowd explosion is withheld until the result ECU; this is the structural payoff of delaying the wide crowd audio.
The takeaway: structure close-up coverage as a three-phase information sequence: face (psychological state), wide (event physics), ECU (physical result). Each phase carries distinct information that the other two cannot. Map audio to the same phases: ambient tension on face, motion/wind on wide, full crowd release on the result ECU. This gives the model a temporal audio-visual contract rather than a single mood instruction. The ECU on impact is only powerful because the face close-up preceded it — close-up technique is a sequencing discipline, not a framing choice.
3. The probe-lens food macro — physical camera entry through steam and oil
See the full prompt on scenic.sh →
"Extreme macro, garlic and chili meeting hot oil, lively golden bubbles, steam rising in a soft backlit plume, 120fps slow motion with every bubble crisp, probe lens push-in through the steam. Orbiting macro around the pour. Close-up tracking the arc."
Why this works: At 114 likes, this prompt's defining technique is the "probe lens push-in" — a camera move that implies the lens is physically entering the cooking environment, not framing it from outside. A probe lens (also called a periproduct lens) is an extreme close-up tool with a very long barrel that can get within centimeters of hot, wet, or chaotic subjects. Specifying it in an AI video prompt encodes both the optic behavior (nearly zero perspective distortion at micro distances, extremely shallow DOF with circular bokeh) and the camera's spatial relationship to the subject (inside the steam, not above it). The prompt describes the probe lens push-in through steam, not toward steam — a preposition that changes the model's understanding of the camera's position relative to the subject.
The prompt sequences five distinct close-up setups in 15 seconds by assigning each a specific camera move rather than holding a static tight frame: the probe lens push-in through steam (Scene 5, garlic in oil), the macro around the pour (Scene 7, sauce ribbon), the close-up tracking the arc (Scene 3, butter toss), the top-down locked shot (Scene 2, herb prep), and the slow push-in finale (Scene 10, hero dish). The speed-ramping technique — "120fps slow motion with every bubble crisp" for the probe lens shot, real-time for the chef's hand movements — differentiates the close-up shots from the wide shots acoustically and temporally. Every scene transitions via "hard whip-pan," which means the close-up cuts feel percussive and kinetic rather than contemplative.
The takeaway: specify "probe lens push-in" rather than "close-up camera move" to encode that the camera is entering the subject's micro-environment. The preposition ("through steam" not "toward steam") is a spatial directive. Assign a distinct camera move to each close-up setup so they read as five different shots, not one tight frame held across multiple events. Use speed-ramping at macro scale (120fps slow motion for the garlic-in-oil ECU, real-time for the hand movement) to create temporal contrast within the sequence — the slow-motion close-up feels like a revelation, not just a different angle.
4. The texture-first fashion detail — three-layer skin, fabric, and liquid grammar
See the full prompt on scenic.sh →
"Shot on 35mm lens, f/1.8 aperture for soft background bokeh. Visible skin pores, light dew, subtle texture on silk and leather fabrics. High-definition capture of water beads breaking against her face and neck, slow cascade effect, glistening hydrated skin. Macro shot of hands: fastening a heavy, chunky gold chain necklace, clicking a leather handbag strap."
Why this works: At 93 likes, this prompt builds close-up technique on a material grammar of three distinct texture registers: skin (pores, dew, hydration), liquid (water beads breaking against the face, cascading, glistening), and fabric (silk sleepwear, leather robe belt, chunky chain texture). Each register requires different lighting and focal distance to resolve correctly, and the prompt implies a sequence through them: the bed wake-up captures skin texture in warm indirect light, the shower captures liquid physics against skin in high-humidity conditions, the accessorizing scene captures material texture (chain, leather strap) in the sharp side-light of a walk-in closet. The three registers are spread across three environments, which gives the model a reason to change the lighting between close-up shots.
The lens spec (35mm at f/1.8) is a softer DOF choice than a 100mm macro — it produces character-level close-ups with background separation rather than object-level macro isolation. At f/1.8 on a 35mm focal length, the background of a face close-up is still recognizable as a bathroom or bedroom; at 100mm f/2.8, it would disappear entirely. This is the difference between "texture in context" (fashion) and "texture in isolation" (product). The phrase "visible skin pores" is a precision directive, not a style descriptor — it tells the model to render skin at enough resolution that epidermal structure is legible, which is what separates a cosmetics-grade close-up from a standard portrait.
The takeaway: build a three-layer texture grammar (skin, liquid, fabric) and distribute the layers across distinct environments so that the model has a lighting-change motivation for each close-up shot. "Visible skin pores" is a resolution instruction, not a style choice — it tells the model to render at epidermis scale rather than at face scale. Choose your focal length by the texture-in-context vs texture-in-isolation decision: 35mm f/1.8 retains environmental context; 100mm f/2.8 eliminates it. Fashion close-ups almost always want context; product close-ups almost always want isolation.
5. The ECU as emotional punctuation — one tight moment that carries a scene
See the full prompt on scenic.sh →
"Shot 2 (The Impact & Visor): Extreme Close-up. The ship takes a massive kinetic hit to the hull; a violent jolt shakes the camera frame. Victra pulls back, her face illuminated by flickering red strobes, and whispers: 'Good luck.' She reaches up and slams Darrow's heavy metallic helmet visor down over his face with a loud mechanical clack."
Why this works: At 73 likes, this 10-second prompt demonstrates the most economical use of ECU in a short video: one extreme close-up shot at the structural climax of a two-shot scene. The first shot is a medium shot of the kiss and the ship's corkscrew roll — a wide-enough frame to establish the chaos of the space battle and the spatial relationship between the two characters. The second shot, labeled "Extreme Close-up," is where the scene's meaning is concentrated: the visor slam, the "loud mechanical clack," the flickering red strobes on her face at extreme proximity. The helmet closing is a mechanical event, but it stands in for departure, danger, and finality. At medium distance, this event would read as an action beat; at ECU, it reads as farewell.
The structural decision here is that the ECU contains both a facial reaction (her face illuminated by red strobes) and a mechanical action (slamming the visor). The two elements at extreme close range create a compound image: human emotion and machine response in the same frame. The sound design carries the ECU's full weight: "loud mechanical clack" followed immediately by "another massive impact jolt" — the mechanical closure and the external chaos happen within the same ECU. The camera-frame rotation (the entire 360-degree corkscrew) is the macro-context that makes the ECU legible; without the physical chaos of the rotating ship, the visor slam would be a small action. The ECU is powerful because everything around it is large-scale and chaotic.
The takeaway: use ECU as punctuation, not as mood. A single ECU shot at the structural turning point of a scene (a departure, a decision, a physical threshold being crossed) carries more weight than a scene written entirely in tight frames. The ECU works because the medium shot preceded it — contrast is the mechanism. Include a compound image in your ECU (face + physical object) when you need to encode both emotional state and physical event in one frame. Pair the ECU with a distinctive sound event ("loud mechanical clack") so the tight frame has acoustic grounding — without it, a close-up can feel like silence.
Seedance close-up prompt cheat sheet
Across all five, the structural principles that make Seedance close-up prompts work:
- Macro lens spec as DOF contract (100mm f/2.8) — specify the lens and aperture rather than "close-up." The lens encodes the depth-of-field behavior across the shot. ECU at 100mm f/2.8 isolates the subject completely; pulling back resolves the environment gradually — the reveal happens through optics.
- Three-phase sports ECU sequence (face → wide → ECU impact) — close-up is a sequencing discipline. The face close-up carries psychological state; the wide shot carries event physics; the ECU on impact carries the result. Each phase encodes different information. Map audio phases to match: ambient tension on face, physical motion sound on wide, full payoff on result ECU.
- Probe lens push-in ("through steam" not "toward steam") — the preposition encodes the camera's spatial relationship to the subject. A probe lens entering the cooking environment reads differently than a macro lens framing it. Use "push-in through" to tell the model the lens is inside the subject's space.
- Three-layer texture grammar (skin / fabric / liquid) — assign each texture register its own environment and lighting condition so close-up shots read as distinct shots, not the same tight frame with different subjects. "Visible skin pores" is a resolution directive; f/1.8 at 35mm retains environmental context while isolating texture.
- ECU as single punctuation moment at the structural climax — one ECU shot after a medium-shot setup carries more weight than a scene in continuous tight frames. Include a compound image (face + mechanical object) so the close-up encodes both emotional state and physical event. Pair it with a distinctive sound event for acoustic grounding.
Browse the Scenic cinematic gallery for more camera technique examples, or see Seedance tracking shot prompts for follow-shot camera movement. Read how to write Seedance 2 prompts for the complete cinematic prompting guide.