Seven runs, one song, one bassist. Six runs attached the recording as an audio reference and came back with a new melody and new timing. The seventh carried the same recording inside a black-screen video. It synced.
Tested September 13, 2026, on Seedance 2.5 through Tonpit. Every number here is from that project's generation log.
The audio reference is a voice, not a track
Seedance puts only two kinds of input on its timeline: a clip you edit and a clip you extend. Everything else, audio included, is a hint the model composes from. Its own guide calls an audio reference "the timbre in Audio 1". The voice, not the performance.
| You ask for | What the model takes | Result |
|---|---|---|
| A character with this voice | A timbre to imitate | Works |
| A character performing this recording | A timbre to imitate | New melody, new timing |
We tried the obvious fixes: writing out the lyrics, demanding the melody and tempo stay exact, trimming the audio to ten seconds and setting the generation to ten. Nothing moved. The audio was never on the timeline.
Seven runs, one variable
Same character sheet, same location, same song, same shot. Only the way the recording was attached changed.
| Run | Recording attached as | Length | Outcome |
|---|---|---|---|
| 1 | Audio reference | 12s | Re-composed |
| 2 | Audio reference | 12s | Re-composed |
| 3 | Audio reference, fixed close-up | 12s | Re-composed |
| 4 | Audio reference, lyrics in the prompt | 12s | Re-composed |
| 5 | Audio reference, longer generation | 14s | Re-composed |
| 6 | Audio reference, trimmed to exactly 10s | 10s | Re-composed |
| 7 | Black-screen video with the same 10s audio | 10s | Synced |
Run six is the control. Same audio, same length, still through the audio slot, still re-composed. The slot was the problem.
Bind a video reference for its sound only, and the recording lands on the clock.
Six steps
1. Carry the recording inside a video reference. A black frame with the audio on it is enough. Bind it for the sound only:
@Video1 is the song, a black-screen video that carries the audio. Use it for the music and the vocal only, never for its picture, and keep the original audio from @Video1 unchanged. The Bassist sings along to the song in @Video1 as if lip-syncing.
2. Write the recording's timeline, not only its words. The model does not hear what is said or when. Write it, inside the shot, in order: what happens before the first word, each line, each pause and its length, what happens after the last word.
he looks down at the bass for a second and a half, lifts his eyes to the lens
and says {first line}, a pause of a second with his mouth closed and a small
nod, {second line}, then his mouth stays closed and he plays until the endA silence direction binds to the sentence it sits in. "Mouth closed" in a sentence that goes on to a nod keeps the mouth closed through the nod. Keep the lines and pauses together as one passage, and give gestures their own sentences.
3. Match the length in whole seconds. Trim the audio to an integer length, then set the generation to that number. Any gap the model has to fill, it fills with its own music.
4. Never extend a performance. An extension invents a new song from the boundary. A longer song is several clips of ten to twelve seconds, cut together afterwards.
5. Keep the mouth readable from the first frame. Full or medium shot, frontal or three-quarter. In the clip above the camera starts wide, but the face is clear and the first line waits a second and a half for the push-in.
6. Lay the original recording under the picture in the edit. The output is still generated audio. This fixes the mouth and the timing, not the mix.
Dialogue and sound in Seedance 2.5
- Four bracket types.Music in (), sound effects in <>, dialogue in {}, subtitles in 【】. Angle brackets are for sound effects only.
- Five to ten words per line. Longer speech splits across shots.
- One spoken language per piece. Mixing degrades pronunciation.
- Sound goes inside the shot it belongs to, never in a list at the end.
- Say no to the score, or the model adds one:
No BGM, no voiceover. Generate only environmental sounds, action sounds, and the dialogue written above.
The clip above, in full
A spoken line instead of a sung one, the carrier road from the start, one run. Note where the timeline is written: before the first word, the pause between the lines, and what the hands do after the last one.
A cinematic live-action opening shot: Kai, a young man in a cream cardigan, sits in a leather lounge chair in a sunlit minimalist room with a bass guitar on his lap, speaks calmly to camera, then plays. [Asset bindings] @Image1 is Kai: a man in his late twenties with short dark curly hair, a fade, and a trimmed beard. Refer to it for his face and hair only; his clothes come from @Image2 and his pose from this text. @Image2 is Kai's outfit: a cream ribbed knit cardigan buttoned over a white t-shirt, brown wool trousers, white socks, dark brown leather loafers. He wears exactly these. @Image3 is the location: a minimalist room with a brown leather lounge chair, a stone side table with a lamp and a small pink ball, a framed print leaning on the wall, an olive tree, hard sunlight through a window casting leaf shadows on the plaster wall. The entire shot takes place here; refer to it for the environment, the light and the color grade only. The camera chooses its own framing and starts much wider than this image. @Video1 is the soundtrack, a black-screen video that carries the audio. Use it for the sound only, never for its picture, and keep the original audio from @Video1 unchanged. Kai speaks the line and plays the bass in exact sync with the audio in @Video1, as if lip-syncing to it. [Scene setting] Kai sits deep in the lounge chair from @Image3, relaxed, his right leg crossed over his left knee, ankle resting on the knee, and the sunburst electric bass guitar resting on top of the crossed leg, his left hand on the neck. The room is exactly as in @Image3, the chair at its center, the side table with the pink ball to his left. Light: the hard afternoon sun from the window falls across him and the wall, the same light as @Image3, warm and natural, with the leaf shadows moving faintly. Live-action realism, real scale, shallow depth of field, fine film grain. [Shot list] Shot 1: One continuous shot, the camera does not cut on its own. It opens on a wide shot from across the room, three-quarter angle at chest height: Kai is small in the frame, sitting with his right leg crossed over his left, the bass resting on it, the whole room around him, the full height of the plaster wall, the olive tree, the leaning print, the side table and lamp, and the concrete floor in front of him. The one camera move: a slow, steady push in across the room, arriving at a medium shot on him only at the very end. He is alive from the first frame: in the first second and a half he is looking down at the bass, settling it on his crossed leg with a small adjustment of his left hand on the neck, then he takes an easy breath, lifts his eyes to the lens and speaks without hurry, small easy smile, and from then on he never looks away from the lens. Dialogue (Kai, warm, low, unhurried): "If you need a motion design video for your product", then a pause of a second and a half in which his mouth is closed but he stays alive, a small nod, holding the look to camera, then: "here's a short and simple way to make one yourself". After the last word his mouth stays closed, a half-second beat, then his right hand drops to the strings and he plays a slow, deliberate bass line until the end, eyes still on the lens, head nodding slightly with the groove. Sound: the audio of @Video1 exactly as it is, with quiet room tone underneath. [Overall requirements] Kai is calm and confident, at home in the chair, always in gentle motion, never frozen; no comedy, no mugging. His right leg stays crossed over his left for the entire shot. English dialogue only. His face stays stable without deformation throughout, and his mouth reads clearly. The cardigan, trousers and loafers from @Image2 stay exactly as they are in every frame. The bass stays a sunburst electric bass and is never redesigned. No BGM, no voiceover, no added music; only the audio of @Video1 and room tone. Do not add subtitles; avoid generating any text or subtitles. Do not generate a logo or watermark.