What a Voice Actor Actually Does in an AI Dubbing Studio
The assumption is that AI dubbing replaced voice actors. In our studio the opposite happened: every language version of every campaign begins with a native voice actor's performance, and the AI's job is to carry that performance into another person's voice. The actor's timing, dialect and emotional choices survive all the way to air. What changed is what their performance carries, and the craft, if anything, got more exacting.
The session looks familiar, the brief does not
The recording session looks like any dubbing session: a performer at a microphone, a director, the film on a screen, take after take against the picture. The brief is where it differs. A traditional dub actor performs a character for the audience to hear as them. In a speech-to-speech session, the actor performs knowing their voice will be re-rendered in the on-screen star's voice. They are lending everything except their timbre: the Tamil rhythm, the comic beat, the softness on the closing line. We describe it as driving, the actor drives the performance, and the cloned voice is the body it arrives in.
That is precisely how the NutriChoice campaign worked: regional artists performed Bengali, Kannada, Tamil and Punjabi scripts, and Aamir Khan's consented voice clone delivered their performances. Four artists' work, one star's voice, every version alive because a human performed it.
Why the craft got harder, not easier
- The picture is the metronome. The actor performs to the original edit's rhythm, hitting emotional beats where the on-screen face places them, because the lip re-animation will follow their audio exactly. A beat performed late is a face animated late.
- Dialect is the deliverable. Casting is for the market, a Chennai register, a Punjabi warmth, because dialect authenticity is exactly what the clone cannot invent and the audience cannot be fooled on.
- Restraint reads louder. With the star's voice carrying the output, overacting doubles: the performance and the timbre both amplify. The best speech-to-speech actors perform with a precision closer to film acting than classic dub work.
The consent line
Two consents hold this system up, and we treat both as preconditions. The star consents to their voice being cloned and approves its use per campaign, with contract coverage for languages and territories. The voice actor works with full knowledge of how their performance is used, credited and paid as the performer they are. An industry that gets these two consents right gets to keep the trust that makes the technology usable; the corners cut elsewhere are how the whole category earns suspicion. Our position is simple: the technology borrows voices, and borrowing requires asking.
The wider pipeline around the session, adaptation, re-animation, native review, broadcast QC, is described in What is Lipsmatch, and the four failure modes the human performance protects against are in Why dubbed ads feel off. If you are planning a campaign and wondering what the session for your film would look like, come watch one.