How to record a voice over: the method, and what a studio changes

Voice actor at the microphone, headphones on, during a take

The file plays back clean. No hiss, no pops, the level sits where it should, every word is clear. And something is still wrong. The read does not carry. The person it was made for will not be able to say why. They will simply stop listening.

Most pages that explain how to record a voice over tell you which button to click. They are not wrong. The technical side of a voice over recording is learnable, and faster than people expect. What they leave out is that a technically clean recording can miss its target completely. That fault rarely shows up in headphones. It shows up in distribution.

The technique is worth having. What decides the result sits either side of it: the preparation before, and the direction during.

Prepare the script for the ear

A text written to be read is not automatically a text that can be said. Long sentences that look fine on a page turn into tunnels with nowhere to breathe in front of a microphone. Read your script out loud, all of it, standing up. You will hear where it snags.

Cut anything that does not fit inside one breath. Replace the words your mouth trips on. Move the important information to the front.

Then mark up your copy: the breaths you want to hear, the words that carry the meaning, the pronunciation you expect for proper nouns, acronyms and figures. That detail wastes more session time than anything else.

Finally, decide the intention before you open the microphone. Three adjectives are enough to frame a read: warm, instructive, reassuring, for example. If you cannot pick them, four questions produce them.

  • Who are we talking to?
  • To do what?
  • Where will the message play?
  • And what do we absolutely not want to hear?

The last one is the most useful, and it is almost always the one that gets skipped.

Treat the room before you treat the sound

This is the advice that changes the most and gets written the least. A room holds two problems that get confused constantly: stopping noise from getting in (traffic, neighbors, ventilation, a refrigerator), and stopping your own voice from bouncing off the walls and coming back into the microphone slightly late, which produces that hollow conference-room sound. Acoustic foam only addresses the second. It absorbs reflections, it does not stop a scooter going past.

Against noise, the real decision is the room and the hour: an interior space, no window onto the street, doors sealed, noisy appliances off for the length of the take. Against reflections, anything soft works for you: hanging clothes, heavy curtains, thick rugs. A well-filled closet makes a better recording space than most offices.

One test settles it. Clap your hands in the middle of the room and listen to what follows. If the sound dies flat, you are fine. If you hear a tail, add absorption and try again.

And keep this paradox in mind: the more sensitive a microphone is, the more it exposes the faults of the room. A modest microphone in a dead space beats a high-end one in a bare room. It is the part of a voice recording setup that costs the least and returns the most.

The take: distance, angle, headroom

Distance first. The useful order of magnitude is around 15 to 20 centimeters between the mouth and the microphone, adjusted for the microphone, the timbre and the volume of the voice. Too close and the voice thickens, and every mouth noise comes with it. Too far and the room walks in with you.

Angle next. Speak slightly across the axis of the microphone rather than straight into it. The consonants that project air, P, B, T and D, otherwise send a gust that hits the diaphragm and produces the dull thud known as a plosive. The pop filter, the round screen between mouth and microphone, solves part of that. The angle solves the rest.

Headroom last. Set the input level so the voice is firm without ever approaching the ceiling, including on your loudest lines. A recording that is too quiet can be brought up afterwards. A clipped one cannot be repaired.

Record in short blocks, and keep a safety take

Doing the whole thing in one pass is tempting. In practice, a mistake near the end sends you back to the top, the voice tires, and the edit becomes archaeology.

Record paragraph by paragraph. When you stumble, do not stop the recording: leave a clean pause, snap your fingers so the moment leaves a visible marker in the waveform, and pick the sentence up from the beginning. That snap saves a great deal of time when you sort the takes.

On the number of takes, instinct misleads. The first is often the truest: it has the freshness, it is not yet over-controlled, and it is very often the one kept after many others have been heard. On a short script, two or three versions with different intentions give you a real choice. On a long text, limit yourself to one or two complete passes and accept a few small imperfections. A tired voice is audible.

Then keep a safety take: one extra pass, recorded when everything already seems good, filed away untouched. You will come back to it the day someone asks you to change the rhythm of one sentence, long after delivery.

Naming and filing: the thankless work that saves you

Name the file when you create the session: project, language, version, take number. Keep the raw recordings, before any processing, in a folder you never touch again. Deliver in the format asked for, and ask when nobody specified it.

The reason is simple: clients come back. Months later, someone wants one sentence changed inside a training module, and wants the correction to sound exactly like the rest. With the original take and your settings, that is a quick pickup. Without them, it is a new session, and the voice may no longer be quite the same.

What is really at stake rarely comes down to the microphone

Apply everything above and you will get a clean recording. Then comes the moment when the recording has to carry something: a campaign, a brand promise, a training course that colleagues will follow module after module. What is missing then is not a better microphone. It is direction.

Directing does not mean watching the sound. Engineering takes care of the signal. Voice direction listens for something else. Does this sentence mean what it is supposed to mean? Does the actor believe what they are saying? Does the pace leave the listener time to understand?

What a directed session actually adds

An outside ear, first. Alone in front of a microphone, what you mostly hear is your own technical faults: a word badly articulated, a mouth click, a breath sitting too high. You redo for the wrong reason. Voice direction makes you redo a sentence because the intention is not in it, not because an s whistles.

Listening that goes past the voice, second. In the booth, breathing says more than the text: it tells you whether the person is tense, tired, or genuinely settled into what they are saying. A good director hears it before it reaches the performance and acts before the session bogs down. A short break, a glass of water, and if the diction starts to come apart, a few tongue twisters before going back in.

Balance, too. Too much direction and the actor executes instructions instead of playing. Not enough and the session goes in every direction at once. You let actors correct themselves when they can, and you never interrupt mid-sentence unless something is technically wrong: a sentence broken in half is rarely played as well the second time.

Delivery matched to the channel, last. The same text is not prepared the same way for a radio commercial, a web video, a lobby kiosk or a training module: the expected reference level changes, and a file set for a video platform plays too quietly on television. When the voice sits with music, you carve out of the music the band of frequencies where intelligibility is decided, so the voice comes forward without anyone pushing the volume. Those calls belong to the brief far more than to the export, which is why our project types do not follow the same production chain.

From the booth

One thing no software tutorial mentions: you cannot ask for a laugh. Ask for one and you get an acted laugh, audible in half a second. So we make the person laugh for real, and we record what comes after.

Several voices, several sessions: consistency is not improvised

This is where self-recording hits its clearest limit. A training module recorded in three sittings, weeks apart, with a voice that changed in between. A campaign whose hero film and its cutdowns do not sound like they live in the same room. A project with several actors, each in their own studio, with their own microphone and settings. The listener will not name the problem. They will feel something is off, and the credibility of the whole thing drops a notch.

It gets handled upstream, with one technical frame given to everybody before the first take: the same distance to the microphone, the same formats, raw files delivered with no processing at all (processing applied to a take can never be taken back out), the same naming for everyone.

It gets caught downstream by picking one reference take and aligning the others to it, with a limit: correct too much and the voices turn artificial. A slight natural difference sits better than forced uniformity.

Holding a project together over time, across several sessions and several languages, is a large part of what a voice over agency does. Fifteen years of sessions have taught us one thing above all: consistency problems get solved at the brief, and far more rarely in the edit.

What the edit cannot fix

The edit repairs a lot. An over-present breath, a mouth click, a badly measured gap, an unstable level: all of that is treated cleanly. Three problems resist.

An intention that is off. If the read does not believe what it is saying, no processing changes that. You can make a sound more flattering. You cannot manufacture a voice that is convinced. The only remedy is to do it again, and better to notice it in the session than at approval.

A false rhythm. Meaning lives in the silences as much as in the words. A read that is too fast stays perfectly intelligible and gives nobody time to understand. You can stretch a gap in the edit. You cannot put back a breath that was never taken.

A room that rings. Reverberation lives inside the same signal as the voice. Tools attenuate it and almost always leave a trace, and continuous background noise reduces on the same compromise. How far that rescue can honestly go is a subject of its own.

So, should you record it yourself?

Recording a voice over at home is entirely justified when the stakes allow it. A script test, a timing mock-up, an internal video, a pilot to be approved before the real thing: the method above will get you there properly, and it is a genuine skill.

When the voice becomes one of the things carrying the message, the stakes move: the right intention, held from one end to the other, across every session and in every language, then delivered in a format that holds up on the channel you are aiming at. That is a studio’s job, and it is prepared well before the first take.

If you have a project in mind, the simplest thing is to talk it through: the text, the intended use, the languages, the deadlines. Tell us what you need and get a quote. What comes back is a scoping of the project, not a price list.

Book the booth

A directed take, a master matched to your distribution channel, and consistency when the project runs several voices. Tell us what you need.

Get a quote

Read also