You've written a script that looks polished in the document. Then the narrator reaches the second paragraph, runs out of breath, stumbles over a product name, and discovers the final call to action no longer fits the slot. That failure usually doesn't begin in the recording booth. It begins when the script treats voice over as ordinary page copy.
A strong voice over script is a delivery brief. It has to serve the message, the performer or synthetic voice, the runtime, the visual edit, and the listener's limited attention at the same time. The most reliable process is simple: define the job, write for the ear, measure every line, add usable direction, and test the finished copy aloud.
Table of Contents
- Start the Script by Locking Purpose, Audience, and Runtime
- Write Copy That Sounds Spoken, Not Written
- Time Every Line to the Second
- Pick a Skeleton That Fits the Use Case
- Add Voice Direction and Hand Off Cleanly
- Edit, Test Aloud, and Run the Final Audit
- A Repeatable Workflow and Quick Answers
Start the Script by Locking Purpose, Audience, and Runtime
Before opening a blank document, write three answers at the top of the brief:
- Purpose: What single action should the listener take?
- Audience: Who exactly is hearing this?
- Runtime: How many seconds does the finished voice have?
Those constraints shape every later decision. A voice over aimed at a technical engineer can use product architecture and integrations. A script for a non-technical manager needs to explain the practical outcome first. The same SaaS product may therefore require completely different language, even when the feature list stays unchanged.

Turn the brief into a usable decision
Suppose you're writing a 60-second explainer for a project-management tool. The action is trial signup. The listener is a manager who wants visibility across projects but doesn't want a technical implementation lecture. That audience changes the copy:
- “Connect your existing API environment” sounds like an engineering instruction.
- “See every project, deadline, and owner in one place” speaks to the manager's daily problem.
- “Automate your data architecture” describes a system.
- “Spend less time chasing updates” describes a benefit.
The audience also determines tone. Managers may respond to calm, efficient language. Engineers may expect precision about workflow, permissions, or integrations. Neither version is universally better. Each is better matched to its listener.
Fill in the intake before drafting
Use a short intake list so the brief answers practical production questions:
- Objective: What must the listener do, understand, or remember?
- Listener profile: What does this person already know?
- Distribution channel: Is the audio part of an advert, explainer, course, podcast, presentation, or private feed?
- Mandatory phrases: Which product names, legal lines, or calls to action must remain unchanged?
- Forbidden words: Which claims, expressions, jargon, or competitor references are prohibited?
- Runtime: What is the hard ceiling, including the closing beat and any room for music?
Runtime isn't a suggestion. Professional narration is commonly paced at about 130 to 170 words per minute, so a 60-second script usually lands around 130 to 160 words. A 30-second script often lands around 65 to 75 words, while two minutes commonly allows about 260 to 320 words, as outlined by this voice over word count guide. Locking the time early gives you the calculation needed to control the draft instead of discovering the problem after recording.
Write Copy That Sounds Spoken, Not Written
Written copy can rely on visual scanning. A reader can stop, reread, and jump back to the beginning of a sentence. A listener gets one pass, so the voice over needs clean syntax, familiar words, and intentional breathing room.
A useful working range is roughly 15 to 20 words per sentence, with some voice over guidance recommending that sentences stay below 20 words for clarity. You'll find this principle discussed in captivating podcast narration tips, where spoken delivery is treated as a different writing discipline rather than a page being read aloud.
Use the ear as your editor
Consider a written sentence like this:
“Designed to provide organizations with a centralized environment through which teams can efficiently monitor, evaluate, and optimize the status of distributed operational initiatives.”
That sentence may look professional, but it asks one breath to carry too many clauses. A spoken rewrite could be:
“Bring every project into one clear view.
See what's moving.
Fix what's stuck.”
The rewrite uses shorter units and gives the performer room to land each idea. It also creates natural points for a visual edit or a change in emphasis.
Replace abstract business language with words people use aloud:
- “Utilize” becomes use.
- “Commence” becomes start.
- “At this point in time” becomes now.
- “A solution that enables teams to collaborate” becomes “a tool that helps teams work together.”
Contractions usually help. “You'll see the difference” sounds more natural than “You will see the difference,” unless the expanded form is needed for emphasis or legal clarity.
Mark what the voice could misread
Pronunciation notes prevent avoidable pickups. Add a parenthetical guide for brand names, uncommon names, technical terms, and acronyms:
- GIF (say: jif)
- SaaS (say: sass)
- Nguyen (say: win)
For a human narrator, the note removes guesswork. For an AI voice, it can guide pronunciation when the engine supports inline instructions or alternate spellings. Keep the note visually distinct from the spoken copy so nobody accidentally reads it aloud.
Avoid these common traps:
- Stacked clauses: Split one complex sentence into two or three lines.
- Semicolons: Replace them with a full stop or a deliberate pause.
- Em dashes: Use a period or comma instead. They can create inconsistent delivery and may be unsupported by some narration tools.
- Acronyms: Decide whether the listener should hear letters, a word, or a spelled-out phrase.
- Nested conditionals: Replace “If the team has already configured the workflow, it can…” with “First, check the workflow. Then choose the next step.”
Every line should pass a read-aloud test in under one breath. If you stumble, slow down, or need to restart, the sentence is giving you useful editorial feedback.
Time Every Line to the Second
Timing turns creative copy into a deliverable. The basic calculation is words divided by words per minute, then adjusted for pauses, breaths, music, visuals, and the final call to action.
The following planning table uses the established voice over ranges from the script timing guide. These are drafting targets, not substitutes for a timed performance.
Words Per Minute by Voice Over Format
| Format | Target WPM | 15 sec | 30 sec | 60 sec | 90 sec |
|---|---|---|---|---|---|
| Deliberate e-learning or documentary | 130 | 33 words | 65 words | 130 words | 195 words |
| Conversational narration | 150 | 38 words | 75 words | 150 words | 225 words |
| Energetic commercial | 170 | 43 words | 85 words | 170 words | 255 words |
Technical or instructional reads often target 120 to 140 WPM, standard narration commonly sits around 130 to 160 WPM, and energetic commercial delivery can reach 160 to 180 WPM, according to the same timing reference. A script that appears short on the page can still run long once pauses and delivery are included.
For a 60-second conversational spot, start with a working ceiling near 150 words. Then reserve room for intentional beats, breaths, and a closing logo bed or end card. If the call to action needs a clear pause before it, remove words from the setup rather than forcing the final line into a rushed read.
Calculate the cost of a pause
At 150 WPM, a half-second beat consumes about one spoken word's worth of time. That doesn't mean you should mechanically delete one word every time you add a pause. It means you should treat emphasis as part of the budget.
Compare:
“Your team sees every deadline in one place.”
With a beat:
“Your team sees every deadline.
In one place.”
The second version may feel more persuasive, but the pause changes the rhythm. If the line sits beside a visual transition, the beat may be worth the space. If the soundtrack already has a strong musical accent, the extra pause may be unnecessary.
For a 15-second explainer, don't cut randomly from every sentence. Remove entire ideas. Keep the problem, the strongest benefit, and the action. A secondary feature, supporting example, or repeated promise should disappear before the core message does.
You can also use a text-to-audio workflow such as generating audio from text to hear whether the draft's pacing works in practice. Mark pauses consistently, for example (beat, 0.5 sec) or (pause, 1.0 sec), so a human narrator, editor, and AI system interpret the intention the same way.
Pick a Skeleton That Fits the Use Case
Structure controls what the listener hears first, what they remember, and where the action lands. A commercial isn't a compressed course, and an e-learning introduction shouldn't behave like a sales spot.
| Use Case | Beats / Structure | Typical Length | Target WPM | Seconds per Beat |
|---|---|---|---|---|
| Commercial | Hook, benefit stack, single CTA | 30 to 60 seconds | 160 to 180 | Hook 5 to 10, benefits 15 to 35, CTA 5 to 10 |
| Explainer | Problem, solution, how it works, outcome | 60 to 90 seconds | 130 to 160 | Each major beat gets a planned visual interval |
| E-learning or podcast intro | Objective, content beats, recap, next step | Longer format | 120 to 140 | Each beat receives a timestamp |
Commercial structure
Open with the listener's tension, not the company history. Then present the strongest benefit, add only the supporting benefits that help the listener believe it, and finish with one clear action.
A workable sequence is:
- Hook: Name the problem or create a relevant question.
- Benefit stack: Move from primary outcome to supporting proof or feature.
- CTA: Tell the listener exactly what to do.
The CTA should land on the final spoken beat, not compete with a crowded ending. If the audience needs a URL, phone number, or product name, leave enough space for it to be understood.
Explainer structure
An explainer earns attention by answering a sequence of questions:
- What problem is happening?
- What changes with this product or process?
- How does it work?
- What outcome does the listener get?
Use transitions that match the storyboard. “First,” “then,” and “finally” can help when the visuals show a process. Don't narrate every visible movement. Let the screen demonstrate simple actions while the voice explains the reason those actions matter.
For longer study, training, or podcast material, a beat-based structure is more useful. Start with the learning objective, divide the content into clear topics, recap the main points, and give the listener a next step. Timestamp each beat before drafting so the script grows around the content plan instead of expanding until the runtime breaks.
For additional patterns, these podcast script templates can help you compare openings, transitions, and recurring segment structures. The quick rule is practical: below 90 seconds, use a commercial or explainer skeleton. Above three minutes, default to explicit learning or content beats.
Add Voice Direction and Hand Off Cleanly
A script should let the talent or AI narration tool make informed choices without requiring a phone call. Direction belongs close to the line it modifies, but it shouldn't interrupt the spoken copy.
Use a consistent bracket style on its own line:
(warm, conversational)(pace picks up)(beat, 1 sec pause)(land the final word)(slower, instructional)
Specific direction works better than vague notes. “Sound natural” leaves the interpretation open. “Calm, clear, and reassuring, with a short pause before the benefit” gives the reader something they can execute.

Build a handoff top sheet
Put the production facts before the script:
- Project name
- Total runtime
- Target delivery style
- Audience and objective
- Pronunciation guide
- Tone reference link
- Required pickup contact
- Revision and approval rules
For AI narration, keep formatting machine-readable. Write numerals as spoken-friendly words when a synthetic voice might interpret digits unpredictably. Flag homographs such as “lead,” “object,” or “record” when context alone may not produce the intended pronunciation. Avoid unsupported symbols, decorative punctuation, and stacked parentheticals.
A multilingual pipeline needs a different kind of note. Idioms, jokes, metaphors, and compact slogans often don't transfer well. Mark them for adaptation and provide an approved alternate meaning, rather than asking a translator to preserve a phrase that only works in English.
Separate direction from narration
A clean handoff might look like this:
Direction: (warm, conversational)
Narration: “Your next update shouldn't take an entire afternoon.”
Then:
Direction: (slight emphasis on “one place”)
Narration: “Bring the work together in one place.”
This layout reduces accidental reads of production notes. It also gives an AI operator a clear place to convert direction into supported controls.
For presentation-based production, guidance on doing a voiceover on Google Slides can help align narration with visual delivery. Include the practical details in the handoff, such as the requested file naming pattern, audio format, pickup contact, and which changes require a new approval.
Edit, Test Aloud, and Run the Final Audit
Many drafts sound fine in the writer's head because the writer supplies the missing pauses and corrects awkward phrasing while reading. The microphone doesn't make those corrections. A real read-aloud pass exposes problems that a screen hides.
Cut for structure before sound
First, ask whether each line earns its time. A commercial line should support the hook, benefit, proof, or CTA. An instructional line should teach, orient, reinforce, or move the learner to the next action. If it serves none of those functions, cut it.
Then read the remaining copy at performance pace. Listen for:
- Tongue-twisting sound combinations.
- Repeated hard consonants across neighboring words.
- Long strings of technical terms.
- Awkward breaths and sudden speed changes.
- Phrases that sound formal only because they were written, not spoken.
Split a difficult line across a breath, replace a dense synonym with a familiar word, or move a proper noun to a sentence where it can receive emphasis. Don't preserve elegant wording if the narrator has to fight it.
Use a final approval gate
Run this audit before recording or generation:
- Runtime fit: The complete read fits the locked ceiling.
- Tone match: Direction describes a specific, achievable delivery.
- Pronunciation: Names, acronyms, and technical terms have approved guidance.
- Breathing room: Important claims and transitions have usable pauses.
- Listener clarity: Each sentence communicates one manageable idea.
- Visual alignment: The narration doesn't describe an action before the screen is ready.
- AI compatibility: Unsupported punctuation, symbols, and ambiguous spellings are removed.
- Multilingual readiness: Idioms and culturally specific phrasing are flagged for adaptation.
- Approval status: The client or content owner has a clear sign-off line.
Log lines that run over or under by more than half a second during the timed read. That threshold is useful because small differences compound across a sequence, especially when voice direction, animation, or music cues depend on precise entrances and exits.
Treat the audit as the gate between draft and recording. It isn't a cosmetic polish step. It determines whether the production team receives usable copy or a problem that becomes more expensive once audio and visuals are involved.
A Repeatable Workflow and Quick Answers
A reliable voice over process doesn't need to feel elaborate. It needs to happen in the same order every time, with each stage producing an artifact the next stage can use.
- Lock the brief: Write the purpose, audience, channel, mandatory language, restrictions, and runtime.
- Choose a skeleton: Select the commercial, explainer, or learning structure that matches the use case.
- Draft on paper: Write the listener-facing copy without polishing every adjective.
- Time to runtime: Count words, read aloud, and remove or expand ideas based on actual delivery.
- Add voice direction: Mark pace, tone, emphasis, pronunciation, and intentional pauses.
- Run the audit checklist: Confirm clarity, timing, handoff quality, and approval status.
Give each stage a fixed working window that fits the project, rather than allowing the first draft to absorb the entire schedule. The script is ready to hand off when the brief is complete, the structure is visible, the timed read fits, and the final audit has no unresolved production questions.
Quick answers for working writers
How many words should a 60-second script contain?
Standard narration commonly lands around 130 to 160 words per minute, so a 60-second script usually sits in that range. Reserve words or time for pauses, breaths, music, and the closing action when the edit requires them. See the voice over word-count benchmark before locking the draft.
Should you number lines?
Number lines when the script must align with scenes, slides, captions, or revision comments. For straight audio, clean paragraphs with direction on separate lines are often faster to read.
What happens when the client changes the runtime?
Don't squeeze or stretch every sentence. Recalculate the word budget, then remove or restore complete ideas while protecting the hook, core benefit, and CTA.
How should you handle multiple narration voices?
Give each voice a clear label and separate their lines visually. Define the role, tone, pace, and handoff rules for each speaker before writing dialogue.
Which scripts need pronunciation glossaries?
Create one for branded terms, names, specialist vocabulary, acronyms, and any word that may change pronunciation across languages. AI narration benefits from the same preparation, even when the voice sounds highly natural.
For genre-specific inspiration, these script examples for narration are useful for comparing how commercial, explainer, and educational copy behave on the page. Study the structure, then test your own lines aloud. Natural delivery comes from decisions you can verify, not from making the document look polished.
Rooy Development creates AI-generated podcast episodes from selected sources, with structured scripts, natural two-host narration, recurring delivery, and multilingual production support. Visit Rooy Development to turn your source material into timed, listener-ready audio without rebuilding the workflow manually for every episode.
