Creatos Logo
Buy License
AI Notes

AI voiceover script: pronunciation and timing handoff template

Prepare an AI voiceover with separate spoken copy, pronunciation notes and picture cues. Includes a 51-word example and a transparent timing calculation.

Published
4min read
Filed under
Three separate voiceover inputs: a 51-word spoken script, Luma pronunciation notes, and picture cues measured after rendering.
Original Creatos diagram of the article’s illustrative workflow. It does not show measured model output.

Sources checked:

Prepared with AI assistance and source review on September 16, 2026. Examples are illustrative, not provider benchmark results.

On this page

An AI voiceover brief needs three things: the words to say, the pronunciation of unusual terms and the places where the audio must meet the picture. Keep those separate. Otherwise a direction such as “pause here” can become part of the spoken script, or a corrected product name can change the timing of every later scene.

StepFun's StepAudio 3 Gen technical report describes a model that supports several kinds of audio generation, including speech and voice design. It is useful context for the expanding choice of audio tools, but a capability list does not tell you whether your brand name or final sentence will sound right. StepAudio 3 Gen technical report

This article provides a voiceover handoff sheet, with a fictional product example. It is not a listening test of StepAudio 3 or a claim that Creatos generates this audio.

Write a script you can actually record

Consider a short demonstration for a fictional note-taking app called Luma Note. The invented pronunciation is “LOO-muh note.” The video shows a saved idea, a linked image and an exported board.

Here is the spoken script:

A good idea rarely arrives with everything attached.
Save the note. Add the image that explains it.
Keep the source beside the draft, so you can find it again.
When the board is ready, export it for the next conversation.
That is Luma Note: a place to keep the work together.

The example contains 51 words if punctuation is ignored and “Luma Note” counts as two words. At an assumed 130 words per minute, the words alone take about 23.5 seconds: 51 divided by 130, multiplied by 60. That is a planning estimate, not an audio measurement. Pauses, emphasis and the selected voice will change it.

If you have only 20 seconds of picture, shorten the script before trying to force the voice faster. For example, remove the first sentence and begin at the action: “Save the note.” Do not promise a precise running time until you have rendered and listened to the audio.

Add a pronunciation sheet outside the spoken copy

TermIntended readingCheck in the rendered audio
Luma NoteLOO-muh note, invented brand pronunciationTwo words; no extra syllable
AIThe letters A and IDoes it become a single word?
v2“version two” in this exampleDo not read “vee squared”
1.5 GB“one point five gigabytes”Decimal and unit remain intact

Only include terms used in the final script. The extra rows illustrate how to prepare a reusable sheet for later videos. A plain-language pronunciation note is a production instruction, not a promise that a particular model will obey it. Use a tool's documented pronunciation controls if available, and listen to the result.

Keep directions in a separate field:

Delivery: calm, conversational; no announcer-style finish.
Pronunciation: Luma = LOO-muh (fictional brand).
Read only the approved spoken-copy field.
Do not speak headings, timing notes or filename labels.
Return uncertainties before generating another take.

Tie the audio to visible actions

Make a cue sheet after the first render. Until then, label all timings “planned.” The editor needs to know which words must line up with a visible change, rather than receive an arbitrary instruction to fill exactly 30 seconds.

CueVisible eventTiming status before the render
“Save the note”Note appears on the boardPlanned; align to the spoken phrase
“Add the image”Reference image is addedPlanned; allow the action to finish
“export it”Export menu opensPlanned; keep the result visible

After rendering, replace the planned cues with measured timecodes from the actual file. Keep the original audio untouched and save revisions with a take number. If only the product name is wrong, generate or record a replacement and check the surrounding transition; do not silently swap a full new take under an already edited video.

Listen once for meaning, once for the edit

The meaning pass checks every number, name, negation and missing word against the approved script. The edit pass checks breaths, clipped endings, unwanted silence and the relation to the picture. Listen to the exported video as well as the audio file; a correct voiceover can still be cut short on the timeline.

For a shared working sheet, put the spoken copy, pronunciation table and cue notes in separate Creatos text nodes. The flow-editor guide covers organizing a canvas and exporting a visual reference. Keep the rendered audio and final timing decisions in your audio or video editor.

Prepared on September 16, 2026. The script, invented brand and timing calculation are original examples; no provider-generated take was used to establish the estimates.

Sources