Skip to content

QC multiple takes of the same line

When a director has an actor deliver a line several times, you can QC every take without doing anything special. Each take is already its own audio file, so each one gets its own transcription and its own comparison against the expected line. The only thing you have to get right is the script CSV, which needs one row per audio file rather than one row per line.

VoiceQC does not need to know that a group of files are takes of the same line. It checks files. Choosing which take ships happens in your own pipeline. What you get from here is which takes are usable.

This is the step people get wrong. Each take needs its own row in the script CSV, repeating the same expected text.

A tracking sheet often has one row per line and a column per take:

Line IDScriptTake 1Take 2Take 3
L042You know who I am.L042_t1.wavL042_t2.wavL042_t3.wav

VoiceQC needs it the other way up, one row per file, with the line repeated:

filename,script_text
L042_t1.wav,"You know who I am."
L042_t2.wav,"You know who I am."
L042_t3.wav,"You know who I am."

Repeating the text is correct. Nothing checks that script lines are unique.

Download takes-template.csv to start from a working file. See Script CSV format for the full rules on quoting, headers and multi-line text.

  1. Drop your audio files and the CSV on the upload step. Order does not matter, because they join on filename.

  2. On Confirm your uploads, read the report before continuing. It separates files that matched, script rows with no audio file, and audio files with no script row.

  3. Fix anything listed there before you continue, while it is still free to do so. After this point the items are spent.

This report matters more with takes than with single reads. Every line multiplies the number of filenames you have to get right, so one mistyped suffix shows up here as an orphaned script row next to an orphaned audio file.

Turn both built-in Auto-rules on and let them sort the batch.

The Non-Verbal rule flags takes that are not really speech. Multi-take recording produces a lot of these: false starts, breaths, slates, room tone, and an actor warming up before the real read.

The Mismatch rule flags takes whose transcription does not match the script. How far a take diverges tells you what went wrong:

What you seeWhat it usually means
A word or two offThe actor dropped, added or altered a word. A performance note.
Nothing matches at allThe wrong audio is under that filename: files swapped during delivery, a take that was never recorded, or a mistyped suffix.

The second case is the one that is hard to catch any other way. A total mismatch on L042_t3.wav means the file is wrong rather than the read, and a shifted or duplicated file is easy to deliver by accident when there are five takes a line.

Review the flagged takes with Highlight differences turned on and step through them in Focus view. Bind keyboard shortcuts to the Property Values you use for verdicts to speed up a pass. Takes that no rule flagged do not need your attention.

None of these announces itself, so check for them first.

Two files with the same name become one item. A Session item is identified by its filename with the extension removed. Uploading two files with the same base name does not create two items and does not raise an error. The second file replaces the first. This is easy to hit if you keep takes in a folder per line:

L042/take1.wav
L043/take1.wav

Uploading flattens the folders, so both arrive as take1.wav and you get one item instead of two. Put the line in the filename rather than only in the folder name.

Line numbering that restarts collides the same way. Most deliveries number lines per character, or per language, so Betty/001.wav and Anne/001.wav both arrive as 001.wav and become one item. EN/L042.wav and FR/L042.wav do the same. Put the character or the language in the filename, or give each one its own Session.

Re-uploading a filename updates that take rather than adding one. If you re-record a line and deliver the fix under its original name, VoiceQC replaces the audio on the existing item and transcribes it again. That is what you want when you are correcting a bad file. It is not what you want if you meant to keep the original take next to the new one, because the earlier transcription is replaced along with the audio. Give the new recording its own name when you need to compare the two.

The extension is ignored. L042_t1.wav and L042_t1.mp3 are the same item, for the same reason. Deliver one format per take.

Capitalisation counts, for now. L042_T1 and L042_t1 are currently treated as two different items, so a CSV typed by hand will not match audio named by a DAW if the casing drifts. The Confirm step warns you when two names differ only by capitalisation. Match them exactly. This is being fixed.

Column order matters. Filename first, script text second. A reversed file is rejected outright rather than read the wrong way round.

Duplicate rows are last-wins. If the same filename appears twice in your CSV, the later row wins and the Confirm step tells you which rows it replaced.

Audio uploads are .wav only, and one script CSV holds up to 10,000 rows. Takes reach that ceiling about five times faster than single reads. A 2,000-line script at five takes each is exactly 10,000 rows.

Every take is its own Session item, because the unit of work is transcribing and checking one file. A 400-line script at five takes each is 2,000 Session items, not 400.

Plan for that before a large batch. Many studios bill their own customer per delivered line with all takes included, so the two models count differently. If you only need QC on the takes you are shipping, upload those. If you are looking for a file mix-up among all of them, upload all of them. Check Plans and limits first.