A second of silence at the head of fifty narration clips is nearly a minute of learners waiting for something to happen, and it is the most common reason a slide feels sluggish. This finds the speech in every file and trims to it, leaving a little room so the first consonant is not clipped.
Free · No account · Your files never upload
Trimmed on your computer and handed back as WAV. Nothing is uploaded.
Finds where the speech starts and stops in each file and trims to it, leaving a configurable amount of room either side. Handed back as WAV.
A second of silence at the head of a clip is barely noticeable on its own. Across fifty narration files it is nearly a minute of learners looking at a slide waiting for something to happen, and it is the commonest reason a course feels sluggish without anybody being able to point at what is wrong. Doing it by hand is fifty trips through an audio editor.
Cutting at the exact first audible sample removes the attack from the opening consonant, and a plosive that starts mid-burst sounds abrupt and slightly wrong. Sixty milliseconds is about the shortest gap a listener does not notice, and it is the default here for that reason rather than as a safety margin.
Lower keeps quiet room tone, which is usually what you want, because a recording that cuts to digital silence between phrases sounds unnatural. Higher cuts closer and risks clipping a softly spoken first word. If a file comes back marked as silent all the way through, either the recording failed or the threshold is set too high for it.
It trims the ends. It does not remove pauses inside the take, which is an editing decision rather than a cleanup one, and which changes the pacing the narrator intended.