Filler word counter

Paste a transcript, an SRT or a VTT. The filler word counter returns how many vestigial sounds it contains — um, uh, erm, ähm — how many context-dependent words like so and actually, and how many minutes of runtime those words occupy, derived from your own speaking rate rather than an average.

Given a transcript or caption file you paste, produce the filler-word count and the runtime those words cost, in under one second.

4 vestigial · 0m 05s

73 words at 46 words a minute. Another 6 are context-dependent — counted separately below, because “so” opens sentences and “actually” sometimes means something. Cutting every one of them would take back 0m 13s.

wordcountkind
um2always filler
uh2always filler
so2depends on the sentence
like1depends on the sentence
yeah1depends on the sentence
basically1depends on the sentence
actually1depends on the sentence

Vocabulary: um · uh · uhm · erm · hm · mhm · äh · ähm — and like · so · yeah · basically · actually · halt · quasi · sozusagen · genau · ja · also

The count is the whole count. Removing them is a different job: every one you delete leaves a joint, and a joint cut on the transcript's own timing is what makes an edit sound clipped.

Next artifact: the same recording with those seconds gone.

Never's selects pass removes filler and dead air using this same vocabulary, then moves each resulting join off the transcript timing onto the measured acoustic onset so the joins are inaudible.

Read why the joins matter

Your transcript is counted in your browser. It is never sent anywhere, never stored, and there is nothing to delete.

What a filler word costs, in runtime and in attention

A filler word costs twice. It costs the runtime it occupies, which the filler word counter measures directly from the speaker's own rate, and it costs the sentence's momentum, which nothing measures but everyone hears. In a 95-second clip with 20 vestigial fillers at a normal speaking rate, the first cost is around eight seconds — roughly a tenth of the clip, spent on nothing.

The German half of the vocabulary is here because the same recording often contains both languages. Äh and ähm are the direct equivalents of uh and um; halt, quasi and genau behave like basically and actually, which is to say sometimes they carry meaning and sometimes they are furniture.

How to count and cut filler by hand

  1. Work from a transcript, not the timeline. Reading is faster than scrubbing, and a transcript makes the pattern visible — most speakers have one or two habits, not twenty.
  2. Search rather than scan. Search the transcript for each sound in turn. Scanning misses the ones inside sentences, which are the majority.
  3. Mark, then cut in one pass. Cutting as you read means re-finding your place forty times. Mark every timestamp first.
  4. Listen to each join. This is the step people skip and the one that decides whether the edit sounds edited. Every removal leaves a joint, and a joint placed on the transcript's own timestamp lands inside the next word.

Where the manual way breaks

Removing filler by hand breaks at the joins, not at the finding. Finding twenty ums in a transcript takes five minutes. Cutting them leaves twenty joints, and each one has to be listened to, because a cut placed on a word's transcript timestamp lands 50 to 100 milliseconds inside the word and clips its attack.

That gap is the whole reason AI-edited clips sound choppy, and it is measurable rather than a matter of taste — the cut-point checker measures it on your own audio, in your browser. The Editing Tax dataset puts removing filler, dead air and bad takes at 45 minutes for a single 12-minute recording, and almost all of that is the listening, not the finding.

Never's selects pass uses this same vocabulary to decide what a removed span was, then moves each resulting boundary outside the measured acoustic onset before the frame-grid snap, so the joins are inaudible without anyone auditioning them one at a time.

Filler word counter · made with Never
Read the transcript

Filler word counter Paste a transcript, an SRT or a VTT. The filler word counter returns how many vestigial sounds it contains — um, uh, erm, ähm — how many context-dependent words like so and actually, and how many minutes of runtime those words occupy, derived from your own speaking rate rather than an average.

Frequently asked questions

Which words does the filler word counter treat as filler?

Two lists, kept apart. Vestigial sounds — um, uh, uhm, erm, hm, mhm, äh, ähm — are filler every time. Context-dependent words — like, so, yeah, basically, actually, and their German equivalents — are counted separately, because so opens legitimate sentences and actually sometimes contradicts something.

Why does it need the runtime?

To turn a count into seconds using your own speaking rate rather than an invented average. Total words divided by runtime gives words per second for this speaker on this recording; the filler count divided by that is the time those words occupy. Paste a caption file and the runtime is read from the last cue.

Should every filler word be cut?

No, and a tool that removed all of them would make the edit worse. A beat before a difficult sentence carries meaning. Over-tight cutting is the most reliable tell of an AI edit — which is why the two lists are reported separately instead of as one number to minimise.