Skip to main content
← Back to Blog
9 min readBrassTranscripts Team

How Speech Quality Is Measured for Transcription

You cannot fix what you cannot measure, and for a long time "good audio" was a matter of opinion. BrassTranscripts starts from a more useful fact: recording quality can be scored before transcription, and that score predicts how accurate the transcript will be. The audio you upload already carries measurable signals of noise, distortion, and dropouts, and each of those signals maps to a specific way a transcript can go wrong.

This post explains how speech quality is actually measured in the field, what the numbers mean, and why a quality score is really an accuracy forecast in disguise. BrassTranscripts maintains a Curated Authority Index of the primary research behind AI transcription, including the audio quality research this article draws on, so you can check every claim at the source.

Quick Navigation

Quality Can Be Scored Before You Transcribe

BrassTranscripts treats audio quality as a measurable property of the recording, not a subjective impression formed after reading a bad transcript. Modern speech-quality models are non-intrusive, meaning they estimate perceived quality from the degraded recording alone, without needing a pristine reference copy to compare it against.

That distinction matters more than it sounds. Older quality measures required the original clean signal to measure how far a recording had drifted from it, which is useless in the real world where no clean copy of your meeting or interview exists. A non-intrusive model looks only at the file you actually have. The leading example is NISQA, described in "NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets" by Gabriel Mittag, Babak Naderi, Assmaa Chehadi, and Sebastian Möller of TU Berlin (2021), published at arxiv.org/abs/2104.09494. Because it needs no reference, a score like this can be produced for any recording before it is ever transcribed.

What a Mean Opinion Score Actually Measures

A Mean Opinion Score, or MOS, is a single number that captures overall perceived speech quality, and BrassTranscripts uses it as shorthand for how clear a recording sounds to a listener. Historically the MOS was gathered by asking human listeners to rate audio on a simple scale and averaging their responses; the achievement of modern models is predicting that human rating automatically.

The overall MOS is a useful summary, but a summary hides detail. A recording that scores poorly could be too quiet, too noisy, distorted, or full of dropouts, and the single number treats all of those as the same generic "bad." NISQA was built specifically to reject that oversimplification. Alongside the overall MOS, it predicts four separate quality dimensions, so the score tells you not just that a recording is degraded but in which way it is degraded. That is the difference between a warning light and a diagnosis.

The Four Dimensions of Speech Quality

BrassTranscripts finds the four-dimension breakdown more actionable than any single rating, because each dimension names a defect you can actually go and fix. NISQA predicts Noisiness, Coloration, Discontinuity, and Loudness in addition to the overall MOS.

Each dimension isolates one failure of the recording chain. Noisiness reflects background sound layered over the speech, the hum, chatter, or traffic competing with the voice. Coloration reflects distortion of the frequency balance, the tinny or muffled quality that comes from a poor microphone or aggressive compression. Discontinuity reflects interruptions in the signal, the gaps and dropouts that appear when audio is lost in transit. Loudness reflects level problems, speech that is too quiet or clipped from being too hot. A recording can score well on three and fail the fourth, which is exactly the information a single MOS throws away. The four-dimension model was introduced in the NISQA work by Mittag and colleagues (2021).

Why Quality Predicts Transcript Accuracy

The reason to measure any of this is that input quality is the single biggest driver of transcript accuracy, and BrassTranscripts treats a low quality score as an early warning that the transcript will need scrutiny. Clean input audio consistently produces the most accurate output, a principle so well established it borders on obvious once you have compared a studio recording against a phone call of the same conversation.

What the dimensional view adds is a prediction of the failure mode, not just the failure. A low Discontinuity score suggests packet-loss gaps, the kind of dropout that erases whole words and leaves the transcript missing content entirely. A low Noisiness score suggests background noise, which tends to produce misheard words rather than missing ones, as the model tries to resolve speech buried under competing sound. Coloration problems blur the acoustic distinctions between similar-sounding words, and loudness problems can push quiet speech below the threshold where it registers at all. Each dimension predicts a different way the transcript will disappoint you, which is far more useful than a vague sense that the audio was "not great." Our guide to what determines transcription accuracy covers why the recording matters more than the brand of software.

How the Measurement Handles Real-World Audio

A quality model is only trustworthy if it was tested on the audio people actually record, and BrassTranscripts values the NISQA research precisely because it was built on real conditions rather than laboratory tone. Its corpus contains more than 14,000 speech clips spanning a wide range of distortions, including live recordings made over mobile phone, Zoom, Skype, and WhatsApp.

That coverage is what makes the scores relevant to your uploads. A meeting recorded through Zoom, an interview captured on a phone, and a voice note sent over WhatsApp each degrade in characteristic ways, and a model trained on those exact channels predicts their quality reliably rather than guessing. The crowdsourced datasets behind NISQA were designed to capture this real-world variety, which is why its predictions hold up outside the lab. For a deeper look at the vocabulary around all of this, see our AI transcription glossary.

How to Use Quality Measurement Before You Upload

You do not need to run a research model to benefit from thinking in these terms, and BrassTranscripts recommends listening to your recording through the four-dimension lens before you transcribe. Ask whether the speech is buried in noise, whether it sounds distorted, whether the audio cuts in and out, and whether it is loud and clear, because those four questions map directly to the dimensions research uses to predict quality.

Where you can, fix the problem the diagnosis points to rather than transcribing and hoping. If dropouts are the issue, re-export from the original source or record locally instead of over a call. If noise is the issue, move somewhere quieter or closer to the microphone. Our practical walkthroughs of audio quality secrets for perfect transcription and fixing the audio problems ruining your transcripts cover the specific fixes. When you are ready, the surest test is your own audio: BrassTranscripts shows a 30-word preview of every transcript before purchase, so you can confirm accuracy on the exact file you care about rather than trusting any score in the abstract.

Frequently Asked Questions

Can speech quality be measured before transcription?

Yes. Non-intrusive quality models estimate perceived speech quality from the recording alone, without needing a clean reference copy to compare against. BrassTranscripts treats this the same way a professional would: the recording carries measurable signals of clarity, noise, and dropouts that exist before a single word is transcribed, and those signals predict how the transcript will turn out.

What is a Mean Opinion Score?

A Mean Opinion Score, or MOS, is a single rating of overall perceived speech quality, traditionally averaged from human listeners and now predictable by models. The NISQA research goes further than one number, predicting four separate dimensions alongside the overall score. BrassTranscripts finds the multidimensional view more useful because a bad recording is rarely bad in only one way.

What are the four dimensions NISQA predicts?

NISQA predicts Noisiness, Coloration, Discontinuity, and Loudness in addition to an overall MOS. Each isolates a different defect: Noisiness captures background sound, Coloration captures frequency distortion, Discontinuity captures gaps and dropouts, and Loudness captures level problems. BrassTranscripts maps these to distinct transcript failure modes, so a low score points to the specific problem to fix.

Does a low quality score mean a bad transcript?

Usually, and it also tells you why. A low Discontinuity score suggests packet-loss gaps that erase whole words, while a low Noisiness score suggests background noise that produces misheard words. BrassTranscripts shows a 30-word preview of every transcript before purchase, so users can confirm accuracy on their own audio rather than relying on a score alone.

Where does call and meeting audio fit into quality measurement?

Modern quality models are trained on real communication conditions, not just studio recordings. The NISQA corpus includes clips captured over mobile phone, Zoom, Skype, and WhatsApp, which is why its scores translate to everyday audio. BrassTranscripts sees the same patterns in practice, where the recording channel often matters more than the microphone.

About BrassTranscripts

BrassTranscripts is a pay-per-file AI transcription service with no subscription. Pricing is simple: $2.50 for files 1 to 15 minutes long, and a flat $6.00 for anything 16 minutes and up, at any length. Every transcript includes automatic speaker identification and downloads in TXT, SRT, VTT, and JSON, with support for 99+ languages. Advanced AI transcription handles the recording; you just upload it. Before you pay, a 30-word preview lets you check the quality of your specific transcript, so you can confirm the result on your own audio first. When you are ready, upload a file and see the preview yourself.

Ready to try BrassTranscripts?

Experience the accuracy and speed of our AI transcription service.