Skip to content

Noise & Masking

An audible signal succeeds only if a listener can detect it, distinguish it from competing sounds, understand what it means, and act in time. Increasing volume addresses only part of that chain. The signal can still be masked by nearby frequencies, confused with another alert, distorted by the speaker, missed by a person with hearing loss, or suppressed because the device is silent.

Design for noise and masking by controlling the soundscape, measuring the delivered signal, and providing persistent visual or tactile equivalents. Audio should add information and urgency; it should not become a single point of failure.

Masking occurs when one sound raises the threshold at which another sound can be heard. The result depends on frequency content, level, timing, and the listener—not just a single decibel reading.

Sounds that occur together compete within the auditory system. Competition is usually greater when their energy occupies nearby frequency regions. A broadband fan, crowd, road rumble, or music track can cover part of speech or a notification even when its overall level seems moderate.

A strong sound can reduce perception of a quieter sound immediately before or after it. A short notification placed against a loud transition, door slam, musical accent, or another alert may be missed even if a sound-level meter shows an adequate average.

Competing speech is especially disruptive because it is meaningful and changes over time. Two voices with similar location, level, and timbre are harder to separate than a voice and a steady mechanical sound. Product tests that use only pink noise can therefore overestimate intelligibility in offices, homes, or public spaces containing conversation.

The audio file is not what reaches the listener. The operating system may mix it with other apps; the output may be a small speaker, one earbud, hearing device, Bluetooth system, or remote display; and the listener’s thresholds may differ between ears and across frequencies.

NIDCD explains that perceived loudness is affected by duration, frequency, and environment, and that the decibel scale is logarithmic. ASHA’s hearing-loss guidance shows why “make it louder” is not a complete solution: hearing loss can have different frequency configurations rather than one uniform attenuation.

Signal-to-noise ratio (SNR) is the difference between the level of the wanted signal and competing sound, measured using a defined method and time window. A positive SNR means the signal level is higher; it does not by itself prove the signal is intelligible or recognisable.

For recorded speech, WCAG 2.2 Success Criterion 1.4.7 provides a useful Level AAA benchmark: prerecorded audio-only content that is primarily speech must have no background sound, allow it to be turned off, or keep background sound at least 20 dB below the foreground speech, apart from brief exceptions. That criterion has a defined scope; do not apply its 20 dB number as a universal alarm or live-conversation standard.

For safety, public address, or professional communication, use the domain’s assessment method. ISO 9921:2003 covers assessment of speech communication for alerts, danger signals, information messages, and general communication. IEC 60268-16:2020 defines the Speech Transmission Index (STI), test signals, measurement, and prediction methods. These methods account for more than a media file’s peak level.

Measure the context before designing the cue

Section titled “Measure the context before designing the cue”

Lists such as “coffee shops are 70 dB” are too crude for a product requirement. Sound levels vary with location, time, crowd, equipment, measurement distance, frequency weighting, and averaging. Measure or obtain representative recordings from the actual contexts in scope.

Document:

  • output path and listening position;
  • background sources and whether they are steady, intermittent, or speech-like;
  • sound level, frequency weighting, and time weighting;
  • frequency spectrum or octave-band data where masking matters;
  • whether the product can control, pause, or duck competing audio;
  • users’ ability to change volume, output, captions, and haptics; and
  • the consequence and required response time if a cue is missed.

NIOSH’s noise guidance distinguishes a sound level at a point in time, an eight-hour time-weighted average, and a cumulative noise dose. The metric must match the question. An occupational exposure measurement is not automatically the right measure for a 300-millisecond notification, and a peak measurement does not establish speech intelligibility.

Inventory every sound before composing new ones:

ClassExampleRequired behaviour
Safety-criticalEvacuation or collision warningGoverned by relevant safety standards; redundant and tested in context
ConsequentialPayment failure, medical workflow exceptionDistinct cue plus persistent visual state and recovery path
StatusUpload complete, message sentSubtle optional cue plus visible confirmation
Ambient or expressiveMusic, atmosphere, decorative responseMust never mask speech or required alerts; user-controllable

Do not use a louder sound merely to make a low-priority event feel important. A crowded alert vocabulary teaches users to ignore all of it.

For each audible cue:

  • use a recognisable temporal pattern, not pitch alone;
  • separate events by rhythm, duration, timbre, and context;
  • avoid relying on a very narrow or very high-frequency component;
  • keep the onset clear without creating a startling transient;
  • reserve the most salient pattern for the most important class; and
  • use the platform’s established sound where it already conveys the intended meaning.

Frequency separation can help, but “choose a frequency outside the noise” is rarely sufficient. Real environments and speech are broadband, and a frequency that escapes one background may be poorly reproduced by a phone speaker or inaudible to a listener with high-frequency hearing loss.

When the product owns both foreground and background tracks:

  • remove nonessential background audio during instructions;
  • let users turn the background off;
  • duck it smoothly before speech or a consequential cue;
  • avoid two spoken streams at once;
  • limit reverberation and effects on speech; and
  • restore the background smoothly after the message.

When another app owns the audio, follow platform audio-session conventions. Apple’s playing-audio guidance describes how audio categories affect mixing, interruption, silent mode, and background playback. Do not seize audio focus for decorative feedback or unexpectedly stop content the user chose.

Raising level to overpower the environment can increase hearing risk. WHO’s safe-listening guidance emphasizes that risk depends on sound level, duration, and frequency of exposure and recommends reducing level, using well-fitted noise-cancelling headphones in noisy conditions, and taking listening breaks.

If users routinely maximize volume to understand the product, improve recording, noise reduction, mixing, captions, and output options. Do not describe occupational limits as safe consumer playback targets: NIOSH defines its 85 dBA eight-hour recommended exposure limit for workplace noise.

Audio may be unavailable because the user is Deaf or hard of hearing, the device is muted, the speaker is covered, the output changed, the room is noisy, or playing sound would be socially inappropriate.

Pair a sound with a visual state that remains long enough to find and understand:

  • notification banner with clear action and status;
  • inline error beside the affected field;
  • badge or activity log for events that can be reviewed later;
  • progress state that changes visibly on completion; and
  • optional haptic pattern that respects system settings.

Apple’s accessibility guidance recommends augmenting audio cues with visual cues and, where appropriate, haptics. Haptics are not a complete replacement: they can be disabled, unavailable, or difficult to distinguish.

WCAG 2.2 Success Criterion 1.2.2 requires captions for prerecorded audio in synchronized media at Level A. Captions include relevant non-speech information and speaker identification, not dialogue alone. Live synchronized media is covered by Success Criterion 1.2.4.

Also provide a transcript when it helps searching, scanning, translation, or review. Make player controls keyboard accessible and give users independent control of captions and volume. Do not autoplay audio that interferes with screen-reader speech; WCAG’s Audio Control criterion addresses audio that plays automatically for more than three seconds.

Never say only “continue after the beep” or distinguish actions only as “the high tone” and “the low tone.” WCAG 2.2 Success Criterion 1.3.3 requires instructions not to rely solely on sensory characteristics including sound. Name the action and show the state.

Worked example: a transit disruption alert

Section titled “Worked example: a transit disruption alert”

A journey app initially plays the same short chime for a platform change, a promotional message, and a completed download. The platform-change banner disappears after three seconds. In a moving train, users miss the chime or cannot tell what it means.

The team defines the outcome:

A rider can detect, identify, review, and acknowledge a platform change with audio disabled, while using the device speaker or headphones, and in representative transit noise.

The design and measures become:

Failure modeDesign responseVerification
Cue is maskedUse a short, platform-appropriate consequential cue; pause app-owned speech firstDetection rate across recorded and live representative conditions
Cue is confused with low-priority eventsRemove promotional sound and give platform changes a unique rhythm and statusIdentification without looking, with no misleading associations
Banner disappearsPersist the change in the journey timeline until reviewedSuccessful review after a delayed response
Audio is unavailableShow lock-screen notification, in-app banner, changed platform label, and optional hapticComplete the task with device muted and haptics off
Message is not actionableState old platform, new platform, effective time, and next actionComprehension and correct-route selection
Repeated alerts encourage unsafe volumeKeep user volume control; improve message and redundancy rather than forcing levelPlayback behaviour and safe-listening review

Safety decisions for an actual transport service also require the operator’s standards, public address design, and operational risk process. An app cue cannot replace an official warning system.

Use several backgrounds rather than one generic noise track:

  • steady mechanical or ventilation noise;
  • competing speech;
  • intermittent transients;
  • app-owned music or effects;
  • another notification near the cue; and
  • quiet conditions, where an overly harsh cue may become unacceptable.

Calibrate playback and document the output device, position, level method, and room. A laptop playing a phone recording at an arbitrary volume is not a repeatable test.

Signal typePrimary measures
NotificationDetection, correct identification, response time, false alarms
Spoken instructionWord or sentence intelligibility, task comprehension, action accuracy
MediaCaption accuracy, foreground/background level, listening effort, user control
Safety messageDomain-required intelligibility and coverage method, response accuracy

Measure with the actual encoded asset and output chain. Compression, automatic gain control, speaker protection, Bluetooth codecs, and mono downmix can all change the result.

Recruit participants with relevant hearing experience and test the accessibility modes they use. Filtered simulations can expose a cue that depends entirely on one frequency region, but simulation is not evidence that people with hearing loss can understand the experience.

Check:

  • device muted, low volume, and changed output route;
  • one channel only and mono downmix;
  • captions at enlarged text sizes;
  • hearing aids or supported hearing devices where appropriate;
  • visual-only and haptic-off completion; and
  • notification history after an interruption.
  • Is each sound’s purpose and priority documented?
  • Were real or representative contexts measured rather than assigned generic dB labels?
  • Does the signal remain distinguishable in steady, speech-like, and intermittent noise?
  • Can app-owned background audio be reduced or disabled?
  • Is speech assessed with an appropriate intelligibility method?
  • Does every meaningful sound have a visual or text equivalent?
  • Are captions complete, accurate, synchronized, and user-controllable?
  • Can the task be completed with sound and haptics disabled?
  • Does the product avoid pushing users toward unsafe listening levels?
  • Were the final asset, encoding, device, output route, and mono behaviour tested?

  • Frequency Ranges — Design spectra that survive devices and hearing variation
  • Hearing — Apply the broader hearing-accessibility model
  • Haptics — Add optional tactile feedback without replacing visual status