/Entity Specific Speech Recognition Using Key Term Based Acoustic Model Tuning
Abstract

A system and method for audio processing. A method includes training an acoustic model over a plurality of training iterations by, at each of the plurality of training iterations: applying the acoustic model to features extracted from training audio data in order to output a set of acoustic model predictions; applying a language model to at least a set of key term sample data in order to output a set of language model predictions, wherein the key term sample data demonstrates use of a plurality of key terms; clipping the training audio data into a plurality of clips based on the acoustic model predictions and the language model predictions; and tuning the acoustic model via a machine learning algorithm using the plurality of clips.

Full Text

What is claimed is:

A system and method for audio processing. A method includes training an acoustic model over a plurality of training iterations by, at each of the plurality of training iterations: applying the acoustic model to features extracted from training audio data in order to output a set of acoustic model predictions; applying a language model to at least a set of key term sample data in order to output a set of language model predictions, wherein the key term sample data demonstrates use of a plurality of key terms; clipping the training audio data into a plurality of clips based on the acoustic model predictions and the language model predictions; and tuning the acoustic model via a machine learning algorithm using the plurality of clips.
Timeline
Filed
06/01/2026
Published
09/24/2026
Granted
Not Available
IPC Codes(2)
G10L 15/06:Creation of reference templates; Training of speech recognition systems, e.g. adaptation to the characteristics of the speaker's voice (takes precedence G10L 15/14)
G10L 15/02:Feature extraction for speech recognition; Selection of recognition unit