Describe to me how an audio model works. In early days, you tried to replicate it exactly like you would replicate it with the human body. So you would try to completely reproduce a machine that would create a vocal tract effectively. Then that progressed into trying to create a digital signal for speech. Bell Labs was one of the first to try to create a structured set of signals that would represent the speech. And that is a first precursor to what we would do today. Then you would try to stitch in phonemes, effectively different sounds of how we would speak humans, and then try to concatenate them together.
Based on the most probabilistic approach of the next word, you would effectively try to bring the phonemes from your library of phonemes and bring them together. And then down to the modern history where now we predict tokens, but instead of predicting them on the text space, you predict them on a phoneme space. you would effectively try to bring the phonemes from your library of phonemes and bring them together. And then down to the modern history where now we predict tokens, but instead of predicting them on the text space, you predict them on a phoneme space.