There are a lot of problems in speech besides transcription. I should have been clearer about what I meant.
Even for people that do want transcription, though, often you can handily beat the accuracy of off-the-shelf services by tuning the models to the domain.
I'm pasting some copy from our website in part because I can't talk too much about specifics, but here's some examples:
"
High quality speech to text transcription, including in the very difficult areas of conversational human-to-human speech: phone calls, voicemails, meetings, etc.
Classification of speech & language: identifying gender, age, regional accents, level of education, etc.
Voice front-ends for various applications and devices, including natural language interfaces: heads-up displays, robots, smartphone apps, etc.
Speech synthesis (a.k.a. TTS) to generate high-quality synthesized speech from text.
Analysis of audio: detecting different types of noise or speech: dog bark, shouting, gunshot, water running, cars, etc.
Customization of speech recognition to processed signals, such as proprietary compression or noise reduction algorithms
Fine-grained analysis of speech and language, for giving feedback in language pathology or language learning.
Analysis of speech & language to detect various health conditions: stroke, dementia, depression, schizophrenia, etc.
Information extraction from text or audio: phone numbers, dates, entities and relationships, etc.
" [0]
Sure, but best possible scenario you end up with a half-baked siri: interesting technology with gimmicky applications. Hardly a good use case to show case how ML can help a startup. The cost/benefit is huge for the general case using speech recognition.
> 'There are a lot of problems in speech besides transcription. '
I'm trying to be vague because I don't want to violate NDAs and it's a pain to think up examples I'm not currently working on, but there are a lot of applications of machine learning to the _speech signal_ (rather than to the problem of speech to text; there are many other things you can predict) that are super interesting. Some examples are listed in the text I copied.
> 'a half-baked siri'
There are so many instances where speech-to-text is super useful for applications other than consumer facing CRUD apps. In a lot of varied industrial applications people are doing very cool things with speech; there are a lot of applications that need very targeted speech systems.
Interesting, sure. Useful or profitable is a separate point that needs a discrete argument. I'd argue it's still in the gimmick phase unless interface IS your product.
Even for people that do want transcription, though, often you can handily beat the accuracy of off-the-shelf services by tuning the models to the domain.
I'm pasting some copy from our website in part because I can't talk too much about specifics, but here's some examples:
"
High quality speech to text transcription, including in the very difficult areas of conversational human-to-human speech: phone calls, voicemails, meetings, etc. Classification of speech & language: identifying gender, age, regional accents, level of education, etc. Voice front-ends for various applications and devices, including natural language interfaces: heads-up displays, robots, smartphone apps, etc. Speech synthesis (a.k.a. TTS) to generate high-quality synthesized speech from text. Analysis of audio: detecting different types of noise or speech: dog bark, shouting, gunshot, water running, cars, etc. Customization of speech recognition to processed signals, such as proprietary compression or noise reduction algorithms Fine-grained analysis of speech and language, for giving feedback in language pathology or language learning. Analysis of speech & language to detect various health conditions: stroke, dementia, depression, schizophrenia, etc. Information extraction from text or audio: phone numbers, dates, entities and relationships, etc. " [0]
[0] http://cobaltspeech.com/our-products.html