The, Um, Psychology of, Like, AI-Generated Voices; Effective Altruists Strike Back at OpenAI
OpenAI’s forthcoming artificial intelligence-powered voice assistant, which allows users to speak with ChatGPT on their phone, has generated tremendous buzz in the industry, and not just because of ScarJo. That’s because developing AI models that can understand human voices and generate realistic AI voices is a lot harder than you might think.
While text chatbot users may be OK with waiting a few seconds for written answers, a few seconds of silence on the phone can feel awkward. Audio AI also faces other unique challenges, such as needing to differentiate between the person speaking and other voices that might be in the background, or understanding when a person is done speaking versus when they’re just pausing mid-sentence. And that’s without getting into quirks like bad audio quality, including echoes that sometimes happen when a caller turns on speaker-phone mode, or the person’s use of lingo such as prescription drug names.
To handle these issues, AI audio startups have spent a lot of time researching a scientific field outside machine learning: human psychology.