How do computers talk with voices that sound like people?
Computers break examples of human speech into small bits of sound, then put them together to sound like people do.

Curious Kids is a series for children of all ages. If you have a question you’d like an expert to answer, send it to CuriousKidsUS@theconversation.com.
How is a text-to-speech voice made? – Sarah G٫ age 11٫ Seguin٫ Texas
When you talk to computerized assistants like Siri or Alexa, they reply in voices that sound very human. But how do computers, smartphones and apps actually talk like a person? They use a technology called text-to-speech.
When you speak, your lungs push air up your windpipe and through the vocal cords in your throat. That makes the vocal cords vibrate, which creates sound. Your brain tells your mouth, tongue and lips to shape that sound into words.
I’m a computer engineer who researches how computers create realistic experiences for people. A computer simulates this process of making spoken words. It sends electrical signals to a tiny speaker, which vibrates really fast. The vibrations push against the air surrounding the speaker, creating sound waves. Software in the computer controls those electrical signals in order to shape the sound waves to create speech.
If a computer wants to say “Hello, how are you?” it breaks each word into bits of sound called phonemes. Phonemes are the smallest building blocks of speech, such as the sounds “sh,” “short i” and “p” to say “ship.” The software creates the phonemes and groups them in the correct order to form words: “Heh” “lo” “how” “r” “u”? The human brain and mouth create and connect phonemes, too.
Robot speech
Way back in the 1700s, inventors tried to make machines work like your lungs and throat do. They used bellows – a big bag that a person could squeeze – to push air from inside the bag through pipes, whistles and leather tubes. The sounds that came out were squeaky, weird and creepy.
The first electronic speech machines, called synthesizers, were built in the 1930s. One famous machine was called Voder, which made its debut at the 1939 World’s Fair in New York City. Voder looked like an organ. A person had to press electronic buttons, keys and foot pedals to get it to gasp out basic phrases like “Good morning!”
Computers began speaking by putting together phonemes in the 1960s, creating stiff, robotic speech.
Pieces to a puzzle
Old computer voices often sounded robotic and choppy, like “He-llo-hu-man-I-am-a-com-pu-ter.” This happened because older software programs had to stitch together small sounds that had been mapped out from recorded voices. The maps, called spectrograms, look like graphs with peaks and valleys representing how strong each tone was at each instance when a sound happened.
The programs put the maps together like pieces in a puzzle and turned them back into sounds. That method worked, but it sounded very unnatural.
Today’s computers use machine learning, a kind of artificial intelligence, to sound like a person. Engineers and scientists train an AI program by giving it many hours of recordings of real people talking. The machine learning software analyzes patterns in the speech. That includes everything from how people breathe to when they laugh and to how their voices sound higher when they get excited.
The patterns allow the computer to shape phonemes into words and sentences in all the subtle ways people do.
This advanced technology even allows a sophisticated AI computer to listen to a recording of your voice for just a few seconds, learn your exact speech patterns and copy it. It can then say sentences you have never actually spoken, in a voice that sounds just like yours.
Helpful and harmful voices
Text-to-speech software helps millions of people every day. In cars, it gives directions so drivers can keep their eyes on the road. It can read news articles and websites for people who are blind, and it can give a voice to people who cannot speak.
Like some other powerful technologies, the software can also be misused. Advanced software can create highly realistic fake voices, sometimes called audio deepfakes. These synthetic voices can sound almost identical to a real person’s voice by learning from a short sample of their speech.
Scammers can use this technology to impersonate family members, co-workers or celebrities. They can make phone calls or leave voice messages that try to fool people into believing bad information. Scientists and engineers are working to make tools that can identify fake voices to help stop scammers.
So the next time you hear a phone, computer or video game speak with a human voice, you know how it was able to it without having lungs, vocal cords, lips or a tongue!
Hello, curious kids! Do you have a question you’d like an expert to answer? Ask an adult to send your question to CuriousKidsUS@theconversation.com. Please tell us your name, age and the city where you live.
And since curiosity has no age limit – adults, let us know what you’re wondering, too. We won’t be able to answer every question, but we will do our best.
Tam Nguyen receives funding from National Science Foundation.
Read These Next
A Colorado jail hasn’t allowed in-person visits in nearly 20 years – here’s what research says about
Nearly all county jails across Colorado’s Front Range allow prisoners to visit with their loved ones…
First mRNA flu shot approved by FDA bodes well for improving drugs of the future – though a few hurd
From COVID-19 vaccines to Moderna’s mRNA flu vaccine, using mRNA as medicine has shown promise. But…
Why Gen Z turns everything – even murder – into a joke
A psychologist explains why dark political humor doesn’t prove young people are nihilistic, but rather…


