A Nigerian software Engineer at Google, Ijemma Onwuzulike has announced the creation of the first Igbo speech-to-text model.
Known as IgboSpeech, this Automatic Speech Recognition (ASR) model is specifically designed for Igbo language, showcasing a Word Error Rate (WER) of 29.
Onwuzulike shared a preview of IgboSpeech on social media, highlighting its potential to revolutionize the way Igbo language is processed and understood by technology.
“The first Igbo speech-to-text model is here!” she exclaimed in her announcement. She emphasized the team’s commitment to further refining the model and plans to release it alongside the Igbo API.
This development marks a significant step forward in the preservation and promotion of the Igbo language, offering new opportunities for integration into modern technology and digital platforms.
“The first Igbo speech-to-text model is here!
“Please welcome IgboSpeech an ASR model tailored to Igbo speech and text with a 29 WER
“We plan on improving the model and releasing it with the Igbo API https://t.co/3TXqklSRt9
https://x.com/i/status/1807804737380552704
What is text to Speech technology
Text-to-speech (TTS) is a type of assistive technology that reads digital text aloud. It’s sometimes called “read aloud” technology. With a click of a button or the touch of a finger, TTS can take words on a computer or other digital device and convert them into audio.
How text-to-speech works
The voice in TTS is computer-generated, and reading speed can usually be sped up or slowed down. Voice quality varies, but some voices sound human. There are even computer-generated voices that sound like children speaking.
Some TTS tools also have a technology called optical character recognition (OCR). OCR allows TTS tools to read text aloud from images.
