Video-Based Lip-Reading Training for the Deaf
Author: University of East Anglia
Published: 10 Sep 2009 - Updated: 30 Jun 2026
Publication Details: Peer-Reviewed | Announcement
Table of Contents:
Synopsis - Definition - Overview - Insights, Updates - Related Content
Synopsis
This research, a peer-reviewed pilot study from the University of East Anglia, is the first to benchmark an automated lip-reading system against human lip-readers, and its findings carry weight for the deaf and hard of hearing community because they were prepared for presentation at an international academic conference on auditory-visual speech processing. The work is useful to people with hearing loss, seniors and educators alike because it challenges the long-standing teaching method of spotting static lip-shapes, showing instead that the movement and full appearance of speech gestures matter, and it points toward practical video-based training that helped participants improve after only a few hours - findings published in the Proceedings of the International Conference on Auditory-Visual Speech Processing (AVSP) 2009.*
At a Glance
- 1 - The study compared a machine-based system against 19 human lip-readers.
- 2 - The automated system scored an 80 percent recognition rate, while human viewers reached only 32 percent on the same task.
- 3 - Machines could read speech using only simple features representing the shape of the face, whereas human readers needed full video of people speaking.
Topic Definition
- Lip-Reading
Lip-reading, also called speechreading, is the skill of understanding spoken language by visually interpreting the movements of a speaker's lips, face and tongue, often combined with facial expression and context, rather than relying on sound. It is widely used by people who are deaf or hard of hearing as a way to follow conversation, and while it is a demanding skill that takes considerable practice to develop, research suggests the way it is taught - particularly the use of dynamic, video-based methods over static images - can meaningfully affect how quickly and how well a person learns it.
Overview
A new study by the University of East Anglia (UEA) suggests computers are now better at lip-reading than humans.
The peer-reviewed findings will be presented for the first time at the eighth International Conference on Auditory-Visual Speech Processing (AVSP) 2009, held at the University of East Anglia from September 10-13.
A research team from the School of Computing Sciences at UEA compared the performance of a machine-based lip-reading system with that of 19 human lip-readers. They found that the automated system significantly outperformed the human lip-readers - scoring a recognition rate of 80 percent, compared with only 32 percent for human viewers on the same task.
Furthermore, they found that machines are able to exploit very simplistic features that represent only the shape of the face, whereas human lip-readers require full video of people speaking.
The study also showed that rather than the traditional approach to lip-reading training, in which viewers are taught to spot key lip-shapes from static (often drawn) images, the dynamics and the full appearance of speech gestures are very important.
Using a new video-based training system, viewers with very limited training significantly improved their ability to lip-read monosyllabic words, which in itself is a very difficult task. It is hoped this research might lead to novel methods of lip-reading training for the deaf and hard of hearing.
"This pilot study is the first time an automated lip-reading system has been benchmarked against human lip-readers and the results are perhaps surprising," said the study's lead author Sarah Hilder.
"With just four hours of training it helped them improve their lip-reading skills markedly. We hope this research will represent a real technological advance for the deaf community."
Agnes Hoctor, campaigns manager at the RNID, said:
"This research confirms how difficult the vital skill of lip-reading is to learn and why RNID is campaigning for people who are deaf or hard of hearing to have improved access to classes. We would welcome the development of video-based or online training resources to supplement the teaching of lip-reading. Hearing loss affects 55 percent of people over 60 so, with the aging population, demand to learn lip-reading is only going to increase."
The AVSP conference is being held in the UK for the first time since its inception in 1998. The University of East Anglia will host cutting edge researchers including psychologists, engineers, scientists and linguists from as far afield as Australia, Canada and Japan.
As part of the conference, delegates will take part in a Visual Speech Synthesis Challenge in which a number of visual speech synthesizers, or 'talking heads', will battle it out to determine the most intelligible and visually appealing system.
AVSP runs as a satellite conference to Interspeech 2009 which will be held in Brighton. Topics under discussion will include: machine recognition of audiovisual speech; the role of gestures accompanying speech; modeling, synthesis and recognition of facial gestures; and speech synthesis.
Keynote speakers will be Dr Peter Bull of the University of York who will be exploring The Myth of Body Language and Prof Louis Goldstein of the University of Southern California whose presentation is entitled Articulatory Phonology and Audio-Visual Speech.
Comparison of human and machine-based lip-reading by Sarah Hilder, Richard Harvey and Barry-John Theobald is published in the Proceedings of the International Conference on Auditory-Visual Speech Processing (AVSP) 2009 on Thursday September 10 2009.
The research will be presented on Saturday September 12 at the International Conference on Auditory-Visual Speech Processing (AVSP) 2009 at the University of East Anglia.
How Lip-Reading Errors Happen, Revealed by Network Science: University of Kansas researchers mapped 20,000 English words to show why some are far harder to read on the lips than others.
Insights, Analysis, and Developments
Editorial Note: It is worth remembering that this study dates to 2009, and the field of automated speech recognition has moved considerably since then, yet the central insight remains relevant: how we teach a skill matters as much as the effort a learner puts in. By questioning the value of drilling static lip-shapes and demonstrating that dynamic, video-based practice produces measurable gains in just hours, the UEA team gave teachers and learners a clear reason to rethink traditional methods, and with hearing loss affecting more than half of people over 60, the case for accessible, well-designed training resources has only grown stronger with time.*
Attribution/Source(s): This peer reviewed publication was selected for publishing by the editors of Disabled World (DW) due to its relevance to the disability community. Originally authored by University of East Anglia and published on 10 Sep 2009, this content may have been edited for style, clarity, or brevity.
* Editorial additions by Ian C. Langtree.