AI's ability to process spoken languages has advanced rapidly over recent decades, making automatic translation, dictation and conversational interfaces feel effortless to hearing users. Yet that revolution has not reached the world's more than 200 sign languages — or the estimated 70 million Deaf and hard of hearing people who use them.
Google DeepMind has introduced SL2T, a massively multilingual sign-language-to-text translation model aimed at changing that. The company describes it as a breakthrough in quality and generality, and the headline is this: sign language AI is moving out of the lab and into consumer products for the first time.
Where it runs
SL2T powers two features on the Pixel 11: sign-to-text dictation in Gboard, and Live Transcribe. It starts with American Sign Language (ASL) to English, with more devices coming soon and additional languages to follow.
The experience mirrors dictation for hearing users. A Deaf user can sign to their phone anywhere they would normally type:
- Searching the web
- Drafting messages or documents
- Asking Gemini to answer queries or execute tasks
- Signing responses in Live Transcribe instead of typing back and forth
According to the company's testers, signing in ASL is faster, more natural and more delightful than typing in English.
Why this was hard
Sign languages are the primary languages of Deaf communities around the world and a cornerstone of Deaf cultural identity. There is great diversity among Deaf people in proficiency at signing, speaking, reading and writing, which is why supporting access across all modalities matters.
Despite the opportunity, progress in the field has been slow. DeepMind points to two core challenges.
The first is a translation problem. Transcribing speech is a sequential mapping from sound to text within the same language. Sign languages are independent, natural languages with their own distinct grammars and lexicons, so they require true machine translation rather than a sequential process of sign-to-word transformations.
The second is a vision problem. The model must learn to "see" and understand physical movement. Sign languages convey meaning through simultaneous movements of the hands, arms, torso, head and face. Tracking all of that at once, accurately, is structurally different from following a single stream of audio.
The company also notes that widespread misconceptions about how these languages actually work have contributed to the slow pace of progress.
Why it matters
The weight of this announcement lies less in the technical achievement than in the distribution. Academic work on sign language translation has run for years; what has been missing is that work running on a device in millions of pockets. The distance between a model scoring well in the lab and a user reading a label in a shop has always been the hardest part of this field.
For hearing users, dictation is an alternative to typing what you would say. For a Deaf user, sign-to-text dictation is a way to communicate in their own first language — not having to write in a second one. The testers' verdict that it is "faster and more natural" describes exactly that.
The limits are equally clear. One language pair (ASL to English) and one device family are supported today. No timetable has been given for the remaining 200-plus sign languages.