Sign-to-text arrives on mainstream phones—and that changes the baseline
Google DeepMind’s SL2T is an AI model that turns live American Sign Language captured by a phone camera into written English text, powering a new ASL to text feature in Pixel 11 that lets people sign instead of type anywhere they would usually enter text.
The most important part of the Pixel 11 launch is not a faster chip or a sharper camera. It is that sign language AI has finally shipped as a standard phone feature. SL2T now sits inside Gboard and Live Transcribe on the Pixel 11 family, marking the first time a sign language translation app has moved from research demo to consumer keyboard and accessibility tools. This is not a novelty add-on; it is a statement about who counts as a default user. For Deaf signers who think in sign, being forced to type in English has always been a tax on speed and self-expression. Pixel 11 quietly removes part of that tax.

Signing instead of typing: why this matters more than another camera spec
On Pixel 11, SL2T powers sign-to-text dictation anywhere a user would normally type: a Deaf signer can dictate a web search, draft a message or document, or query Gemini by signing to the phone’s camera instead of typing. In Live Transcribe, they can sign responses during in-person conversations rather than pecking out replies on a keyboard. Testers reported that using ASL felt faster and more natural than typing in English, which should surprise no one who has ever watched fluent signers converse at full speed.
This is where the new accessibility AI technology earns its hype: it respects language preference instead of treating signing as a problem to work around. Voice dictation normalized talking to phones for hearing users; SL2T does the same for people who communicate visually. The friction it removes is not just physical effort but constant code-switching. When your native language is finally accepted as valid input, your phone starts to feel less like a workaround and more like it was built for you.
The tech leap: from video to text without giving up your privacy
Under the hood, SL2T is ambitious. Google trained the model on more than 100,000 hours of signing data spanning over 50 sign languages, with about a quarter of that data in ASL. According to Google, “Training jointly across languages, dialects, and proficiency levels produced a model that outperforms single-language systems” on internal tests. On the FLEURS-ASL benchmark of complex ASL-to-English translation, SL2T scores 70 BLEURT zero-shot, rising to 74 BLEURT for one-handed signing and 85 BLEURT on assistant-style PRESTO-ASL phrases, with 64 percent exact-match accuracy on a fingerspelling dataset.
Crucially, Pixel 11 does not stream raw video to the cloud. An on-device model tracks 130 key points on the face, body, and hands frame by frame; only these geometric coordinates leave the device, and the original camera feed is discarded immediately as a privacy safeguard. From that landmark stream, SL2T generates English directly, skipping the old gloss layer that chopped sign languages into limited word labels and discarded facial expressions and spatial grammar. This design matters beyond benchmarks: it shows that mainstream accessibility AI can be real-time and privacy-aware instead of “send everything to the server and hope for the best.”
Limits, governance, and the risk of over-selling accessibility AI
For all its promise, SL2T is far from a universal interpreter—and Google admits it. The model can stumble on regional signs, slang, complex ASL grammar, rapid fingerspelling, and meanings carried mostly through facial expressions or head movement. Vendor-reported examples show substitutions for rare signs, dropped classifier constructions, tense errors, and even “ghost text,” where the system produces words when no one is signing or when a second person enters the frame.
Instead of pretending SL2T is magic, Google and its advisory committee have framed version 1.0 as an assistive draft-generation tool for low-stakes, informal settings such as messaging, search, navigation, note-taking, and casual one-on-one conversations. It is explicitly not a replacement for human interpreters in legal, medical, educational, or employment contexts. That restraint is the right call. Accessibility AI technology is most harmful when it is oversold—when imperfect models are pushed into high-stakes situations they cannot handle. The value of SL2T today lies in making everyday phone use less exhausting, not in mediating critical conversations where nuance is non-negotiable.
Where SL2T goes next—and what it signals for all phones
SL2T’s launch is framed as a starting point, not a finished product. Near-term updates already planned include dictation of punctuation, newlines, emojis, and basic editing commands, plus conversation memory across sequential clips in Live Transcribe. Longer-term research aims at better sensitivity to facial expressions and eyebrow position, support for multiple signers, and richer data from regional ASL dialects and Black ASL. The company also plans to keep working with its advisory committee as it adds more sign languages and explores AI that can generate sign language as well.
The bigger story is that a major phone now treats sign language as a first-class input alongside touch, keyboard, and voice. SL2T is the first sign language AI shipped in consumer Android phones, and that sets a bar competitors will struggle to ignore. In a few years, the “Pixel 11 features” list may not be what matters most; the expectation that any flagship should include a serious sign language translation app might. When access stops looking like an app download and starts looking like a default system capability, accessibility moves from accommodation to assumption—and that is where it belongs.







