What the score covers
SL2T, introduced by Google DeepMind, was trained on more than 100,000 hours of data across more than 50 sign languages, and roughly a quarter of that data is American Sign Language. The feature that reached the phone works in one direction only, from American Sign Language to English. The single numeric result the company reports is a zero-shot score of 70 BLEURT on the FLEURS-ASL sd-test benchmark, and that benchmark tests the same language pair.[1]
The issue here is scope. A zero-shot 70 BLEURT says how close a translation output comes to its reference text; it says nothing about the other 50 languages in the training set, and nothing about the filming conditions met on a phone. The benchmark runs over a cleanly recorded evaluation set, while the product runs through the camera of a handheld phone. No published figure closes that gap.[1]
Google's own list
What is interesting is that Google points to this gap first. The post says that optimizing academic benchmarks alone "doesn't guarantee usability in real-world applications", and then lists what was worked on separately: minimizing streaming latency, preventing hallucination on non-signing inputs, ensuring fairness for the 10 percent of signers who are left-handed, and improving performance for one-handed signing used while the other hand holds the phone. None of those four headings carries a measurement.[1]
The quality of the shipped feature therefore rests on the one measurement that does not cover it. The same post accepts that errors persist on rare signs, rapid fingerspelling, passive constructions and tense without context, but describes them at the level of examples rather than rates. Another reading is available: Deaf user studies were run and a joint impact report was written with AISLAC, and the numbers may sit there. Until those values appear alongside the announcement, though, what the feature shipped on the Pixel 11 does condition by condition stays unknown from outside.[1], [2]
What would make this testable?
Writing on August 1 about the 70.1 average LG AI Research reported across 24 benchmarks, I argued that the distribution under a single figure had become measurable from outside because the weights were released under an open licence. The situation here is the closed form of the same problem. SL2T also comes with a single figure, but neither the weights nor results separated by condition have been published, so for now there is no way to measure the distribution from outside.[1], [3]
What would change that is specific and narrow: a document giving error rates separately for the four headings Google itself listed. Error rates for left-handed and one-handed signers, the rate of invented output on non-signing video, streaming latency, and error rates on rare signs and fingerspelling. If a technical document or impact report carrying those separated results appears by the end of December 2026, the quality of the feature on the phone becomes a testable claim for the first time; if it does not, 70 BLEURT will keep circulating on its own as evidence of performance for a product it did not measure.[1]