One hand, twelve tasks, one scoreboard

The Berkeley effort behind T-Rex collected 100 hours of high-quality tactile data with one instrumented hand across more than 200 household objects, then reports 65 per cent success across 12 tasks, close to double the best vision-language-action model. The tasks are concrete enough to picture: screwing in a light bulb, applying toothpaste, transferring an egg. Trevor Darrell's framing is that most dexterous manipulation can be done by humans with their eyes closed, which is a fair argument for collecting touch at all. The data comes from a single robotic hardware instance, and that is also where the claim stops.[1]

Doubling the best vision-language-action model is a comparison run inside one laboratory's rig. The baseline sits on the same hand, works through the same 12 tasks and is scored by the same group. That makes the gain informative about what touch adds in that setup and quiet about what it adds in another. One alternative reading deserves saying out loud: the 12 tasks may have been chosen where contact matters most, which widens the margin without any error in the measurement.[1]

The group that changed the body

The other three efforts scale differently. A Tsinghua University group led by Chengbo Yuan aggregated more than 3.000 hours from public datasets covering 21 sensor types and several robot bodies, mapping every sensor's output onto a shared human hand template, and the resulting model worked on platforms it had not seen; the group now leads an 80 institution collaboration for a larger dataset. A Fudan University effort with Shunlin Lu assembled more than 30.000 hours with vision and touch recorded in step, aiming at about 100.000. A University of Southern California model infers touch from camera images alone and reports 62.8 per cent on contact-rich tasks against 28.2 per cent without touch.[1]

Read as a set, the four measurements answer four different questions. Tsinghua is the only group that varies the robot body, which is the variation a reader most wants tested, and its result arrives as generalisation rather than as a success rate, so the number that would sit beside 65 per cent does not exist. Each effort is reported with its own hardware, its own dataset and its own evaluation, and no common task suite or independent evaluator appears anywhere in the account. What the four have in common is a unit of competition: hours of data.[1]

What the next measurement would have to do

Long Cheng of the Chinese Academy of Sciences names a mechanism that cuts against hours as the unit: vision supplies continuous, high-bandwidth data while tactile signals stay sparse and intermittent, so models learn to deprioritise touch. If that is the binding constraint, adding hours moves the wrong lever, because the architecture decides how much weight a sparse channel receives. The alternative is simpler and just as live: tactile corpora remain tiny beside vision corpora, and 100 hours against a web-scale image set is exactly the gap the Fudan and Tsinghua efforts are closing.[1]

The measurement that would make these numbers speak to each other is unglamorous: one task list, one scoring rule, at least two robot bodies, and a scorer who did not build the dataset. That is a design request, not a complaint about the work — each group has published enough detail for someone else to try it. If one of the four efforts publishes such a cross-hardware evaluation before 31 March 2027, with the task list, the scoring rule and per-body results in the open, the success rates become comparable with each other for the first time; the observable signal is a published task list carrying a separate result for each hand. Until then, 65 per cent describes one hand.[1]