Robotics Risk Engine

Robotics Research — the science of actuation, sensing & embodied AI

An independent crew studying the COMPONENT SCIENCE upstream of the humanoid / physical-AI deployment wedge — actuation (electric quasi-direct-drive, artificial muscles), sensing (vision-based tactile), manipulation (dexterous in-hand), embodied AI (vision-language-action foundation models), locomotion (sim-to-real perceptive control), and power/endurance. It grades each domain by REPRODUCED capability — not demo theater — and tracks the R&D that would drive the next inflection and reprice the picks-and-shovels bottlenecks. 3-tier write-ups; failed ideas kept; findings bridge into the wealth AI (Parallax), the telescope dream, and the robotics wedge's BOM-layer reads.

*In plain English: Robots can now walk over wild terrain reliably — a genuinely solved, reproduced skill — and are learning to follow spoken commands. But dexterous hands and all-day battery endurance are NOT solved, and full-body humanoid demos often borrow the credibility of solved locomotion to imply unsolved manipulation. This crew keeps that line bright and honest.
For the kids: Robots got really good at walking and even hiking, and they're learning to understand what you say. But they're still clumsy with their hands and run out of battery fast — and we say so instead of pretending.
viable = reproduced / crediblemodeled = single-lab / plausibleopen = early / contested— capability is graded by reproduced results, not demo theater.

Domains — actuation, sensing & embodied AI

reproduced / credibleElectric quasi-direct-drive actuators (commodity modules)actuation

HELD viable — the cleanest independence check in the lens. At least three genuinely independent hardware lineages: MIT (Wensing et al., T-RO 2017; Katz et al., ICRA 2019), ETH (ANYdrive/ANYmal), and MPI-IS/NYU/LAAS (Grimminger et al., RA-L 2020, an open module rebuilt at multiple institutions outside MIT's lineage), plus purchasable commodity modules and dozens of adopting labs. Remove any one lineage and the grade survives.

★ For the kids: Robot leg muscles you can buy in a box. Lots of different teams, in different countries, built their own versions and they all work.

Technical: Wensing, Wang, Seok, Otten, Lang & Kim, IEEE T-RO 33(3):509-522 (2017); Katz, Di Carlo & Kim, ICRA (2019); Kau et al. (Stanford Doggo), ICRA (2019); Grimminger et al., IEEE RA-L 5(2):3650-3657 (2020); Hutter et al., ANYdrive/ANYmal (ETH, 2016+); CubeMars/T-Motor AK-class adoption 2020-2026.

Frozen prediction: By end of 2027, a teardown of a sub-$30k commercial humanoid or quadruped shows commodity QDD modules or Mini-Cheetah/ODRI-derivative actuators in the majority of its joints.

single-lab / plausibleHASEL electrohydraulic artificial musclesactuation & power

Held at modeled: one peer-reviewed robotic leg (Buchner 2024), a commercial supplier (Artimus Robotics), and EPFL (Shea lab) electrohydraulic variants. Walking/hopping integration is unreproduced outside the Keplinger/Katzschmann lineage; all demonstrations use off-board kV supplies; no independent life-test publication.

★ For the kids: Oil-filled pouches that squeeze when electrified. They still need a big power box on the floor.

Technical: Acome et al., Science 2018; Kellaris et al., Sci. Robot. 2018; Rothemund et al., Adv. Mater. 2021; Buchner et al., Nat. Commun. 2024. Open: onboard HV conversion mass, dielectric life >10^6 cycles, self-clearing failures.

reproduced / credibleVision-based tactile sensing (GelSight)sensing

HELD viable on the LAB axis only, with the commercial leg recorded as failing. Lab reproduction is strong: TacVerse (2026) benchmarks seven vision-based tactile sensors across three sensing principles and four independent design families on 106,800 labelled images. The products are choosing otherwise — the first mainstream industrial-gripper tactile fingertip is capacitive, and the loudest humanoid fingertip programs advertise force-type sensing. The standing commercial prediction stays frozen, unreworded, and marked AT RISK / trending false; the viable grade is forbidden from laundering the …

★ For the kids: Lots of labs can build the tiny-camera touch pad and it works — but the ones going on sale use a different, simpler kind of touch.

Technical: Yuan, Dong & Adelson, Sensors 2017; Ward-Cherrier et al., Soft Robotics 2018 (TacTip); Taylor et al., ICRA 2022 (GelSlim 3.0); Lambeta et al., RA-L 2020 (DIGIT); Wei et al., TacVerse, arXiv:2606.25877 (2026). Counter-evidence on the commercial leg: Robotiq TSF-85 (capacitive), Jan 2026.

Frozen prediction: AT RISK (trending false), left frozen and ungraded: 'By end 2027, at least one shipping commercial robot hand integrates camera-based tactile fingertips as standard.'

single-lab / plausibleDexterous in-hand reorientation (the consecutive-success spread between labs)manipulation & dexterity

Held modeled. 12x spread in consecutive reorientations across labs (ViserDex > 25 vs unified-hand-action-space 2.0) is the signature of single-lab capability.

★ For the kids: One hand flips a block 25 times in a row, another two; until several get the high number it is 'promising'.

Technical: arXiv 2604.11138; arXiv 2607.03570; Lin et al. arXiv 2502.20396; Chen et al. Science Robotics 2023.

Frozen prediction: By 2027-06-30 at least two independent labs report >= 10 mean consecutive reorientations of held-out objects from RGB-only on an open LEAP-class hand. p=0.45.

single-lab / plausibleVision-language-action foundation modelsembodied AI

Held modeled. Reproduced in simulation by vla-eval (six codebases); every real-robot number remains developer-own.

★ For the kids: Big robot brains improve, but only their makers have measured them on real robots.

Technical: vla-eval arXiv 2603.13966; GR00T N1.6; RynnBrain 1.1; pi*0.6 arXiv 2511.14759.

Frozen prediction: The first independent >=50-trial evaluation of open-weight pi0.5 in an unseen home, if published by 2027-08-26, reports success below 60% of the paper figure (P=0.6).

reproduced / credibleSim-to-real perceptive legged locomotion (the reproduced recipe)locomotion

Viable stands. No single ingredient; a reproduced three-part recipe: privileged teacher -> proprioceptive student (Lee et al., Sci. Robot. 2020, ETH), massively parallel GPU simulation (Rudin et al., CoRL 2021), online adaptation from proprioceptive history (Kumar et al., RMA, RSS 2021). Reproduced on ANYmal, Unitree A1/Go1, Mini Cheetah and in shipped commercial controllers. Perceptive extension reproduced by Miki et al. 2022 (ETH), Agarwal et al. 2022 (CMU), Cheng et al. 2023 (CMU parkour).

★ For the kids: Robots learned to walk by practising thousands of copies at once inside a computer, then carrying that practice into the real world.

Technical: Elevation-map encoding with belief-state estimation (Miki 2022). The 'single ingredient' framing in the engine's own pantry question (origin=ENGINE) is not supported by the record.

Frozen prediction: Through 2027 every new sim-to-real legged controller reaching a commercial platform uses all three ingredients; a shipped controller omitting parallel sim or teacher-student/adaptation falsifies this.

early / contestedPower and energy density (the runtime ceiling)actuation & power

The reproduced facts are narrow: Joule heating dominates legged electric cost at low speed (MIT Cheetah) and no artificial-muscle class simultaneously matches mammalian muscle on specific power, strain, efficiency and cycle life (two independent review tabulations agree). Everything above that is vendor specification or cell datasheets with no robot-integrated pack documented.

★ For the kids: Robots carry batteries and motors. Both are heavy for how much work they can do, and that is the main reason robots get tired fast.

Technical: Seok et al., TMECH 2015; Kashiri et al., Frontiers in Robotics and AI 2018; Mirvakili & Hunter, Adv. Mater. 2018; Zhang et al., IEEE Trans. Robotics 2019 (venue corrected from 'Chemical Reviews'). Si-anode 400-500 Wh/kg cells (Amprius datasheets) are not literature evidence for a robot pack.

reproduced / crediblePrecision strain-wave & cycloidal reducers (the gearing bottleneck)actuation

HELD viable on CAPABILITY, with the uncited price claim excised. Decades-old, multi-vendor, reproduced precision gearing; the grade is earned by the mechanism and its industrial track record. The trade-press economic assertion ('Chinese reducers at 40-60% of incumbent pricing') now lives at open in precision-reducer-price-collapse with an explicit teardown falsifier. Capability viable; economics unproven; the two no longer share a grade.

★ For the kids: The special gearboxes that make robot joints strong are genuinely old, reliable technology. Whether they suddenly got cheap is a different question, now in the 'not proven' pile.

Technical: Musser, US Patent 2,906,143 (1959); Ghorbel, Gandhi & Alpeter, ASME J. Mechanical Design (2001). Price evidence: none of publishable quality.

Frozen prediction: Through end of 2028, strain-wave or cycloidal reducers remain the dominant rotary joint reduction in the majority of joints across published humanoid teardowns.

early / contestedPlanetary roller screws (humanoid linear actuation)actuation & power

Demoted from modeled. Humanoid use (Tesla Optimus, 2024-2025 entrants) rests on marketing video and supplier press; no peer-reviewed characterization (efficiency, backdrivability, life) of a roller-screw humanoid joint exists. Peer-reviewed roller-screw evidence is aerospace/exoskeleton EMA literature at a different duty cycle.

★ For the kids: A roller screw turns spinning into pushing, like a very strong bolt. Some robot companies say they use it in their robot legs, but nobody has shown the test results.

Technical: Ma et al., Mechanism and Machine Theory 2015 (load distribution); aerospace flight-control EMA reports. Scripted vendor demos are not capability evidence.

Frozen prediction: By end of 2028, planetary roller screws (not ball screws or rotary-only drives) will be the dominant linear-actuator mechanism disclosed in at least two of the three highest-volume commercial humanoid programs.

single-lab / plausibleTwisted-and-coiled polymer (TCP) artificial muscles (thermally driven only)actuation & power

Held modeled; scope note only (electrochemical branch split out). Reproduced material FoM across many labs; ~1% thermal-cycle efficiency and slow cooling keep it off robots.

★ For the kids: Fishing line that shrinks when heated; heating and cooling it fast wastes almost all the energy.

Technical: Haines et al., Science 343:868 (2014); Mu et al., Science 365:150 (2019).

Frozen prediction: By 2028-12-31 no commercial robot joint ships with a thermally driven TCP muscle as primary actuator. p=0.90.

single-lab / plausibleDielectric elastomer actuators (DEA)actuation

CONFIRMED modeled. Reproduced across many labs for 25 years, but multi-kV drive electronics, dielectric breakdown, and load-cycle reliability keep DEAs research-grade — no humanoid ships DEA joints. Distinct mechanism from the existing HASEL electrohydraulic-muscle domain.

★ For the kids: A soft rubber sheet that squishes and stretches when you put electricity across it — fast like a muscle, but it needs dangerously high voltage.

Technical: Pelrine, Kornbluh, Pei & Joseph (SRI), Science 2000, 287(5454):836-839. Reproduced widely; open reliability limits are electrical breakdown, high-voltage driver mass/cost, and load-cycle life.

Frozen prediction: Through 2028, no commercially deployed bipedal humanoid will use dielectric-elastomer artificial muscles as a primary joint drive.

single-lab / plausibleActuator thermal limits (the continuous-torque ceiling)actuation

HELD modeled, cross-linked to its structural blocker. The thermal ceiling is real physics and the published fix is peer-reviewed but single-lineage. The domain cannot move in EITHER direction until actuator-measurement-standards resolves: with no shared continuous-torque protocol under stated ambient and duty cycle, a vendor cannot be shown to have beaten the ceiling and cannot be shown to have hidden it. Publishing peak numbers is the equilibrium a missing referee produces.

★ For the kids: Robot muscles get hot and have to slow down. Someone found a way to cool them, but because nobody measures heat the same way, we cannot check who actually fixed it.

Technical: Paine & Sentis, Actuators 4(3):182-202 (2015); Chignoli et al., 'The MIT Humanoid Robot' (2021); no cross-lab continuous-torque-density benchmark exists (Actuators 15(6):301, 2026 proposes one, unadopted).

Frozen prediction: By end of 2027, at least one shipping or publicly torn-down commercial humanoid uses liquid-cooled joint actuators; and through 2027 no vendor publishes independently verified CONTINUOUS joint torque ratings for a full humanoid under stated ambient and duty cycle.

reproduced / credibleMagnetic / Hall-effect taxel tactile skinssensing

PROMOTED modeled -> viable with CORRECTED rationale: the proposal claimed three independent lineages, but ReSkin (CMU/Meta) and AnySkin (NYU) share a lead author — that is one lineage. What actually clears the bar: two truly independent lineages (ReSkin/AnySkin family + Waseda uSkin, commercialized by XELA and deployed on iCub), commercial availability, external user groups, and pretrained encoders (Sparsh-skin). Meets the >=2-lineages-plus-external-use bar with no margin; any retraction re-opens the grade.

★ For the kids: A stretchy sticker full of tiny compasses feels pushes and slides — different labs build and swap these skins like phone cases.

Technical: Bhirangi et al. CoRL 2021 / ICRA 2025; Tomo et al. RA-L 2018 (uSkin/XELA); Sparsh-skin 2025.

Frozen prediction: By end 2027, magnetic-taxel skins ship in >=2 commercially sold hands/grippers from different vendors.

early / contestedLarge-area stretchable electronic skinsensing

HELD open, with trade-show claims explicitly refused. The strongest capability result is real-time whole-body tactile control on a humanoid — but that is the TUM HEX-o-SKIN lineage in roughly its fifteenth year, unreproduced by any independent lab, and the field's own ICRA 2026 workshop states that current solutions are 'often limited to small-area arrays or fingertip sensors'. CES 2026 produced a full-body-skinned humanoid unveiling and a hundred-dollar whole-humanoid e-skin cost claim with no third-party measurement, no task result and no published spec; under this lens those are NOT eviden…

★ For the kids: One lab has covered a robot in skin and it works, but nobody else has copied it — and the shiny trade-show skins come with no proof at all.

Technical: Armleder et al. (TUM ICS), Advanced Intelligent Systems 2025 (HEX-o-SKIN lineage, Mittendorfer & Cheng, T-RO 2011); 'Whole-Body Compliant Control ... Flexible Tactile Electronic Skin', IJHR 2022; ICRA 2026 workshop 'Towards Large-Area Tactile Sensing Skins'; Bao lab, Science 2023; NUS ACES 2019. Refused as evidence: CES 2026 vendor unveilings.

Frozen prediction: Through 2028, whole-body tactile skin remains absent from every commercially deployed humanoid pilot and collision safety stays joint-torque- or safety-skin-based; and no CES-class low-cost whole-body skin claim is substantiated by a third-party-measured spec sheet within 24 months of announcement.

single-lab / plausibleEvent-based / neuromorphic vision sensingsensing

HELD modeled. Sensor commercially reproduced (Prophesee/Sony, iniVation, Samsung); strongest closed-loop robot payoff (UZH ~3.5ms quadrotor dodge) is originating-lab. No third-party fair benchmark where event vision beats frame vision on a manipulation/locomotion task; no shipping control loop standardizes on it.

single-lab / plausibleCross-sensor tactile representation learning (automated architecture search matched, did not beat, the expert)manipulation & dexterity / touch

Held modeled. TacEvo (arXiv 2606.30109, verified) matches the expert baseline on ViTacTip force regression; 96.0% is generation trainability, not sensing accuracy. Single sensor, single lab.

★ For the kids: A computer tried thousands of fingertip-brain designs; its best was about as good as the human one.

Technical: AbuSadeh, Wei & Zhang, Imperial 2026; ViTacTip arXiv 2402.00199.

Frozen prediction: By 2027-06-30 no automated tactile architecture search reports > 10% force-regression improvement over an expert baseline on a second sensor family. p=0.70.

reproduced / credibleLearned grasping of novel objects (parallel-jaw / suction)manipulation

CONFIRMED viable — the one manipulation capability both reproduced AND deployed at scale. Predicting robust parallel-jaw/suction grasps of unseen objects in clutter from depth is reproduced (Dex-Net GQ-CNN, Contact-GraspNet) and underpins live commercial bin-picking/parcel induction (Amazon, Covariant, Ambi). Deliberately sidesteps in-hand dexterity (object is regrasped, not manipulated within fingers) — which is why grasping is solved while dexterity is not.

★ For the kids: Robots are genuinely good now at picking up things they have never seen and dropping them in a bin — that part is done and used in real warehouses.

Technical: Mahler et al. Dex-Net 2.0, RSS 2017 (arXiv:1703.09312), 93% adversarial; Sundermeyer et al. Contact-GraspNet, ICRA 2021 (arXiv:2103.14127); deployed in commercial induction/bin-picking.

Frozen prediction: Through 2028, at-scale commercial robotic manipulation stays dominated by learned parallel-jaw/suction grasping; multi-fingered dexterous hands appear in <10% of deployed manipulation stations.

reproduced / credibleVisuomotor imitation learning (diffusion policy / action chunking)embodied-ai

HOLDS viable, with an EMBEDDED CLAIM EXCISED. The viable grade is earned by the technique itself — diffusion policy and action chunking are reproduced across dozens of independent labs on independent hardware, and TRI's Large Behavior Model study is the rare blind, randomized, powered evaluation (N=50 real / 200 sim per condition, paired Barnard/Welch, Bonferroni correction, Bayesian credible regions). What does NOT inherit that grade is the SCALING claim riding inside it — 'improves predictably with pretraining scale' is a single-organization result on a single organization's data and evalua…

★ For the kids: Lots of teams built this kind of robot brain and it works — that part is real. Whether it keeps getting better the bigger you make it is one company's finding, not everyone's.

Technical: Chi et al., Diffusion Policy, RSS 2023 / IJRR 2024; Zhao et al., ALOHA/ACT, RSS 2023; TRI, 'A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation', Science Robotics 2026 (doi 10.1126/scirobotics.aea6201; arXiv:2507.05331), lbm_eval suite (49 tasks) open-sourced. Scaling claim: TRI-only, no external replication.

Frozen prediction: By 2028-06-30, at least one group unaffiliated with TRI publishes a blind, randomized, N>=50-per-condition real-robot comparison of a generalist policy against a single-task baseline; absent that, the multitask-pretraining scaling claim remains single-organization and ungraded.

reproduced / credibleOpen anthropomorphic dexterous hand hardware (LEAP-class)actuation

KEPT viable (existing domain; grade confirmed at viable, scoped strictly to HARDWARE): LEAP Hand independently built and used as the standard platform across many labs; Allegro and Shadow commercially available for a decade. The scope discipline is what saves this grade — hardware access is reproduced; control software is not, and inherits nothing from this promotion.

★ For the kids: Any robot lab can now buy or 3D-print a good robot hand cheaply — the hard part left is teaching it.

Technical: Shaw, Agarwal & Pathak, LEAP Hand, RSS 2023 (open hardware, independently rebuilt); Wonik Allegro (2013-); EyeSight Hand 2024; F-TAC Hand, Nature Machine Intelligence 2025 (single-lab, excluded from the viable scope).

Frozen prediction: By end of 2027, at least three independent labs will publish multi-axis in-hand reorientation results on sub-$5k open-hardware hands, confirming the bottleneck has fully shifted from hand hardware to policy learning.

early / contestedContact-rich manipulation & assembly sim-to-realmanipulation

HELD open (existing grade confirmed): NVIDIA IndustReal/Factory line is the strongest evidence but remains essentially single-lab-centered; independent reproductions on shared NIST task boards are sparse. Clear promotion trigger defined.

★ For the kids: Simulator-taught snap-together assembly is starting to work on real parts — mostly in one lab.

Frozen prediction: Promotes to modeled when >=2 non-NVIDIA labs reproduce sim-to-real insertion >=90% on a shared NIST-class board; by end 2027 this happens for insertion but not multi-step assembly.

reproduced / credibleGPU-parallel simulation & sim-to-real infrastructureembodied-ai

HOLDS viable, scope stated precisely: the grade rests on massively parallel LOCOMOTION sim-to-real, reproduced across dozens of independent labs since the Isaac Gym lineage, not on the 2025 engine consolidation. Newton (NVIDIA + Google DeepMind + Disney Research, Linux Foundation) and MuJoCo-Warp are substrate news, not reproduced results, and carry no grade of their own yet. The unsolved part is not throughput but contact-model fidelity: manipulation, not walking, is where the reality gap still bites, which is why contact-rich-manipulation-sim2real stays open alongside this viable grade.

★ For the kids: Robots practise millions of times inside one graphics card before they try the real thing — that part really works for walking, not yet for fiddly hand jobs.

Technical: Rudin, Hoeller, Reist & Hutter (ETH Zurich), 'Learning to Walk in Minutes Using Massively Parallel Deep RL', CoRL 2021; Newton differentiable physics engine (2025, contributed to the Linux Foundation) and MuJoCo-Warp GPU backend — new substrate, ungraded; contact-fidelity gap documented in physics-engine reality-gap benchmarks (MuJoCo vs Bullet vs Flex vs SOFA).

Frozen prediction: Through 2027, no published contact-rich assembly policy transfers zero-shot from a GPU physics engine to hardware at >80% success across >=10 novel part pairs without real-world fine-tuning.

single-lab / plausibleLearned world models for robot policy learninglearning

HELD modeled (existing grade confirmed): DreamerV3 (Nature 2025) peer-reviewed and widely reproduced in sim; V-JEPA 2-AC real zero-shot pick-and-place is single-lab; sim results alone don't earn viable under this study's rules.

★ For the kids: The robot imagines what happens next before acting — good in games, still being tested in the real world.

early / contestedHuman-video pretraining for manipulationlearning

CONFLICT RESOLVED to open (existing grade confirmed; proposals carried both open and modeled): three groups report downstream gains in kind (LAPA ICLR 2025 — peer-reviewed; GR00T N1 and V-JEPA 2 — company technical reports), but no controlled ablation anywhere isolates video pretraining against a same-compute robot-data baseline, even within a single lab, and two of the three legs are vendor-published. Gains-in-kind without controlled comparison cannot ground modeled for a claim whose entire content is 'video helps relative to baseline.' The frozen third-party-ablation gate is the promotion t…

★ For the kids: Robots watching people-videos to learn chores sounds great, but nobody has fairly tested how much it really helps.

Technical: Ye et al. ICLR 2025 (LAPA); NVIDIA arXiv:2503.14734; Meta V-JEPA 2-AC 2025.

Frozen prediction: By end 2027, a lab that did not train the model publishes a controlled public-benchmark ablation of video/latent-action pretraining vs a no-video baseline — hit promotes to modeled; absent that, stays open (P4 companion: no >=2x multi-lab sample-efficiency win by 2028-06-30).

early / contestedTeleoperation vs autonomy (demo-provenance gap)adversarial skeptic

Held open. 1X CEO-quoted launch statement (Engadget 2025-10-29) is the first vendor-declared teleop share; The Robot Report editorial struck as a disclosure. No intervention rate published anywhere.

★ For the kids: One company now says people will drive its home robot at first; a news column suggesting it is not the company saying it.

Technical: Tesla We, Robot 2024-10-10 (press-reported remote assistance); 1X NEO 2025.

Frozen prediction: No ICRA/RSS/CoRL provenance-disclosure requirement in 2027 author guidelines. p=0.80.

early / contestedNarrow-task humanoid deployment (structured pilots)deployment

HELD open with 2026 numbers that look like progress and are not gradeable evidence. Vendor disclosures now include an eleven-month single-workstation body-shop run (~84s cycle, >1,250 cumulative hours) and >65,000 cumulative operating hours across nine customer facilities. Every one is a cumulative total with no denominator: no interventions per hour, no downtime, no task mix, no third-party audit. Cumulative hours are a marketing unit — a numerator that grows monotonically no matter how badly the fleet performs.

★ For the kids: Robots really did work in factories for thousands of hours — but only the company counted, and only for one easy job.

Technical: Vendor and press disclosures 2025-2026: Figure/BMW body-shop station (single-task; forearm identified by the vendor as top hardware failure point); Agility Robotics >65,000 hours across nine facilities, >100,000 totes at GXO. No third-party audit of any figure.

Frozen prediction: EMB-P3 unchanged: no independently verified >=60-minute unscripted humanoid work run with <2 interventions/hour by 2027-12-31.

reproduced / credibleLow-cost teleoperation & demonstration-data hardwaremanipulation

HELD viable for the reproduced core. ALOHA/UMI low-cost demonstration-data hardware is rebuilt across many labs worldwide. The hand-scale extension (DexCap, Open-TeleVision) is the newest commoditizing layer and grades modeled until it reproduces the way ALOHA did — but the gripper-scale recipe is genuinely reproduced.

early / contestedAxial-flux joint motorsactuation

HELD open (existing grade confirmed): motor class mass-proven in automotive (YASA/Mercedes-AMG) but robot-JOINT integrations remain single-team prototypes with no cross-lab reproduction or teardown evidence in shipping robots.

★ For the kids: A pancake motor proven in sports cars, not yet in robot joints.

Frozen prediction: Through 2028, no independently verified cooling-matched measurement shows an axial-flux robot joint delivering >=30% continuous torque-density gain over best radial-flux QDD.

single-lab / plausibleElectro-hydrostatic actuators (EHA)actuation

HELD modeled. Multi-lab prototype capability (HYDROiD, U-Tokyo Hydra, IIT/Moog HyQReal) plus aerospace lineage, but the field consolidated on electric and no commercial humanoid uses EHA as primary drive — real prototypes, no winning deployment path.

early / contestedMagnetic gears & pseudo-direct driveactuation

HELD open. 25 years on, robot-relevant torque density and cost still have not beaten mechanical reducers; deployments remain lab/aerospace demos; essentially zero robot joints ship magnetic gears. Contested whether it ever crosses over.

early / contestedHot-swap batteries & dock charging (endurance by logistics) — humanoid scope, DEMOTED modeled -> open (two lenses concur)actuation & power

SKEPTIC: two lenses independently demoted this slug for the same reason; merged. The self-swap capability rests on one vendor video (UBTech Walker S2, July 2025) with no swap count or failure rate; the dock-charging leg rests on a vendor statement about its own deployment (Agility Digit at GXO) with no third-party dock-cycle or uptime figure — the same evidence class graded open in narrow-task-humanoid-deployment and humanoid-reliability-data-vacuum. AMR hot-swap/dock charging stays viable under amr-fleet-autonomy-deployment and may not lift this slug by association.

★ For the kids: One company's video of a robot swapping its own battery is not proof it works every time; nobody outside the companies has counted.

Technical: Re-promotion bar: an independent party reports >= 100 consecutive unattended swaps with a failure rate, or third-party dock-cycle counts and uptime for a humanoid deployment.

Frozen prediction: By 2028-09-01 no independent party publishes >= 100 consecutive unattended humanoid self-battery-swaps with a failure rate. p=0.70.

single-lab / plausibleSilicon-anode high-energy-density cells (cell-level, drone-scale)actuation & power

SKEPTIC: grade modeled held, but the lens's retirement of the '<350 Wh/kg' mechanism clause is REVERSED. Verified: Amprius shipped 450 Wh/kg SiCore cells to drone customers (2025-07-22, vendor release) and states 500 Wh/kg as third-party-validated at cell level; a second-gen 500 Wh/kg cell is announced for commercial availability Q4 2026. All cell-level; pack-level density (typically 60-75% of cell) and cycle life at humanoid duty remain unreported, so a pack-level clause is not contradicted. Evidence is vendor-sourced throughout; modeled is the ceiling.

★ For the kids: A company sells extra-dense battery cells to drone makers; a whole robot battery pack is heavier than its cells, so the old rule about packs still stands.

Technical: Amprius press releases 2025-07-22 and 2026-09-02 (vendor); no independent cycle-life measurement located. Open: cycle life at 1C-2C; whether any robot pack ships these cells UN38.3-certified.

Frozen prediction: Through 2027-12-31 no humanoid vendor discloses a >= 400 Wh/kg cell in a shipping robot pack with a stated cycle-life figure; pack-level density in any third-party humanoid teardown stays < 300 Wh/kg. p=0.75 (skeptic-assigned; the proposing lens froze none).

early / contestedRare-earth-free motors & magnetsactuation

HELD open, citation base upgraded from trade press to literature. Coey's review (Engineering, 2020) states the position: NdFeB is a mature, optimized technology, post-2011 alternatives (Sm-Fe-N, iron nitride, tetrataenite) remain niche or pre-commercial, and the author explicitly identifies robotics as a major NEW source of rare-earth magnet demand — expert consensus runs opposite to the substitution narrative. No robot-grade rare-earth-free joint actuator has been demonstrated anywhere.

★ For the kids: Robot motors need special magnets that mostly come from one country. People are trying to make magnets without those rare metals, but nothing works well enough yet.

Technical: Coey, 'Perspective and Prospects for Rare Earth Permanent Magnets', Engineering 6(2):118-130 (2020); Niron iron-nitride at pilot scale; no published robot-grade rare-earth-free joint actuator.

Frozen prediction: Through end of 2028, every credible humanoid teardown still shows NdFeB in the majority of joint actuators, and no peer-reviewed or teardown-verified rare-earth-free robot joint matches NdFeB torque density within 20% at equal mass.

early / contestedRegenerative actuation & cost of transportactuation-power

Single lab a decade old (Seok et al., MIT Cheetah, T-Mech 2015); no second legged platform has reported recovered-energy fraction. Demotion to open upheld.

★ For the kids: A robot leg could put energy back in the battery when it slows, but only one lab has ever measured how much.

Technical: Needs four-quadrant drives with a bus that can sink current; commodity QDD stacks dump braking energy and never report the fraction.

Frozen prediction: By 2027-12-31 a second independent lab publishes measured regenerated-energy fraction for a legged robot over a walking task (P=0.3).

single-lab / plausibleMultimodal hemispherical tactile fingertips (DIGIT 360 class)sensing

HELD modeled — deliberately NOT demoted early, and the reason is recorded so a later audit does not read it as an oversight. On evidence alone this is one engineering lineage with roughly twenty months of a free-hardware distribution programme producing no located external peer-reviewed manipulation result, which reads as low utility or high integration cost rather than slow uptake. But the component already carries a frozen precommitted DECAY whose failure condition IS demotion to open at end 2027; demoting now would pre-empt a frozen prediction and convert a falsifiable claim into a retroac…

★ For the kids: A company gave away amazing fingertip sensors for free, and so far almost nobody has published anything they did with them. We already made a public bet about it, so we wait.

Technical: DIGIT 360 (Meta FAIR + GelSight, Nov 2024; facebookresearch/digit360 open hardware); predecessor DIGIT: Lambeta et al., IEEE RA-L 2020. Trigger check July 2026: no external peer-reviewed manipulation result on distributed units located.

Frozen prediction: By end 2027, >=3 independent non-Meta labs publish peer-reviewed manipulation results on distributed DIGIT 360 units (promotion trigger). PRECOMMITTED DECAY unchanged: if that count is 0-1 at end 2027, DEMOTE to open.

single-lab / plausibleTactile simulation & sim-to-real for touchsensing

PROMOTED open -> modeled (borderline): multiple independent open simulators (Taxim, TACTO, DIFFTACTILE) plus one demonstrated real transfer recipe (Bristol CoRL 2021) — but per this lens's own rule, tooling reuse is not capability reproduction; transferred policies remain narrow. Modeled rests on the single demonstrated transfer, no further.

★ For the kids: Robots can practice feeling inside a computer — but only simple touches so far.

Frozen prediction: By end 2028, no sim-trained tactile contact-rich policy reproduced across two labs at >80% real success on novel objects.

reproduced / credibleJoint-torque proprioception & collision detection (cobot-grade)sensing

HELD viable. Momentum-observer collision detection is textbook (DLR lineage), shipped in torque-sensed arms from multiple independent vendors (KUKA iiwa, Franka) and certified under ISO/TS 15066 across thousands of installations. Reproduced across labs, vendors, and deployments — the incumbency bar every tactile-skin pitch must beat.

early / contestedPre-touch proximity sensingsensing

HELD open for manipulation, premise revised: pre-touch is entering products through the SAFETY door, not the manipulation door. Capacitive proximity skins already detect a human at under ~20 cm and stop the robot, and 2026 work fuses egocentric tactile and proximity sensing for humanoid collision avoidance. For grasping the fifteen-year pattern holds — each generation re-prototypes from scratch with no shared benchmark and no product integration.

★ For the kids: Robots that can sense you before touching you are being sold as a safety feature, not as a better way to pick things up.

Technical: 'Egocentric Tactile and Proximity Sensors as Observation Priors for Humanoid Collision Avoidance', arXiv:2604.25554 (2026); 'Integrated Grasping Controller Leveraging Optical Proximity Sensors', arXiv:2407.05582 (2024); arXiv:2204.02371 (2022); arXiv:2502.01228 (2025); Bosch APAS.

Frozen prediction: Through 2028, pre-touch proximity sensing ships in commercial robots only as a safety/collision channel; no commercial gripper or hand ships pre-touch as a specified GRASPING input, and no cross-lab pre-touch grasping benchmark exists.

reproduced / credibleCompliant & underactuated grippers (soft end-effectors)manipulation

HELD viable. Reproduced across labs since the jamming gripper (PNAS 2010) and iHY/Yale OpenHand (IJRR 2014), and commercially deployed at scale in food/logistics (Soft Robotics mGrip, acquired by Schmalz 2023). Simple compliant end-effectors do the paid work while five-finger hands do the demos.

early / contestedDeformable-object manipulation (cloth, rope, food)manipulation

HELD open (existing grade confirmed): multi-lab garment progress (SpeedFolding, ClothFunnels) is garment-class-specific; crumpled-state estimation unsolved in general; no cross-lab benchmark with shared items; laundry VLA demos are single-organization and unevaluated.

★ For the kids: Folding laundry is still one of the hardest robot jobs — soft things flop into infinite shapes.

Frozen prediction: By end 2028, no third-party evaluation shows >80% end-to-end mixed-laundry folding (>=5 garment types, unstructured pile).

early / contestedExtrinsic dexterity & non-prehensile manipulationmanipulation

HELD open. Real single-lab progress (HACMan, wall-assisted grasping) but nothing reproduced across labs or deployed — a named gap between what dexterous hardware should enable and what current policies exploit.

early / contestedMobile manipulation (base + arm whole-body household tasks)manipulation

HELD open — the self-audit demotion is endorsed and is the correct application of the rule that maturity grades the CAPABILITY, not the tooling. What is reproduced is hardware and teleoperated-imitation pipelines (Mobile ALOHA, TidyBot, Stretch), already credited at viable under low-cost-teleop-data; crediting them twice inflated a second slug. Autonomous mobile manipulation in unstructured homes remains single-lab curated demonstrations with human resets, and long-horizon-task-chaining gives the mechanism: household mobile tasks are the >=10-step regime where success decays multiplicatively.

★ For the kids: Robots that drive around and use an arm are easy to buy and easy to drive by remote control. Doing real house chores by themselves is not solved.

Technical: Fu, Zhao & Finn, Mobile ALOHA, CoRL 2024; TidyBot (Princeton/Google 2023); Hello Robot Stretch; demotion consistent with the grasp-synthesis sim-inflation precedent.

Frozen prediction: Promotes back to modeled only when >=2 independent labs report autonomous (zero-reset, zero-teleop) success on a shared >=5-task household mobile-manipulation suite; through end of 2027 this does not happen.

reproduced / credibleSim-to-real humanoid whole-body controlembodied-ai

REVISION of existing domain (modeled -> viable; promotion upheld after adversarial check, narrowly scoped to the motion-tracking kernel only): RL-trained whole-body motion tracking sim-to-real is reproduced across 3 genuinely independent lineages (UCSD ExBody, Stanford HumanPlus, CMU/NVIDIA OmniH2O/ASAP) on shared commodity Unitree hardware with open code outside groups have re-run and quantitative tracking metrics. Skeptic conditions binding: (a) census counts 3, not 5, independent lineages; (b) task-purposeful autonomous loco-manipulation inherits NOTHING from this grade (stays open via tel…

★ For the kids: Many different university teams can now make the same store-bought robot copy human dance and walk moves after practicing only inside a computer — checked and repeated, so this one is real. Doing useful jobs with those moves is a different, unproven thing.

Technical: Cheng et al., ExBody, RSS 2024 (UCSD); Fu et al., HumanPlus, CoRL 2024 (Stanford); He et al., OmniH2O CoRL 2024 + ASAP RSS 2025 (CMU/NVIDIA, single lineage); HOVER 2024. Recipe: Isaac Gym/Lab massive parallel sim + retargeting + domain randomization + residual correction, on Unitree H1/G1.

Frozen prediction: Frozen test emb-p3: >=3 additional independent labs publish quantitative sim-to-real WBC tracking results on commodity humanoids by 2027-03-31 (confirming); any documented reproduction failure triggers re-demotion to modeled.

early / contestedReproducible VLA evaluation (benchmark gap)evaluation

SKEPTIC RESOLUTION of a direct grade conflict: two lens batches returned this same slug at different maturities (open after a double-count audit; modeled elsewhere). Resolved to the STRICTER grade, because the demotion argument is uncontested and the holding argument does not answer it. Every referee cited in support — RoboArena, TRI blind A/B, AutoEval, Colosseum, SIMPLER — is already the evidentiary basis of a different component (third-party-policy-evaluation, blind-ab-policy-evaluation, perturbation-robustness-evaluation). Credited once each, the residual unique to this slug is LIBERO (sa…

★ For the kids: We thought robot report cards were finally real — but we had counted the same referee three times, and the test itself turned out to be too easy.

Technical: Residual evidence: LIBERO (NeurIPS D&B 2023, saturated by OpenVLA-OFT arXiv:2502.19645); vla-eval (arXiv:2603.13966) and VLA-REPLICA (arXiv:2605.20774) carry no independent-adoption record; RoboDojo (arXiv:2607.04434, 2026) is new and untested as a community standard. Reassigned away: RoboArena (arXiv:2506.18123), TRI LBM blind A/B, AutoEval, Colosseum (RSS 2024), SIMPLER (CoRL 2024).

Frozen prediction: EMB-P13 (held): re-promotes to modeled only when a single VLA harness accumulates independently-run results from >=5 external labs on an identical fixed protocol by 2027-12-31; viable additionally requires that harness to report per-factor perturbation degradation AND confidence intervals as default output.

single-lab / plausibleSample-efficient real-world RL for manipulation (SERL / HIL-SERL)learning

HELD modeled (consolidates the former real-world-rl-finetuning slug — same HIL-SERL result). SERL/HIL-SERL (UC Berkeley) reach near-100% on contact-rich tasks in ~1-2.5h of real-robot training — a genuine counterpoint to imitation-only trends, open-sourced and re-implemented — but independent cross-lab reproductions remain thin. On the cusp of viable with a preregistered cross-lab-replication promotion trigger.

Frozen prediction: Promote to viable on an independent cross-lab replication of HIL-SERL near-100% contact-rich results outside the origin lab.

early / contestedCross-embodiment data & robot data scaling laws (DEMOTED modeled -> open)manipulation & dexterity / data scaling

Demotion ACCEPTED. Single-lab scaling fit (Lin et al., Tsinghua 2024) unreplicated on disjoint data; the largest 2026 cross-embodiment model (RynnBrain 1.1) calls its own gains 'initial evidence'.

★ For the kids: People hoped piling up data from many robots would help every robot predictably; the newest big study says 'early signs'.

Technical: Falsifier: two independent labs publish scaling exponents on disjoint corpora agreeing within 20%.

Frozen prediction: By 2027-09-05 no two independent labs publish cross-embodiment data-scaling exponents on disjoint corpora that agree within 20%. p=0.80.

single-lab / plausibleDual-system VLA control stacks & onboard inferencelearning

HELD modeled at the floor of the tier, with the evidence base corrected and a note on how close this ran to open. Figure's Helix report is struck — it belongs to the permanently retired vendor-showreel class. What remains admissible is GR00T N1 (arXiv paper with RELEASED OPEN WEIGHTS — a checkable artifact, which is what separates it from a marketing spec sheet) plus Gemini Robotics On-Device (vendor report, discounted). One admissible leg = single-lab = modeled. The slow-VLM/fast-controller split is an architectural CONSENSUS among vendors, not a reproduced RESULT, and if the open-weight art…

★ For the kids: Robot brains are being split into a slow thinker and a fast reflex. Lots of companies say it works; only one of them showed its homework.

Technical: NVIDIA GR00T N1, 2025 (arXiv:2503.14734), System-2 VLM + System-1 diffusion transformer, open weights; Gemini Robotics On-Device (vendor report); Figure Helix (STRUCK — vendor showreel class).

Frozen prediction: EMB-P5 (unchanged): an open-weight fully-onboard dual-system stack (<150 W, >=50 Hz) is reproduced on two embodiments outside the origin lab by 2027-12-31.

single-lab / plausibleThird-party policy evaluation — PROMOTION REVERSED (viable -> modeled)manipulation & dexterity / evaluation

SKEPTIC REVERSAL. The capability claimed as reproduced — 'blind third-party evaluation yields stable rankings' — has been shown once (RoboArena's ~0.92 correlation vs its own oracle, CoRL 2025). RoboChallenge (arXiv 2510.17950, 38 authors, verified) is a second evaluation system, but no cross-referee ranking agreement has been published; the lens's own prediction rob-2026-09-05-p8 (rank correlation >= 0.7 between the two) is unresolved, so promoting on it grades the future. Also, another lens in this pass found the RoboArena public leaderboard rendered empty on 2026-09-05. Modeled until p8 re…

★ For the kids: Two separate referee groups test robot brains blind; nobody has yet checked whether the two groups agree.

Technical: Atreya et al., CoRL 2025 (PMLR v305); RoboChallenge arXiv 2510.17950; RoboDojo arXiv 2607.04434. Promotion bar = a published cross-referee comparison over >= 5 shared policies with Spearman >= 0.7.

Frozen prediction: By 2027-06-30 a published comparison of RoboArena and RoboChallenge rankings over >= 5 shared policies shows Spearman >= 0.7. p=0.55.

early / contestedHumanoid fleet reliability (MTBF vacuum)deployment

Vendor statements (Agility >65,000 h; Figure at BMW) contain no MTBF or intervention rate. Open.

★ For the kids: Robots have worked tens of thousands of hours but the companies do not say how often they break.

Technical: 65,000 h exposure would bound MTBF within 2x if failures were logged.

Frozen prediction: By 2027-03-01 no humanoid vendor publishes an audited MTBF or interventions-per-hour for a paid deployment (P=0.8).

early / contestedHumanoid safety standards & certification (the deployment gate)deployment

HELD open, with the lens's own over-precision corrected. The defensible public position as of mid-2026 is only that the dynamic-stability humanoid standard remains PRE-PUBLICATION — a working draft circulated in 2025, comment period closed, no ratified publication date, industry estimates of 18-36 months to ratification. ISO 10218:2025 still does not cover dynamic stability. The certification gate insurers and plant owners need does not exist in force. Worth tracking separately: the project leadership is drawn from the regulated parties.

★ For the kids: The safety rulebook for walking robots is still being written, so nobody can be officially approved yet.

Technical: ISO/TC 299 work item 25785-1 (dynamically stable industrial mobile robots), working draft circulated 2025, comment period closed; ISO 10218:2025 excludes dynamic stability. Prior 'Committee Draft' stage claim retired as unsupportable precision.

Frozen prediction: Through 2027-12-31, no humanoid is certified to a published, in-force dynamic-stability safety standard, and deployment scale stays capped by insurance and liability rather than by capability.

reproduced / credibleSeries-elastic & compliant force-controlled actuators (SEA)actuation

KEPT viable — this is what the viable bar looks like: introduced MIT Leg Lab 1995, independently reproduced for three decades across NASA Valkyrie, Rethink Baxter/Sawyer, ETH ANYdrive, and commercial exoskeletons; deployment-grade, multi-lab, multi-decade. Honest caveat retained: the current humanoid wave is moving to stiffer QDD + roller-screw joints for bandwidth, relegating SEA to cobots, exoskeletons, and human-proximal machines. Complements (does not duplicate) electric-qdd-actuators.

★ For the kids: Putting a spring inside a robot's joint lets it push gently and safely, like a person's stretchy muscles and tendons.

Technical: Pratt & Williamson, IEEE/RSJ IROS 1995; Radford et al., Journal of Field Robotics 2015 (Valkyrie); Hutter et al., ANYdrive/ANYmal (ETH, 2016+). Trade-off: elastic element caps torque-control bandwidth.

Frozen prediction: Of commercial humanoids publicly announced or shipping 2026-2028, fewer than 25% will use series-elastic actuators as the primary joint drive; QDD and roller-screw architectures will dominate, while SEA keeps the cobot/exoskeleton niche.

single-lab / plausibleShape-memory alloy (SMA) actuatorsactuation

KEPT modeled, completing the artificial-muscle family alongside existing hasel/TCP/DEA entries: SMA is a reproduced MATERIAL capability (many labs, niche products) but NOT a reproduced robot-actuation capability — Carnot-limited 1-5% efficiency, ~1 Hz uncooled bandwidth, hysteresis. Modeled is the ceiling; any future promotion requires a shipping robot joint, not more materials papers.

★ For the kids: Some metal wires shrink like muscles when heated, but they waste almost all their energy as heat and cool down too slowly to move a robot fast.

Technical: Jani et al., Materials & Design 56 (2014); Mirvakili & Hunter, Advanced Materials 30(6), 2018. Three orders of magnitude below electromagnetic joints on power-per-efficiency.

Frozen prediction: Through end of 2029, no commercially shipping humanoid or quadruped will use SMA as its primary joint actuation; SMA robot use stays confined to grippers, valves, and micro-mechanisms.

early / contestedSodium-ion cells for robots (the lithium-supply hedge)power

KEPT open. Cell chemistry is commercially real (CATL/BYD/HiNa), but no published mobile-robot integration exists and the ~30-40% energy-density penalty lands on the binding runtime constraint (see actuator-power-energy-density). Vendor Wh/kg figures (Naxtra ~175) remain unverified in packs.

★ For the kids: A new battery made with cheap salt instead of rare lithium works fine, but it is heavier for the same energy, and walking robots care about every gram.

Technical: Hwang, Myung & Sun, Chem. Soc. Rev. 46, 3529 (2017); CATL Naxtra line (announced 2025, vendor figure). No published mobile-robot integration.

Frozen prediction: Through end of 2028, no commercial humanoid will ship with a sodium-ion main pack; if sodium-ion enters robotics at all, it appears first in charging docks, swap-station buffers, or wheeled AGVs.

early / contestedHigh-C-rate cells & opportunity charging (the third endurance path)power

SKEPTIC DEMOTION UPHELD (modeled -> open). The ingredients are individually reproduced — 4C LFP cells ship in EVs (CATL Shenxing 2023), ML fast-charge protocols are peer-reviewed (Attia et al., Nature 2020), and Digit dock-charges autonomously at GXO (2024, single vendor). But the composed capability this component claims — multi-shift humanoid operation sustained by opportunity charging alone — has never been publicly documented even once. Component-level evidence does not compose into system-level maturity. Promotes to modeled on the first documented multi-shift dock-charged operation.

★ For the kids: Instead of a bigger battery, the robot could take lots of tiny charging naps between jobs — but no robot has actually worked a full double shift this way yet.

Technical: Attia et al., Nature 578, 397-402 (2020); CATL Shenxing 4C LFP (2023); Agility Digit autonomous dock-charging within the GXO RaaS deployment (2024) — one vendor, one site class, no published multi-shift-on-opportunity-charge-alone record from anyone.

Frozen prediction: By end of 2028, at least one humanoid vendor will publicly document multi-shift operation sustained by >=2C opportunity dock-charging without pack swaps; if none does, the fast-charge endurance path is falsified for this cycle and swap logistics wins the interim.

reproduced / credibleSix-axis wrist force/torque sensingsensing

KEPT viable (new; distinct from joint-torque-sensing-collision, which covers joint-level proprioception) — decades of commercial shipping (ATI since the 1990s, Bota, Robotiq, Schunk), reproduced in essentially every manipulation lab, used as ground truth in peer-reviewed calibration studies. Cost and overload fragility are the open engineering axes, not capability.

★ For the kids: Robots have had a solid sense of push-and-twist in their wrists for many years — that part of touch is truly done and sold in stores.

Technical: ATI multi-axis F/T line (1990s-); Bota Systems (ETH spinoff); ground-truth usage in Yuan et al., Sensors 2017 and peers.

Frozen prediction: Through 2028, wrist 6-axis F/T plus joint-torque sensing (not fingertip tactile) will remain the only force-sensing modality present in the majority of commercially deployed humanoid and cobot units.

single-lab / plausibleScalable tactile gloves for human-grasp capturesensing

HELD modeled (existing grade confirmed): capture reproduced (MIT Nature 2019 glove; OSMO 2025) but no cross-lab result shows glove-captured touch data improving robot policies.

★ For the kids: A cheap sensor glove records exactly how people squeeze things so robots can copy our touch.

Frozen prediction: By end 2028, one glove-data-trained policy beats a vision-only-human-data baseline in a published head-to-head — in a single lab only.

single-lab / plausibleEvent-driven (neuromorphic) tactile sensingsensing & tactile

Raise open -> modeled ACCEPTED: sensors built in three labs (Bristol NeuroTac, NUS NeuTouch, Darmstadt Evetac) and Evetac adopted by a second lab. The capability claim (kHz touch improves control) is single-lab with trial count unreported; not viable.

★ For the kids: A few labs built touch sensors that react a thousand times a second; no one has shown twice that a robot works better because of it.

Technical: Ward-Cherrier et al. ICRA 2020; Taunyazov et al. RSS 2020; Funk et al. T-RO 2024; Krohn et al. arXiv 2606.06281 (2026).

Frozen prediction: By 2027-09-05 a lab other than TU Darmstadt reports closed-loop manipulation with event-based touch at >= 500 Hz and >= 30 trials per condition. p=0.35.

early / contestedAcoustic / contact-microphone sensingsensing

HELD open, and the extension is the diagnosis. There are now three lineages — SonicSense passive impact audio, ManiWAV manipulation audio, and Sound of Touch active string excitation (2026) — with ZERO cross-lab comparison between them and no shared protocol. Three independent inventions and no referee is exactly the pattern that keeps a near-free BOM addition permanently pre-competitive.

★ For the kids: Robots can hear what they touch — three different teams found three different ways, and none of them tested against each other.

Technical: Yi, Xing, Manchester & Fazeli, 'Sound of Touch', arXiv:2602.16846 (Feb 2026, Michigan x CMU); SonicSense, CoRL 2024 (Duke); ManiWAV, CoRL 2024 (Stanford/Columbia).

Frozen prediction: Through 2028, no commercially shipped manipulation product uses contact audio as a specified sensing channel, and no published work compares two of the three acoustic-sensing lineages on a shared task.

single-lab / plausibleBimanual dexterous manipulation (two-arm coordination)manipulation

KEPT modeled: ALOHA Unleashed's 78-91% results required ~26,000 demonstrations and exist only in the collecting lab; TRI's LBM study confirms the recipe but no third party has reproduced the success rates on unseen tasks. Single-lab capability at extreme data cost.

★ For the kids: Robots are learning to use two hands together, like tying shoelaces, but they need thousands of practice lessons from people first.

Technical: Zhao et al., ALOHA Unleashed, CoRL 2024 (arXiv:2410.13126); Fu et al., Mobile ALOHA 2024; TRI Large Behavior Models 2025 (blind A/B).

Frozen prediction: Through end of 2027, no policy from any lab or humanoid maker will show >=80% success on >=5 unseen bimanual household tasks in a third-party double-blind evaluation network (RoboArena-class).

early / contestedDexterous multi-finger grasp synthesismanipulation

DEMOTED modeled -> open (self-audit demotion retained): the reproduced artifact is the simulation pipeline (DexGraspNet-class), not real grasping — papers reporting both show consistent double-digit sim-to-real drops on novel objects; maturity grades the capability, not the tooling.

★ For the kids: Great at the video-game version of grabbing; real hands still fumble.

Frozen prediction: Through 2027, real-hardware novel-object success stays >=15pp below the same method's sim success wherever both are reported.

single-lab / plausibleDexterous teleoperation & hand-motion retargetingmanipulation

KEPT modeled (distinct from low-cost-teleop-data, which covers gripper-class rigs and is viable): AnyTeleop/Open-TeleVision/DexCap are open-source and re-used across labs, but retargeting remains lossy (kinematic mismatch, no force feedback) and no commercial adoption is teardown-confirmed. On the gripper-teleop path to viable, not there yet.

★ For the kids: People wear cameras and VR headsets to puppet robot hands, teaching them finger moves.

Technical: Qin et al., AnyTeleop, RSS 2023; Cheng et al., Open-TeleVision, CoRL 2024; Wang et al., DexCap, RSS 2024.

Frozen prediction: By end of 2027, at least one dexterous-hand teleop/retargeting stack will be adopted (published or teardown-confirmed) by >=3 independent labs AND >=1 commercial humanoid data operation.

reproduced / credibleTactile-reactive grasp control (slip detection & grip-force servoing, two-contact grippers)sensing-tactile

Viable held on the NARROW claim only: slip detection and grip-force servoing on parallel-jaw/two-contact grippers, demonstrated independently at MIT (GelSight 2017/18), TU Darmstadt (BioTac 2015), Bristol (TacTip 2021) and CMU/Meta (ReSkin 2021) with different sensing physics. Whole-hand reactive grasp on novel objects is NOT reproduced and stays under dexterous-in-hand-manipulation [modeled]. Skeptic caveat: reproduction is of the capability, not of a shared protocol; any widening of the claim demotes it.

★ For the kids: When a cup starts to slide, a two-finger robot feels the slip and squeezes a bit harder.

Technical: Slip signals: marker shear divergence (optical, 30-60 Hz), vibration band-power (BioTac, kHz), magnetic-field derivative (ReSkin).

Frozen prediction: HELD: by end 2027 at least one commercial humanoid or cobot hand ships slip-reactive grip-force SERVOING (closed loop, not detection alone) as a documented standard software feature tied to integrated tactile sensing.

single-lab / plausibleDynamic high-speed manipulation (throwing, catching, striking)manipulation

KEPT modeled: TossingBot and DeepMind table tennis are peer-reviewed single-lab demonstrations; no dynamic skill has cross-lab reproduction at matched performance, and humanoid relevance is untested. Demonstrated-not-reproduced is precisely modeled.

★ For the kids: Some robots can throw things into the right bins and even play ping-pong against people.

Technical: Zeng et al., TossingBot, RSS 2019/T-RO 2020; D'Ambrosio et al. (Google DeepMind), robot table tennis, 2024 (arXiv:2408.03906).

Frozen prediction: Through 2028, no dynamic manipulation skill (throw/catch/strike) will be reproduced across two independent labs at matched performance on shared task specs.

single-lab / plausibleOpen-weight VLA fine-tuning to new tasksembodied-ai

HELD modeled — the demotion from viable is endorsed as the correct application of the evidence bar. Fine-tuning an open VLA to a new task is widely practised and widely reported to work, but the evidence base is statistically underpowered: modal sample sizes of 10-20 trials per condition, essentially no confidence intervals or paired tests, and sequential-testing analysis showing 20-30 real trials cannot support the comparisons being claimed. 'Many labs report success' is popularity, not reproduction. Restore to viable only on powered, cross-lab evidence.

★ For the kids: Lots of teams say they retrained a robot brain and it worked — but each only tested a handful of times, so we cannot be sure yet.

Technical: Kress-Gazit, Hashimoto, Kuppuswamy, Shah, Horgan, Richardson, Feng & Burchfiel, arXiv:2409.09491 (2024); Snyder et al., STEP, arXiv:2503.10966 (2025). Contrast the LBM protocol: N=50 real / 200 sim per condition, paired Barnard/Welch, Bonferroni correction.

Frozen prediction: By 2027-12-31, fewer than 20% of real-robot VLA fine-tuning papers at ICRA, CoRL or RSS report confidence intervals or paired statistical tests on their headline success rates.

single-lab / plausibleRL post-training of VLAs from real deployment experienceembodied AI

Promotion open -> modeled ACCEPTED with a flag: this is the weakest modeled in the set. Real-world RL post-training of a generalist VLA is one company (pi*0.6/RECAP, developer-evaluated, endurance demos struck as evidence); the multi-lab leg is sim-only. Modeled because sim reproduction exists; reverts to open if no open checkpoint plus third-party count appears by 2027-09.

★ For the kids: The robot practices and a coach marks good tries; only one company has shown it on real robots.

Technical: arXiv 2511.14759; SimpleVLA-RL, VLA-RL (2025, sim).

Frozen prediction: By 2027-08-21 at most one lab other than Physical Intelligence publishes a real-robot VLA RL post-training result with >=1.5x throughput on a >=10-minute continuous task, n>=20 episodes. Probability 0.7.

single-lab / plausibleOnboard robot inference compute (the embedded-TOPS bottleneck)embodied AI / compute

Held modeled. Each speedup result is one group; energy axis split to policy-inference-energy-per-success.

★ For the kids: Small computers on robots are getting faster tricks to think quickly.

Technical: Jetson-PI arXiv 2607.12659; arXiv 2604.24447; EcoVLA arXiv 2608.15502.

Frozen prediction: By 2027-08-21 no >=100B total-parameter embodied model runs as the closed-loop action policy entirely onboard a battery-powered humanoid at >=5 Hz. Probability 0.8.

single-lab / plausibleBlind randomized A/B policy evaluationevaluation-integrity

KEPT modeled: TRI's blind randomized CI-reported protocol and Berkeley's AutoEval are each single-organization; two organizations with DIFFERENT protocols is demonstration, not cross-lab adoption of a standard. Modeled, watching for spread.

★ For the kids: Scientists test robots the way doctors test medicine: nobody doing the scoring knows which robot brain is which, so nobody can cheat, even by accident.

Technical: TRI LBM study, arXiv 2025 (blind randomized A/B, CIs, thousands of rollouts); Zhou et al., AutoEval, 2025 (automated 24/7 real-world eval).

Frozen prediction: By 2027-12-31, at least 3 distinct labs or companies will publish real-robot evaluations explicitly using blind randomized A/B protocols; falsified if fewer than 3 exist.

early / contestedIntervention-rate and autonomy-duration metricsevaluation-integrity

KEPT open (companion to humanoid-reliability-data-vacuum, which tracks MTBF; this tracks interventions-per-autonomous-hour): the single most decision-relevant number for humanoid economics is published by essentially nobody in third-party-checkable form. Until it appears, all 'autonomous fleet' claims are ungraded.

★ For the kids: The honest question for any robot is not 'can it do the trick once' but 'how many minutes until a human has to rescue it' — and right now almost no company will say.

Technical: Proposed contract: mean time between interventions + intervention taxonomy + task-mix disclosure over >=100 continuous hours; CA DMV disengagement reporting as the imperfect template.

Frozen prediction: No humanoid vendor publishes third-party-audited interventions-per-hour data covering >=3 customer sites by 2027-12-31; falsified (happily) if one does.

single-lab / plausiblePerturbation-robustness evaluation (Colosseum-class)evaluation

Demoted viable -> modeled (three batches independently concur). The Colosseum (Pumacay et al., RSS 2024, UW/NVIDIA) is a single-lab simulated suite whose real-robot validation is a small in-lab study; wide citation is not reproduction of its predictive validity.

★ For the kids: One lab built an obstacle course for robot brains. Lots of people run it, but nobody else has checked that doing well on it means the robot copes in a real kitchen.

Technical: 14 perturbation axes; real-robot correlation reported only by the authoring lab on a handful of tasks.

Frozen prediction: By 2027-08-21 no non-author lab publishes a real-robot study of >=5 policies with Colosseum rank preserved (rho>0.7).

early / contestedHumanoid unit economics (the dollars-per-productive-hour vacuum)deployment-truth

KEPT open: sell-side BOM teardowns exist but zero audited customer-side dollars-per-productive-hour data from any real deployment does; RaaS pricing is confidential and excludes intervention labor, supervision, and downtime. Any TAM claim built on unaudited $/hr is speculation.

★ For the kids: A robot only saves money if the whole bill — babysitting humans included — is smaller than the wage it replaces, and nobody has shown that bill yet.

Frozen prediction: By 2027-12-31 no customer (not vendor) will have published an audited all-in cost-per-productive-hour for a humanoid deployment; falsified if one does.

single-lab / plausibleWheeled-legged hybrid locomotion — PROMOTION REVERSED (viable -> modeled)manipulation & dexterity / mobile base

SKEPTIC REVERSAL. The km-scale learned autonomous urban-mission capability is one lab (ETH RSL, Science Robotics 2024). The 'reproductions' offered are shipped products (Unitree B2-W, DEEP Robotics Lynx) — the reproduced quantity is the product, not the published capability, which is exactly the rule applied to quadruped-inspection-autonomy on 2026-08-26 — and the '2026 independent group' (arXiv 2604.23761, verified) lists DeepRobotics among its affiliations, so it is a vendor-co-authored paper, not an independent reproduction. Another lens in this same pass held this slug at modeled; consist…

★ For the kids: Robots with wheels on their feet work well in one lab and are sold as products; a second team without the maker on the paper has not yet shown the same thing.

Technical: Lee et al., Science Robotics 9 (2024); Bjelonic et al., RA-L 2020; Zhao et al., arXiv 2604.23761 (Tianjin/NUS/DeepRobotics, 2026). Re-promotion: a second lab with no vendor co-authorship publishes km-scale autonomous wheeled-legged missions, or manipulation-while-rolling success rates with n >= 50.

Frozen prediction: By 2027-06-30 at least one peer-reviewed paper reports whole-body loco-manipulation success on a wheeled-legged base with n >= 50 trials per task. p=0.60.

single-lab / plausibleHumanoid hardware cost collapse (the reproduction substrate)deployment

KEPT modeled: the price datapoints are hard (H1 $90k 2023 -> G1 $16k 2024 -> R1 ~$5.9k 2025; Berkeley Humanoid Lite ~$5k open-source) but capability-at-price is unproven — sub-$20k platforms have no proven autonomous task capability. The visible causal effect is scientific (cheap Unitree hardware enabled the cross-lab WBC reproduction), not commercial.

★ For the kids: Robot bodies are getting cheap fast — a walking robot now costs about as much as a used car — so many more labs can test the same experiments.

Technical: Unitree G1 May 2024 $16k (verified by purchasing labs); R1 July 2025 ~$5,900 (vendor); Berkeley Humanoid Lite 2025 (~$5k BOM, 3D-printed cycloidal gearboxes).

Frozen prediction: By end of 2027, >=3 independent labs will publish peer-reviewed sim-to-real results on sub-$10k commercial humanoids. Separately: through 2027, no sub-$20k humanoid will clear >=50% success on any third-party manipulation benchmark of >=10 tasks.

reproduced / credibleCompact hydraulic legged actuation (displaced incumbent)actuation

HELD viable: reproduced across multiple independent organizations over two decades (Boston Dynamics BigDog/hydraulic Atlas, IIT HyQ/HyQReal, Sarcos). Capability tier viable; deployment relevance declining after BD's 2024 electric pivot.

★ For the kids: Oil-powered robot muscles really work, but they are messy, so builders are switching to electric.

Technical: Semini et al. 2011; HyQReal 2019 aircraft-tow; BD electric-Atlas transition 2024.

Frozen prediction: By end 2028, no newly launched commercial humanoid uses hydraulic primary joint drive.

single-lab / plausibleLiquid-cooled electric actuatorsactuation & power

Single-lab peer-reviewed evidence (Urata et al., IROS 2010, JSK) of a several-fold continuous-torque gain from liquid cooling on humanoid motors. 2024-2025 vendor claims of liquid-cooled joints are unverified and add no tier.

★ For the kids: Putting water pipes around a robot's motors keeps them cool so they can push harder for longer, like a car radiator.

Technical: Urata, Nakanishi, Okada & Inaba, IROS 2010. Cooling moves the binding constraint to pump/radiator mass and seal life, neither quantified over >1000 h in a robot.

Frozen prediction: By end of 2028, liquid-cooled joint actuators appear in at least two teardowns from vendors with NO Sentis/UT-Austin lineage; Apptronik alone does not satisfy this.

early / contestedWide-bandgap (GaN/SiC) joint motor drivespower

MERGED duplicate slugs (gan-sic-motor-drive-electronics + wide-bandgap-motor-drives) and SKEPTIC-DEMOTED modeled -> open: device physics is EV-proven and 30-50 inverters per humanoid make it relevant, but the robotics claim rests entirely on vendor reference designs (EPC/TI) and automotive analogy — zero robot demonstrations, zero teardown evidence anywhere. Automotive proof does not transfer to the robot-joint claim.

★ For the kids: New chip materials waste less power as heat, so robot joints could get smaller and cooler — but nobody has shown them inside a robot yet.

Technical: Millán et al. IEEE TPEL 2014; EPC/TI GaN servo reference designs 2022-2024; no humanoid-teardown confirmation as of 2026.

Frozen prediction: By end 2028, either an independent teardown documents GaN/SiC joint-drive stages in a shipping humanoid/quadruped, or silicon MOSFET drives are confirmed standard and the domain stays open.

early / contestedMusculoskeletal tendon-driven humanoidsactuation

MERGED duplicate slugs (musculoskeletal-tendon-actuation + musculoskeletal-humanoids) and SKEPTIC-DEMOTED to open: the Univ. of Tokyo JSK line (Kenshiro/Kengoro) is scientifically rich but a single-lab program — venue prestige (Science Robotics) is not an evidence tier; no independent reproduction and no benchmarked capability advantage over rigid-link humanoids exists anywhere. Kept as the contrast case defining what the deployment stack gave up.

★ For the kids: One lab builds robots with string muscles like a human body — amazing to study, but simpler robots still work better.

Technical: Asano, Okada & Inaba, Science Robotics 2017 (Kengoro; perspiration cooling). Single-lab (JSK).

Frozen prediction: Through 2029, no musculoskeletal humanoid exceeds conventional rigid-link humanoids on any independent benchmark, and none is commercially deployed as primary architecture.

early / contestedSupercapacitor / hybrid peak-power bufferingpower

HELD open: physics established (MIT Cheetah regenerative drivetrain, TMech 2015) but supercap-buffered legged platforms remain scattered early demonstrations without cross-lab reproduction.

★ For the kids: A quick-burst energy pocket next to the battery, for jumps.

Frozen prediction: Through 2028, no shipping humanoid discloses a supercapacitor-buffered power architecture.

early / contestedStructural batteries (mass-free energy)power

MERGED duplicate slugs (structural-batteries-robotics + structural-battery-composites), HELD open: Chalmers multifunctional cells ~24 Wh/kg at coupon scale (AESR 2021), Imperial structural supercapacitors — an order of magnitude below packaging-grade Li-ion, no robot integration exists.

★ For the kids: What if the robot's bones WERE the battery? The bone-battery is still weak.

Technical: Asp et al., AESR 2021; Imperial College Greenhalgh group.

Frozen prediction: Through 2029, no commercial robot ships load-bearing structural-battery members; no lab walking robot draws >5% of mission energy from them; cell-level stays <100 Wh/kg at structural stiffness.

early / contestedHydrogen fuel-cell power for mobile robotspower

HELD open: commercial in drones (Doosan 2h+), but legged evidence is concept-grade — Kawasaki CORLEO (2025) is a concept reveal, not a capability; burst-power mismatch forces battery hybridization anyway.

★ For the kids: Hydrogen could let robots work much longer, but for walking robots it is still just a cool idea.

Frozen prediction: Through 2029, no fuel-cell-primary legged robot is commercially deployed with paying customers.

single-lab / plausibleVariable-stiffness / variable-impedance actuatorsactuation

HELD modeled (distinct from series-elastic-actuators, which stays viable): hardware reproduced across academic labs (DLR, IIT, Twente; EU VIACTORS/SAPHARI) but the claimed advantage over software impedance was never proven in deployment; every shipping humanoid uses fixed gearing.

★ For the kids: Robot joints with real adjustable springs inside — store robots do not use them because the springs add weight.

Technical: Vanderborght et al. RAS 2013; Grebenstein et al. ICRA 2011.

Frozen prediction: Through 2028, no commercially shipping humanoid uses mechanically variable-stiffness primary joint drives.

single-lab / plausibleTendon-driven transmission (hands and slim limbs)actuation

HELD modeled: reproduced across Shadow (commercial), DLR Awiwi, UW biomimetic hand — but tendon wear is the unpriced consumable; no independent million-cycle durability result exists.

★ For the kids: Puppet-string robot hands work, but the strings wear out and must be replaced.

Frozen prediction: Through 2028, no independently tested tendon-driven hand demonstrates >1,000,000 grasp cycles without tendon service.

single-lab / plausibleGecko-adhesive & electroadhesive grippersmanipulation

HELD modeled (self-audit demotion from viable retained): reproduced principle across Stanford/JPL (microgravity) and EPFL (electroadhesion), lightly commercialized — but dust sensitivity, surface dependence, and adhesion decay in industrial duty are undocumented; niche end-effector, not dexterity substitute.

★ For the kids: Robot fingers that stick like gecko feet or cling with static electricity.

Frozen prediction: Through 2028, adhesion grippers stay confined to flat/rigid-surface niches; no humanoid hand integrates them as primary grasp mechanism.

single-lab / plausibleElectrostatic clutches (zero-power holding)power

HELD modeled: reproduced across CMU (ICRA 2016) and EPFL/ETH (DextrES, UIST 2018) in wearables — milliwatt torque holding directly relevant to humanoid posture power, but absent from every humanoid joint.

★ For the kids: Static-cling joint locks let a robot hold still for almost no battery.

Frozen prediction: Through 2028, electrostatic clutches ship in exoskeletons/haptics but in no commercial humanoid joint.

single-lab / plausibleWhole-hand high-resolution tactile integration (F-TAC class)sensing & tactile

Raise open -> modeled ACCEPTED on a peer-reviewed single-lab result with 600 trials (PKU F-TAC, NMI 2025). Sharpa Wave press carries zero weight. No second lab.

★ For the kids: One lab built a hand that feels almost everywhere and tested it 600 times; others still have to copy it.

Technical: Zhao et al., Nature Machine Intelligence, June 2025; arXiv 2412.14482. Missing: durability of 17 gels/hand.

Frozen prediction: By 2027-09-05 a second lab publishes a peer-reviewed hand with >= 50% surface coverage by high-resolution touch and >= 200 real trials. p=0.30.

single-lab / plausibleVisuo-tactile policy learning (the 2026 tactile world-action-model cluster)manipulation & dexterity / touch

Held modeled. Five single-lab tactile WAM preprints in three months, no shared benchmark, no third-party evaluation.

★ For the kids: Many teams teach robots to imagine a touch before acting, all on different tests.

Technical: arXiv 2606.08737, 2606.26663, 2607.02840, 2607.20683, 2608.19574.

Frozen prediction: By 2027-03-31 no tactile world-action-model paper reports results on a shared physical tactile benchmark scored by an independent group. p=0.80.

early / contestedTactile evaluation benchmarks (missing referee)sensing

HELD open, with the 2026 evidence upgraded and the trigger sharpened. Real artifacts finally exist — the first openly licensed, physically reproducible 3D-printed texture benchmark, TacVerse's cross-sensor dataset, TacO, Tactile MNIST, and an open two-stage characterization pipeline. It stays open because the 3D-print benchmark itself shows why: within-printer generalisation is strong but CROSS-PRINTER generalisation fails on geometric inconsistency. When even the physical reference artifact is not reproducible across labs, no benchmark can referee.

★ For the kids: People finally made shared test objects for robot touch — and then found that printing the same test object on a different printer changes the answer.

Technical: Shepherd, Herzig, Husbands, Philippides, Johnson & Kimbell (Sussex), arXiv:2606.25886 (June 2026); Wei et al., TacVerse, arXiv:2606.25877 (2026); 'TacO', arXiv:2605.21976 (2026); 'Tactile MNIST' (2025); Lo Preti et al., Advanced Intelligent Systems 2026.

Frozen prediction: Through 2027, no tactile benchmark accumulates published results from >=5 independent external labs on identical protocols; and through 2028 no tactile benchmark demonstrates cross-lab reproduction of its own PHYSICAL reference artifact.

single-lab / plausibleArticulated-object manipulationmanipulation

HELD modeled, and deliberately NOT demoted — the reasoning is recorded so a future audit does not read it as an oversight. The tooling-vs-capability rule that demoted mobile-manipulation would also bite here (sim pipelines reproduced, real-world unseen-mechanism success inconsistent), but this slug already carries a frozen end-2027 falsifier whose failure condition IS demotion to open. Grading it down now would pre-empt a frozen prediction and convert a falsifiable claim into a retroactive one, which the freeze discipline forbids.

★ For the kids: We nearly lowered this grade too, but we already made a public bet about it that gets checked in 2027. Changing the grade early would be cheating on our own test.

Technical: SAPIEN / PartNet-Mobility sim lineage; ETH and CMU real-hardware door opening; frozen falsifier due end-2027.

Frozen prediction: Unchanged and frozen: by end 2027, >=2 independent labs publish >=80% success on unseen doors/drawers/appliances in unstaged buildings, else demote to open.

single-lab / plausibleTask-oriented (functional) graspingmanipulation

HELD modeled (distinct from learned-grasping-novel-objects, which covers stable grasping and stays viable): datasets and methods from multiple labs (TaskGrasp CMU, LERF-TOGO Berkeley) but real-hardware success on novel object-task pairs is modest and protocols differ per lab.

★ For the kids: Grab the knife by the handle, not the blade — robots are learning which part to hold for each job.

Frozen prediction: Through 2027, no third-party evaluation shows >75% task-appropriate grasp success across >=50 novel object-task pairs on real hardware.

early / contestedRobot tool use & improvisationmanipulation

HELD open: planning-based and LLM-prompted demos from strong labs (MIT RSS 2018; RoboTool 2023) but every result is a small curated suite; no reproduced autonomous tool improvisation outside training distribution.

★ For the kids: People grab a stick to reach a toy under the couch; robots are only beginning to figure that out.

Frozen prediction: By end 2028, no independently reproduced result shows autonomous unfamiliar-object tool use across >=10 tasks outside training distribution.

single-lab / plausibleTask and motion planning (TAMP)manipulation

HELD modeled: algorithmic core mature and reproduced across labs on shared formulations (PDDLStream), leading scaffold under learned skills — but perception-grounded TAMP in open worlds remains brittle.

★ For the kids: The robot plans many steps and checks each is physically possible before acting.

Frozen prediction: By end 2027, hybrid TAMP+learned-skill stacks beat end-to-end policies on >=10-step tasks in >=2 independent lab evaluations.

single-lab / plausibleLegged loco-manipulation (whole-body arm-on-legs)manipulation

MERGED duplicate proposals (consistent grades), HELD modeled (distinct from wheeled mobile-manipulation and from humanoid-whole-body-control locomotion): the DIRECTION is reproduced across >=3 independent groups (CMU unified policy CoRL 2022; ETH Pedipulate ICRA 2024; Stanford/Columbia UMI-on-Legs CoRL 2024) but no matched-task cross-lab reproduction exists; commercial Spot arm capability is model-based and operator-assisted. High demo-theater risk.

★ For the kids: Robot dogs use legs, back, and arm together to open doors and grab things — but each robot only knows its own trick.

Frozen prediction: Through 2028, no commercially deployed legged robot performs learned autonomous loco-manipulation as a paid unsupervised capability.

early / contestedDexterous hand durability (wear-item gap)manipulation

HELD open, strengthened by cross-domain evidence. The surgical field, which faces real regulators, treats precision end-effectors as consumables with firmware-capped service lives and prices them as recurring disposables — what an industry looks like when it has actually measured tool wear under duty cycle. Humanoid makers publish DoF counts and payload but zero audited MTBF for any five-finger hand, and the LEAP-class answer to fragility is cheap replaceability rather than measured life. If a hand is a wear item, dollars-per-productive-hour depends on a number nobody has published.

★ For the kids: Surgery robots admit their tools wear out and sell replacements. Robot-hand makers never say how long a hand lasts.

Technical: Shadow hand breakage under RL, DEX-EE robustness redesign, LEAP replaceability, plus dVRK/da Vinci instrument-life capping as the mature-industry contrast; links to humanoid-unit-economics (open).

Frozen prediction: Through end of 2028, no maker publishes third-party-audited MTBF >1,000 hours of continuous dexterous operation for a five-finger hand; the first credible durability disclosure arrives as a warranty or consumable SKU, not a paper.

single-lab / plausibleLatent-action pretraining from actionless videolearning

HELD modeled, evidence base narrowed. Genie is struck as a leg — it learns latent actions for generated 2D video environments, not robot policy pretraining, and citing it inflated a single-lab result into an apparently multi-lab one. GR00T N1's latent-action codes are a vendor technical report. That leaves LAPA (ICLR 2025) as the sole peer-reviewed leg: one lab, its own evaluation. Single-lab plus a real mechanism is the definition of modeled, so the grade survives while its justification does not. Cross-domain analogy is not cross-lab reproduction.

★ For the kids: One team taught a robot by inventing pretend action labels for ordinary videos. It worked for them; nobody else has repeated it on the same test.

Technical: Ye et al., LAPA, ICLR 2025 (KAIST/UW/NVIDIA); Bruce et al., Genie, ICML 2024 (STRUCK — 2D generated environments, not robot control); GR00T N1 (arXiv:2503.14734, vendor report, discounted).

Frozen prediction: Promotes to viable only when a second lab reproduces the LAPA gain on a shared physical bench against a same-compute no-video baseline.

early / contestedWorld models as policy evaluatorsembodied AI / evaluation

Demotion to open ACCEPTED (roster drift corrected; 2026-07-31 ruling stands). Builder-measured correlations only.

★ For the kids: A robot dreaming a chore is not doing it, and the dreamer grades its own dreams.

Technical: arXiv 2512.10675; RoboWorld arXiv 2607.01060.

Frozen prediction: By 2027-08-25 a second group reports Spearman >=0.9 between neural-simulator and distributed real-robot rankings over >=15 policies (P=0.4).

single-lab / plausibleReal-to-sim scene digitizationsimulation

HELD modeled: multiple labs with different pipelines (RialTo MIT RSS 2024, ~67pp robustness gain; ACDC Stanford CoRL 2024 zero-shot transfer), single-lab evidence each; contact fidelity for manipulation is the open question.

★ For the kids: Photograph your kitchen, turn it into a video game, let the robot practice a million times, then do it for real.

early / contestedSafety alignment and guardrails for VLA policiessafety

HELD open (policy-level counterpart of humanoid-safety-certification): SafeVLA and runtime-shield work is early single-lab; no standardized safety benchmark, nothing near certification. Gates every deployment bridge.

★ For the kids: The robot needs a do-not-do-that reflex that works even when its big brain is confused.

Frozen prediction: By 2027-12-31, no VLA-controlled manipulator is certified under an ISO 10218/TS 15066-derived standard for uncaged operation.

early / contestedWhole-body humanoid manipulation VLAslearning

HELD open (deliberately low): Helix and GR00T-on-humanoid are company/single-lab demonstrations with undisclosed take counts and no external evaluation — the showreel class this study refuses to promote. Viable locomotion substrate does not transfer viability to manipulation autonomy.

★ For the kids: Videos of humanoids doing kitchen chores are cool, but until outside testers check them they stay in the maybe pile.

Frozen prediction: Tied to EMB-P3; promotion requires third-party evaluation on unscripted task sets.

early / contestedRobot fleet cybersecurity (undisclosed-access gate)deployment

HELD open: concrete disclosed evidence (Unitree Go1 pre-installed remote tunnel; UniPwn wormable BLE root on Go2/G1/H1/B2, published after vendor non-response); no major humanoid vendor publishes third-party security audits. An unpriced, binding deployment gate.

★ For the kids: Some robots shipped with hidden doors that let strangers take control, and nobody is inspecting the locks.

Frozen prediction: Through 2027, no top-10 humanoid vendor publishes a third-party security audit, and at least one more remotely exploitable vulnerability in a shipping legged/humanoid robot is disclosed.

single-lab / plausibleDemonstration-data quality & curationlearning

HELD modeled (distinct from low-cost-teleop-data, which covers capture hardware): controlled studies (robomimic CoRL 2021; Belkhale NeurIPS 2023; Re-Mix CoRL 2024) show quality/mixture dominate volume — but the lineage is Stanford-heavy, no cross-lab standard metric exists, and flagship VLAs still headline raw hours.

★ For the kids: A few really good practice examples beat piles of sloppy ones — but robot companies still brag about pile size.

Frozen prediction: By end 2027, a flagship VLA release publishes a curation ablation where a <=50%-hours quality-filtered subset matches the full mixture on real-robot eval.

early / contestedCapacitive taxel fingertips (the shipped modality is contested)sensing & tactile

Held open; the 'shipped modality is capacitive' sub-claim retired (blog inference only). No humanoid vendor publishes a third-party-measured fingertip datasheet.

★ For the kids: Companies say their fingers feel a paperclip; no outside lab has measured it.

Technical: Sharpa Wave (camera-based, press); Figure 03 '~3 g' with no modality; Sanctuary Feb-2025 no modality.

Frozen prediction: By 2027-09-05 no listed humanoid vendor publishes a third-party-measured fingertip datasheet with drift or cycle life. p=0.85.

early / contestedPolicy failure detection & runtime monitoringsafety

HELD open, and MERGED with the duplicate slug policy-runtime-failure-detection, which a second lens returned at modeled for the same subject. Merging removes a double-count and resolves to the stricter grade. The strongest single result is now FIPER (NeurIPS 2025): a random-network-distillation OOD score on the observation embedding plus action-chunk entropy, conformally calibrated on success-only trajectories, on five tasks with diffusion and flow policies — genuinely peer-reviewed, with a second group reporting the same recipe class. That is a demonstrated mechanism, not a reproduced capabi…

★ For the kids: Robots are just starting to notice when they are about to mess up — but nobody has published how often the alarm is wrong, so we cannot trust it yet.

Technical: Römer, Kobras, Worbis & Schoellig (TU Munich), FIPER, NeurIPS 2025 (arXiv:2510.09459); 'Can We Detect Failures Without Failure Data?', arXiv:2503.08558 (2025); Ren et al., KnowNo, CoRL 2023 (arXiv:2307.01928), conformal help-seeking with coverage guarantees; Agia et al., Sentinel, CoRL 2024 (arXiv:2410.04640), consistency + progress detectors; Hoque et al., ThriftyDAgger, CoRL 2021; Liu et al., Sirius, UT Austin 2023-2024.

Frozen prediction: EMB-P11 (held): through 2028-06-30, no humanoid or manipulation vendor publishes a runtime failure-detector specification with BOTH false-negative and false-alarm rates measured on a third-party task suite; falsified — happily — if one does.

single-lab / plausibleUnderpowered manipulation evaluation (the n=10 problem)evaluation-integrity

HELD modeled, and MERGED with the duplicate slug real-robot-evaluation-statistics, which a second lens returned for the identical finding off the identical sources. The diagnosis is now multi-source: a 13-paper survey of real-robot VLA evaluations finds modal per-condition trial counts of 10-20 with none reporting confidence intervals or paired tests; an independent five-benchmark audit (LIBERO, CALVIN, SimplerEnv, RoboCasa, RoboTwin 2.0) finds shortcut solvability, absent significance testing, creeping overfitting and data-source dependence. At n=20 a Wilson 95% interval on an observed 80% s…

★ For the kids: Most robot papers try something twenty times and report a percentage. That is like calling a coin unfair after ten flips — and the careful way to count actually needs FEWER tries.

Technical: PhAIL, arXiv:2605.29710 (2026, 13-paper survey, time-to-success CDFs with bootstrap CIs and macro-averaged KS); Jiang, Tan, Wheeler, Sun, Ayalew & Walter, arXiv:2606.04233 (2026, five-benchmark audit); Snyder, Badithela, Matni, Pappas, Majumdar, Itkina & Nishimura, arXiv:2603.13616 (2026, safe anytime-valid inference); Snyder et al., STEP, arXiv:2503.10966 (2025); Kress-Gazit et al., arXiv:2409.09491 (2024, reporting standard); SureSim, arXiv:2510.04354; Agarwal et al., NeurIPS 2021 (statistical precipice); TRI LBM, Science Robotics 2026 (doi 10.1126/scirobotics.aea6201).

Frozen prediction: Through end of 2027, the median per-condition real-robot trial count in accepted RSS/CoRL/ICRA manipulation papers stays <=20-30, and fewer than 20-25% report confidence intervals or paired significance tests.

single-lab / plausibleAutomated demonstration synthesis (language-driven quality-diversity, simulation only)manipulation & dexterity / data generation

Held modeled. Garrabé et al. (arXiv 2608.30983, verified) is Genesis-simulator only; inherits MimicGen-class sim-only status.

★ For the kids: A robot invents many ways to do a job from a sentence, but only inside a computer game.

Technical: 4 manipulation tasks; archive coverage metric, no hardware success.

Frozen prediction: By 2027-06-30 a language-driven QD archive transfers to a real manipulator with >= 50% success on >= 2 tasks in a peer-reviewed venue. p=0.35.

reproduced / crediblePneumatic artificial muscles (McKibben class) — the reproduced muscle the humanoids rejectedactuation

HELD viable. Survives the independence check: independently characterized and modeled at two institutions (Chou & Hannaford, Univ. of Washington 1996; Tondu & Lopez, INSA Toulouse 2000), used to build complete walking/running bipeds by two further independent groups (VUB 'Lucy', Autonomous Robots 2005; Osaka Univ. antagonistic multi-modal biped, RAS 2008), and sold as a catalog part for two decades. Remove any one lineage and the grade survives. Scope binding: the graded capability is 'compliant, high force-to-weight joint actuation from a fixed pressure supply' and explicitly does NOT extend…

★ For the kids: Robot muscles made from rubber tubes that squeeze when you pump air in. They really work — people built walking robots with them twenty years ago — but the robot must carry a noisy air pump, so newer robots use motors.

Technical: Chou & Hannaford, IEEE T-RA 12(1):90-102 (1996); Tondu & Lopez, IEEE Control Systems Magazine 20(2) (2000); Verrelst et al., Autonomous Robots 18(2):201-213 (2005, pleated PAM biped Lucy); Hosoda, Takuma, Nakamoto & Hayashi, RAS 56:46-53 (2008); Festo DMSP commercial since ~1999. Disqualifier is compressor mass and compression-cycle efficiency, not muscle performance.

Frozen prediction: Through end of 2029, no untethered commercial humanoid ships pneumatic artificial muscles as primary joint drive; PAM use stays confined to tethered rigs, rehab/exoskeleton devices and soft grippers.

reproduced / credibleOpen-source torque-control electronics (ODRI / ODrive / moteus / SimpleFOC)actuation

HELD viable. Closed-loop field-oriented current/torque control in a robot joint is demonstrated and reproduced in the strong sense: the ODRI micro-driver and actuator module (RA-L 2020) was published as open hardware and rebuilt across MPI-IS, NYU, LAAS and others into the Solo/Bolt/TriFinger family; SimpleFOC is peer-reviewed (JOSS 2022) with a large independent user base; ODrive and mjbots moteus are shipping commercial-open boards. Note the distinction this lens enforces elsewhere: here independent institutions published their own hardware results, which is reproduction, not mere tool adop…

★ For the kids: The little computer that tells a robot motor how hard to push. You can buy one cheaply or download the design free, so it stopped being the hard part.

Technical: Grimminger et al., IEEE RA-L 5(2):3650-3657 (2020); Skuric, Bank, Unger, Williams & González-Reyes, JOSS 7(74):4232 (2022); ODrive and mjbots moteus commercial-open boards. The control law (Clarke/Park + current loop) is textbook — the reproduced artifact is the packaging.

Frozen prediction: By end of 2028, at least one commercially sold legged robot or dexterous hand under $20k is documented by teardown, BOM analysis or vendor disclosure to ship ODRI/ODrive/moteus/SimpleFOC-lineage joint control; and in every published humanoid joint-module BOM through 2028 the controller board stays under 10% of module cost with reducer plus magnets above 50%.

early / contestedAll-solid-state cells for robots (split from the silicon-anode grade)power

HELD open. Split endorsed: the field's own benchmarking literature is the indictment — Randau et al. (Nature Energy 2020) had to construct a bare-minimum reference cell because published solid-state results are not comparable across labs, and a 2024 Nature Energy follow-up finds all-solid-state cell performance is not reproducible cell-to-cell within a controlled round-robin. Janek & Zeier (2023) document the slipping timelines. Zero published robot integrations as of 2026-07. No robotics solid-state claim may borrow silicon-anode's third-party verification.

★ For the kids: A battery with no liquid inside. Scientists still cannot measure two of them the same way or build two that behave alike — so it is nowhere near being inside a robot.

Technical: Randau et al., Nature Energy 5:259-270 (2020); 'Benchmarking the reproducibility of all-solid-state battery cell performance', Nature Energy (2024); Janek & Zeier, Nature Energy (2023).

Frozen prediction: Through end of 2028, no commercially sold humanoid, quadruped or mobile manipulator ships an all-solid-state main pack; any robotics solid-state claim in this window rests on vendor attestation with no third-party cell test published.

early / contestedPrecision-reducer price collapse (the uncited number holding up a viable grade)actuation

HELD open, and the split is endorsed as exactly the right correction: an ECONOMIC claim was riding inside a CAPABILITY grade. There is no public teardown-verified, unit-priced, like-for-like comparison anywhere — matched ratio, matched rated torque, matched lost-motion and torsional stiffness. Trade press and vendor list prices are attestations, and list price is not transaction price at humanoid volumes. Highest-leverage unverified number in the actuation lens, because the reducer is the largest joint-module BOM line and the humanoid cost-collapse thesis leans on it.

★ For the kids: People keep saying robot gearboxes just got much cheaper. Maybe! But nobody has put two matching gearboxes side by side and shown the receipts.

Technical: Parent-domain evidence (Musser US 2,906,143, 1959; Ghorbel, Gandhi & Alpeter, ASME JMD 2001) supports the mechanism, not the price. Required evidence: matched-spec unit pricing in >=2 independent teardowns or BOM analyses.

Frozen prediction: By end of 2027, at least two independent teardowns or BOM analyses of shipping humanoids document a non-incumbent harmonic or cycloidal reducer under 60% of Harmonic Drive SE / Nabtesco list price at matched ratio and torque; absent that, the price-collapse claim stays unsupported.

early / contestedHumanoid DC bus voltage architecture (the 60 V safety ceiling)power

HELD open. The physics and the standards are settled (IEC 60204-1 / IEC 61140 treat <=60 VDC as safety-extra-low-voltage, keeping a human-adjacent robot out of the high-voltage certification path EV traction occupies); the ROBOT data is absent. Every public humanoid bus figure is vendor attestation, and no peer-reviewed or teardown-verified study of humanoid bus architecture exists. Graded on evidence, not on doubt about the physics.

★ For the kids: Robots run on battery voltage low enough that touching it cannot hurt you. That safety line secretly decides a lot about how the robot is built — and nobody outside the companies has checked it.

Technical: IEC 60204-1:2016 and IEC 61140 (SELV/PELV, 60 VDC touch-safety limit). All humanoid bus figures currently in circulation are vendor-stated; no independent teardown measurement published as of 2026-07.

Frozen prediction: Through end of 2028, every publicly disclosed or torn-down commercial humanoid main bus stays at or below 60 VDC nominal; a shipping humanoid with a >100 VDC traction-style bus falsifies the safety-ceiling read.

early / contestedThe standing-still power tax (holding torque at zero mechanical work)actuation & power

Held open. First vendor number: Walker S2 ~2 h walking / ~4 h standing implies standing ~ half of walking power, but that sum includes hotel load and no source separates them. No third-party shunt measurement.

★ For the kids: Standing still burns about half of walking, because the motors keep pushing and the computer never sleeps.

Technical: UBTech Walker S2 press (July 2025). Separation protocol in study hotel-load-audit-protocol (B-A = holding tax).

Frozen prediction: By 2027-12-31 no third-party measurement shows a commercial humanoid's standing draw below 25% of its 1 m/s walking draw. p=0.70.

early / contestedActuator measurement standards (the missing actuation referee)actuation

HELD open. Robot actuator claims are unfalsifiable by construction: peak torque density is published with no stated ambient, no duty cycle, no thermal soak protocol and no backdrivability or efficiency method, so the '30+ Nm/kg' and '100 Nm/kg' figures circulating in 2026 are not commensurable. A first structured protocol suite exists (Actuators 15(6):301, 2026 — eleven protocols with defined indices) but it is one group's proposal with zero adoption, and Vanderborght et al. made the same complaint thirteen years earlier without effect. Structural blocker on actuator-thermal-limits, axial-flu…

★ For the kids: Every robot company brags about how strong its motors are, but they all measure it differently and none says for how long — so nobody can be proven wrong.

Technical: 'Toward Reproducible Actuator Characterization', Actuators (MDPI) 15(6):301 (2026), doi:10.3390/act15060301 — single-group proposal, no adoption; Vanderborght et al., RAS 61:1601-1614 (2013). No IEEE/ISO/IEC continuous-torque actuator benchmark exists.

Frozen prediction: Through end of 2028, no IEEE/ISO/IEC-endorsed actuator benchmark for CONTINUOUS torque density under stated ambient and duty cycle is published, and no two humanoid vendors report joint torque density under a shared, stated measurement protocol.

reproduced / credibleSafety-rated contact & proximity skin (the large-area skin already deployed)sensing & tactile

Viable HELD: deployed commercial products from at least two vendors with certified performance levels (AIRSKIN pressure-pad PL e / Cat 3 / SIL 3; Bosch APAS capacitive PL d). Precision fix accepted: the highest-certified skin is pressure-based, so the slug name must not imply capacitive is the top tier.

★ For the kids: Robots next to people wear a skin that makes them stop; the touch-pad kind is rated safer than the no-touch kind.

Technical: Vendor certification statements (BG type examination for APAS; TÜV for AIRSKIN). A capacitive skin at PL e is the watch-item.

Frozen prediction: Through 2028, every humanoid or cobot deployment passing a documented close-proximity safety assessment does so on joint-torque sensing and/or safety-rated skin/proximity sensing as the certified stopping channel — never on high-resolution tactile skin.

single-lab / plausibleFiber-optic / FBG force and tactile sensingsensing

HELD modeled. The only force-sensing modality with FDA-cleared, at-scale clinical deployment and independent peer-reviewed outcome studies (da Vinci 5 Force Feedback; two 2025 pre-clinical studies reporting up to ~43% reduction in force on tissue; a 444-procedure real-world thoracic analysis across 73 surgeons). Held below viable because the deployed leg is a narrow tip-force channel inside one product family, and the robot-hand designs are one-off with no shared benchmark. EMI immunity, sterilizability and heat tolerance are the moat, not resolution.

★ For the kids: Surgery robots feel how hard they pull by shining laser light down hair-thin glass threads — the most trusted robot touch sensor in the world is not on a robot hand.

Technical: Intuitive da Vinci 5 Force Feedback (2024); Surgical Endoscopy 2025 studies; one-year real-world FFB thoracic analysis (2025); Yi et al., IEEE T-RO 2026 (SNU); 'Fibertouch', MSSP 2025; clip-on modular FBG attachment, Sensors 2025, 25(19):5943.

Frozen prediction: Through 2028, fiber-optic/FBG force sensing stays confined to surgical, nuclear and other extreme-environment robotics; no general-purpose commercial robot hand or humanoid ships FBG tactile fingertips as standard.

single-lab / plausibleBarometric MEMS taxel sensing (reproduced, never selected)sensing

HELD modeled. The transduction principle is independently rebuilt across several groups (Harvard TakkTile; BaroTac three-axis with slip detection; Toronto barometric slip-detection TCN; soft barometric contact-localization sensors) and was retrofitted onto a commercial gripper thirteen years ago — but the MANIPULATION capability claim is not reproduced, and the market's revealed verdict so far is no. Kept because it prevents the field's cheapest solved transducer being rediscovered as news.

★ For the kids: Cheap air-pressure chips baked into rubber make a great sense of touch that almost nobody actually sells.

Technical: Tenzer, Jentoft & Howe, IEEE RAM 2014 (TakkTile); 'BaroTac', Sensors 2022; 'Learning to Detect Slip with Barometric Tactile Sensors and a TCN', arXiv:2202.09549 (2022, Toronto); soft barometric contact/slip sensor, IEEE 2022.

Frozen prediction: Through 2028, barometric MEMS taxels remain absent from every commercially shipped humanoid hand and every mainstream gripper fingertip product.

early / contestedTactile sensor wear, drift & replaceability (the tactile wear-item gap)sensing

HELD open. Soft sensing interfaces are consumables — gels abrade, magnetized elastomers demagnetize, capacitive dielectrics creep — yet no vendor publishes a cycles-to-replacement figure or a calibration-drift curve. The single serious engineering answer is AnySkin, which decouples the sensing interface from the electronics and demonstrates zero-shot policy transfer onto a REPLACEMENT instance without ever publishing how often replacement is needed. The 2025-26 review literature independently names in-situ drift monitoring, wear-induced drift patterns and replacement time as the missing deplo…

★ For the kids: Robot skin wears out like the sole of a shoe, and nobody will tell you how many touches it lasts.

Technical: Bhirangi, Pattabiraman, Erciyes, Cao, Hellebrekers & Pinto, AnySkin, arXiv:2409.08276 (2024); Lepora, 'Tactile Robotics: Past and Future', arXiv:2512.01106 / IJRR 2026; 2025-26 tactile-sensor reviews on lifecycle and drift metrics; TakkTile's unreplicated 'survives a baseball bat' claim (IEEE RAM 2014).

Frozen prediction: Through 2028, no commercial tactile fingertip or skin vendor publishes a duty-cycle-referenced lifetime spec or a calibration-drift curve; tactile cost per productive hour stays uncomputable from public data.

early / contestedTomographic (EIT / acoustic) skins — the wire-count end-runsensing

HELD open, and honestly so. Several independent groups have built EIT skins and one Cambridge skin combines electrical-impedance with acoustic tomography, but reconstruction is an ill-posed inverse problem, resolution is coarse, drift is severe, and no robot performs any reproduced task because of a tomographic skin. The motivation is real — wiring scales with perimeter rather than taxel count — which is why it keeps being rebuilt.

★ For the kids: Instead of a wire for every touch spot, these skins listen at the edges and work out where you were poked — clever, but blurry.

Technical: Hardman, George Thuruthel & Iida (Cambridge), 2022; 'Artificial skin through super-sensing method...', Scientific Reports 2019; 'Variable sensitivity multimaterial robotic e-skin...', Scientific Reports 2023; 'Data-efficient Tactile Sensing with EIT', arXiv:2411.12658 (2024); 'Flexible EIT for tactile interfaces', arXiv:2411.13306 (2024).

Frozen prediction: Through 2028, no shipping robot uses a tomographic skin as its primary contact-sensing channel, and no tomographic skin is independently reproduced outside its origin lab on a task a taxel array cannot do.

early / contestedFailure detection and autonomous recovery in manipulationmanipulation

HELD open — the capability that separates a demo from a deployment, and it is barely measured. Three distinct approaches exist in single-lab form (post-hoc language explanation, a fine-tuned failure-reasoning VLM, runtime LLM-embedding anomaly detection with reactive fallback), plus 2026 activation-probe escalation work that is simulation-only and a single-author preprint. Most results measure DETECTION, not RECOVERY; there is no shared failure taxonomy, no shared benchmark and no real-hardware recovery statistics at meaningful trial counts. This is the honest reason autonomous manipulation d…

★ For the kids: Robots are getting okay at noticing when they mess up. Getting themselves un-messed-up without a human stepping in is still basically unsolved.

Technical: Liu, Bahety & Song, REFLECT, CoRL 2023 (arXiv:2306.15724); Duan et al., AHA, ICLR 2025 (arXiv:2410.00371); Sinha, Elhafsi, Agia, Foutter, Schmerling & Pavone, RSS 2024 (arXiv:2407.08735); AEGIS, arXiv:2606.06660 (2026, sim-only single-author preprint, cited as early evidence only).

Frozen prediction: Through end of 2028, no third-party evaluation reports >=70% autonomous recovery from >=3 distinct unscripted failure modes over >=100 real-hardware trials on any manipulation platform.

early / contestedLong-horizon task chaining (the multiplicative success wall) — MERGED from two lensesmanipulation & dexterity / evaluation

SKEPTIC: two lenses extended the same slug with different referees; merged. BEHAVIOR-1K NeurIPS 2025 winner scored 26% q-score on 50 simulated household tasks; LongBench (arXiv 2604.16788, 2026) evaluates six policies over 1,000+ real episodes, single lab. The wall is now a number in sim and has a first real referee.

★ For the kids: Ten steps in a row is much harder than one; the best contest team got a quarter of the way.

Technical: Robot Learning Collective BEHAVIOR-1K write-up; Chen et al. arXiv 2604.16788.

Frozen prediction: The top BEHAVIOR-1K challenge q-score in the 2026 edition remains below 50%. p=0.65.

reproduced / credibleStandardized physical benchmark artifacts (YCB objects, NIST task boards)evaluation

HELD viable — one of the few things here that clears the reproduction bar without argument. The YCB Object and Model Set is physically distributed and used across hundreds of labs on multiple continents; the NIST small-parts assembly task boards codify threading, snap-fit, gear-mesh, connector and belt tasks with published protocols and are replicable from published drawings. The graded capability is standardized physical benchmarking itself, and it is demonstrated and reproduced. The scandal is adoption: the artifacts are cheap and optional, so most headline results are still reported on bes…

★ For the kids: There is a standard box of test objects and a standard test board any lab can buy, so results could be compared fairly. Most teams just use their own random stuff instead.

Technical: Calli, Walsman, Singh, Srinivasa, Abbeel & Dollar, IEEE RAM 2015 (YCB); Kimble, Van Wyk, Falco, Messina, Sun, Shibata, Uemura & Yokokohji, IEEE RA-L 5(2):883-889, 2020, doi:10.1109/LRA.2020.2965869 (NIST task boards).

Frozen prediction: Through end of 2027, fewer than 25% of accepted RSS/CoRL real-robot manipulation papers report results on a shared standardized physical artifact at >=50 trials per condition.

single-lab / plausibleTransparent and specular object manipulation (the depth-sensor blind spot)manipulation

HELD modeled. Commodity depth cameras return garbage on glass, clear plastic, chrome and liquids, silently removing a large fraction of kitchen, lab and retail objects from every grasping result reported on opaque YCB items. Four independent lineages attack it and a shared real-world benchmark (TransCG, 57,715 RGB-D images, 51 categories) makes the depth-completion sub-problem genuinely cross-lab — but real grasp success on NOVEL transparent objects under uncontrolled lighting is reported per-lab on per-lab object sets, and no result approaches opaque-object baselines.

★ For the kids: Robot depth cameras basically cannot see glass or shiny metal, so a robot that grabs a mug fine may swipe straight through a drinking glass.

Technical: Sajjan et al., ClearGrasp, ICRA 2020; Ichnowski, Avigal, Kerr & Goldberg, Dex-NeRF, CoRL 2021; Dai et al., TransCG, IEEE RA-L 2022 (arXiv:2202.08471); SR3D arXiv:2505.24305; DCIRNet arXiv:2506.09491; FuseGrasp arXiv:2502.20037.

Frozen prediction: Through end of 2027, no lab independent of a method's originating group reproduces >=90% grasp success on >=30 novel transparent or specular objects under uncontrolled lighting; where both are reported, transparent-object success stays >=15pp below the same system's opaque baseline.

single-lab / plausibleInteractive perception and mechanical search (acting to see)manipulation

HELD modeled, lineage concentration flagged as binding. The paradigm is framed (Bohg et al., T-RO 2017) and mechanical search formalized (Danielczuk et al., ICRA 2019), but the real-hardware evidence sits overwhelmingly in one lab lineage, evaluation is sim-heavy, and object sets are per-paper. This is the mechanism the deployment story needs — bins, shelves, fridges and drawers are occlusion problems before they are dexterity problems — and it is under-invested relative to in-hand reorientation showreels.

★ For the kids: To find something buried in a messy drawer you have to move the other stuff first. Robots are only starting to learn to dig around.

Technical: Bohg et al., IEEE T-RO 2017; Danielczuk et al., ICRA 2019 (Berkeley AUTOLAB); X-Ray and shelf-search follow-ups 2020-2021.

Frozen prediction: By end of 2027, no group outside the originating lineage reports >80% target-retrieval success under >=5-object occlusion on a shared standardized object set; absent that, hold modeled.

single-lab / plausibleAutonomous surgical manipulation (precision dexterity with a real referee)manipulation

HELD modeled, with the split kept clean. INFRASTRUCTURE is reproduced — the da Vinci Research Kit is deployed at dozens of institutions, one of the most reproduced manipulation platforms in existence. AUTONOMY is single-lineage: supervised autonomous suturing, in-vivo intestinal anastomosis and step-level autonomous cholecystectomy all trace to the JHU / Children's National lineage, ex-vivo or animal, never a human patient outside protocol. Internal consistency check retained: 100% on n=8 has a Wilson 95% lower bound near 68% — the same underpowered pattern flagged elsewhere, except the surgi…

★ For the kids: A few research systems have done one whole step of an operation by themselves on animal or lab tissue — genuinely new, but one research group, not the whole field.

Technical: Kazanzides et al., dVRK, ICRA 2014 (JHU); Shademan et al., STAR, Science Translational Medicine 2016; Saeidi et al., Science Robotics 2022; Kim et al., SRT-H, Science Robotics 2025, doi:10.1126/scirobotics.adt5254 / arXiv:2505.10251.

Frozen prediction: Through end of 2028, no autonomous tissue-manipulation subtask is performed on a human patient outside a research protocol, AND no group outside the JHU/Children's National lineage reproduces step-level surgical autonomy at matched success on a shared ex-vivo protocol.

reproduced / credibleMicromanipulation and sub-micron precision manipulationmanipulation

HELD viable. Reproduced dexterity the humanoid narrative ignores because it does not look like a hand: piezo-driven micromanipulators with microscope visual servoing and micro/nano-newton force control routinely perform cell injection, patch-clamp and micro-assembly at sub-micron closed-loop precision; the automated whole-cell patch-clamp robot was open-sourced and independently adopted, and commercial ICSI micromanipulation systems have shipped for decades. Scope stated honestly: fixed base, microscope in the loop, structured 2.5D workspace, known specimen class. Precision manipulation was s…

★ For the kids: There are robots that can poke a single living cell without popping it, over and over, in labs all over the world. They are not hands.

Technical: Zhang, Wang, Liu, Dai & Sun, Annu. Rev. Control Robot. Auton. Syst. 2:181-203 (2019), doi:10.1146/annurev-control-053018-023755; Kodandaramaiah, Franzesi, Chow, Boyden & Forest, Nature Methods 2012; commercial ICSI micromanipulators (Eppendorf, Narishige).

Frozen prediction: Through end of 2028, no general-purpose robot arm-plus-hand demonstrates reproduced sub-10-micron closed-loop placement on novel objects.

early / contestedWhole-body large-object manipulation (arms, chest and non-prehensile contact)manipulation

HELD open. Humans carry a laundry basket with forearms and chest, not fingertips. The one serious research line (Punyo-1, TRI — compliant cut-resistant arm coverings, soft-bubble paws, pressure-sensing chest) is essentially one organization, with no cross-lab reproduction, no shared task specification for 'carry this awkward 12 kg object', and zero deployment. Distinct from extrinsic dexterity in that the ROBOT'S OWN BODY is the manipulator surface, which makes distributed torso/arm skin the sensing bottleneck.

★ For the kids: You carry a big box with your arms and chest, not your fingertips. The one robot built around this is a single research lab's prototype.

Technical: Goncalves, Kuppuswamy, Alspach et al., Punyo-1, TRI, arXiv:2111.09354; Alspach et al., Soft-bubble, RoboSoft 2019; links to large-area-electronic-skin (open) as the enabling sensing gap.

Frozen prediction: Through end of 2028, no commercially deployed robot performs autonomous whole-body non-prehensile manipulation of >10 kg objects using body-contact surfaces as a paid capability.

single-lab / plausibleAction tokenization & action representationlearning

HELD modeled. Multi-lab agreement that per-dimension uniform binning is a bottleneck at high control frequency earns modeled; which replacement wins (DCT-compressed tokens, flow-matching continuous heads, parallel-decoded L1 regression) is UNREPRODUCED — no shared-bench, matched-backbone, matched-data head-to-head exists anywhere.

★ For the kids: A robot's brain has to turn its thoughts into numbers for its arms. How it packs those numbers decides how fast and smoothly it moves — and nobody has fairly raced the packing methods.

Technical: Brohan et al. RT-1/RT-2 (2022/2023); Kim et al., OpenVLA, CoRL 2024 (arXiv:2406.09246); Pertsch et al., FAST, 2025 (arXiv:2501.09747); Black et al., pi0 (arXiv:2410.24164); Kim et al., OpenVLA-OFT, RSS 2025 (arXiv:2502.19645).

Frozen prediction: EMB-P7: by 2027-12-31 no third-party study on a shared physical bench, at matched backbone and training data, shows any single action-representation scheme beating the others by >=10pp absolute success.

single-lab / plausibleAsynchronous action-chunk execution (the inference-latency seam)embodied-ai

HELD modeled. VLAs emit action chunks at 1-10 Hz while joints need 50-1000 Hz, so what happens BETWEEN chunks determines real-world smoothness and contact safety. Four groups attack it independently, which earns modeled — but the strongest asynchronous result is a vendor self-evaluation, and no cross-lab study quantifies the success cost of a given inference delay.

★ For the kids: A robot's brain thinks in slow chunks but its arms need instructions constantly. What it does in the gaps is where a lot of its clumsiness hides.

Technical: Zhao et al., ACT/ALOHA temporal ensembling, RSS 2023 (arXiv:2304.13705); Liu et al., Bidirectional Decoding, 2024 (arXiv:2408.17355); Kim et al., OpenVLA-OFT, RSS 2025; Black et al., Real-Time Execution of Action Chunking Flow Policies, 2025 (arXiv:2506.07339, vendor self-evaluated, discounted).

Frozen prediction: EMB-P8: by 2027-12-31 no independent lab publishes a controlled real-robot study isolating inference-latency cost (task success versus injected delay) for an open-weight VLA at >=2 control rates.

single-lab / plausibleLLM/VLM high-level task planning & affordance groundinglearning

HELD modeled. Reproduced as a PATTERN across at least five independent groups and now shipped inside a commercial VLA's high-level subtask policy — but every evaluation is self-designed on the authors' own task list, reported failures are overwhelmingly LOW-LEVEL execution, and the planner's marginal contribution over a flat language-conditioned policy is never isolated. Simulation planning suites are excluded as evidence.

★ For the kids: Robots can now use a chatbot to write their to-do list. Writing the list turns out to be the easy part.

Technical: Ahn et al., SayCan, CoRL 2022 (arXiv:2204.01691); Huang et al., Inner Monologue, CoRL 2022; Liang et al., Code as Policies, ICRA 2023 (arXiv:2209.07753); Singh et al., ProgPrompt, ICRA 2023; Huang et al., VoxPoser, CoRL 2023 (arXiv:2307.05973); pi0.5 (arXiv:2504.16054). ALFRED/BEHAVIOR/VirtualHome excluded under the sim-inflation rule.

Frozen prediction: EMB-P9: through 2028-06-30, no third-party evaluation shows an LLM/VLM high-level planner adding >=20pp end-to-end real-robot success over a flat language-conditioned policy at matched low-level controller across >=10 unseen multi-step tasks.

single-lab / plausibleVLM spatial grounding as the action interface (pointing / keypoints)learning

HELD modeled. Four-plus independent groups demonstrate the interface (mark-based prompting, relational keypoint constraints, affordance points, pointing supervision at scale), and it is the cheapest known route to open-vocabulary manipulation without collecting robot data. Held below viable because each result is self-evaluated on its own object set, pointing accuracy is a PERCEPTION metric that must never be read as manipulation capability, and depth/occlusion failure rates go unpublished.

★ For the kids: Instead of moving the arm itself, the AI just points at where to grab. Pointing well is not the same as grabbing well.

Technical: Liu et al., MOKA, RSS 2024 (arXiv:2403.03174); Huang et al., ReKep, CoRL 2024 (arXiv:2409.01652); Yuan et al., RoboPoint, CoRL 2024 (arXiv:2406.10721); Deitke et al., Molmo/PixMo, 2024 (arXiv:2409.17146); Gemini Robotics-ER (vendor report, discounted).

Frozen prediction: EMB-P10: by 2028-01-01, no shared third-party physical benchmark reports keypoint/pointing-grounded pipelines beating end-to-end VLAs by >=15pp on novel-object pick-and-place.

early / contestedAutonomous data flywheel & fleet self-improvementlearning

HELD open. The load-bearing assumption under every humanoid business plan is demonstrated only in pieces — self-generated-data iteration, LLM-orchestrated multi-robot collection, autonomous instruction-following improvement, zero-shot deployment from cheap in-the-wild capture. Each is single-lab, gains are in-lab and short-horizon, human labor is typically displaced into resets and supervision rather than removed, and no result shows sustained multi-week compounding on a fixed external eval.

★ For the kids: The plan is that robots practice by themselves and slowly get better. So far the robots still need people nearby to reset things and fix mistakes.

Technical: Bousmalis et al., RoboCat, TMLR 2024 (arXiv:2306.11706); Ahn et al., AutoRT, 2024 (arXiv:2401.12963); Zhou et al., SOAR, CoRL 2024 (arXiv:2407.20635); Etukuru et al., Robot Utility Models, 2024 (arXiv:2409.05865); Zhou et al., AutoEval 2025.

Frozen prediction: EMB-P12: by 2028-01-01, no published result shows an autonomous deployment flywheel delivering >=2x sample efficiency or >=15pp absolute success over a fixed human-teleop dataset baseline, replicated or audited outside the originating lab.

reproduced / credibleBenchmark memorization collapse in VLA evaluationevaluation-integrity

HELD viable, with an explicit guard on how the grade may be used. The maturity grades a NEGATIVE finding's reproducedness, not a capability: four independent groups report the same collapse of headline VLA scores under task-preserving perturbation (>90% to 0.0%; 95% to below 30%; >90% average failure across three physical variations). Two distinct benchmark substrates are involved (LIBERO-family perturbations plus the RLBench-based Colosseum lineage), so this is not four re-skins of one artifact. GUARD: a viable-graded negative finding may never be cited as evidence that any capability exists…

★ For the kids: Some robots ace the test only because they memorized the answers — move one object and they score zero. Four different teams found the same thing.

Technical: Zhou, Xu, Tie, Chen, Zhang, Chu, Zhou & Sun, LIBERO-PRO, arXiv:2510.03827 (rev. 2026-05-25); Fei, Wang, Shi et al., LIBERO-Plus, arXiv:2510.13626, CVPR 2026; Liu, Ruan, Long et al., Eva-VLA, arXiv:2509.18953 (rev. 2026-03-15); LIBERO-X litmus, 2026-02; Pumacay et al., THE COLOSSEUM, RSS 2024.

Frozen prediction: SKP-P1: through 2027-12-31, no flagship VLA release headlines an externally-run perturbation-robust score alongside its standard-benchmark number.

reproduced / credibleLanguage-conditioning shortcut in VLA policiesevaluation-integrity

HELD viable on the reproducedness of the failure finding, with the same guard: it caps grades, it never establishes a capability. Multiple independent groups, on different model families, find VLA policies execute visually plausible trajectories largely without using the instruction — models 'largely insensitive to language variations' and in places ignoring instructions entirely, and manipulation-primitive execution substantially outperforming instruction-conditioned success once object-location correlations are broken. The 'language-conditioned' framing is doing unearned work.

★ For the kids: Robots that look like they are following your words are often just repeating a move they already knew.

Technical: Fei et al., LIBERO-Plus, arXiv:2510.13626 (CVPR 2026); Emukpere, Deffayet & Renders, arXiv:2602.24143 (2026-02-27); RoboSemanticBench, arXiv:2606.02277 (2026); LangForce, arXiv:2601.15197 (2026).

Frozen prediction: SKP-P5: promotion of any VLA component to viable requires a decomposed metric separating primitive execution from instruction-conditioned success, reported by a non-author party; on current evidence no such promotion occurs through 2027-12-31.

single-lab / plausibleSimulation as a proxy for real evaluation (rank-preservation gap)evaluation

HELD modeled. SIMPLER-class sim/real correlation is the field's licence to skip hardware, but independent work finds it overestimates policies because of few environments and biased scenario selection. Statistically principled hybrids exist and are quantified — prediction-powered inference combining large sim with small real testing saved 20-25% of hardware effort — which bounds what today's physics sims are worth as evidence. What is missing is demonstrated rank preservation across policy FAMILIES, so sim-only numbers stay ungradeable.

★ For the kids: Practising in a video game tells you something about the real world, but not enough to declare a winner.

Technical: Badithela, Snyder, Zha, Mikhail, O'Kelly & Dixit, SureSim, arXiv:2510.04354; Chen et al., RoboDojo, arXiv:2607.04434 (2026-07-05); SIMPLER, CoRL 2024.

Frozen prediction: SKP-P4: by 2028-01-01, no published study shows simulation-derived policy rankings matching real-robot rankings at Spearman >=0.8 across >=3 distinct policy families on a shared task set.

early / contestedReproducibility reporting standards in robotics (the provenance gap)embodied AI / evaluation

Held open. vla-eval is the first harness re-running six public codebases; one harness, sim only; training-compute disclosure rare.

★ For the kids: When others re-ran the robot brains, scores mostly matched, but tiny setup mistakes could halve a score.

Technical: arXiv 2603.13966; OpenVLA 21,500 A100-h disclosed.

Frozen prediction: Through 2027, no top-tier robot-learning venue requires a machine-checkable reproduction manifest (action space, frame conventions, preprocessing, seed/trial protocol) as a condition of publishing real-robot success rates.

single-lab / plausibleReplicable low-cost evaluation cells (the copy-this-bench standard)embodied-ai

HELD modeled, with the self-validation caveat made explicit. Building the evaluation cell itself from off-the-shelf parts so any lab can rebuild the exact bench is the first credible answer to 'reproduced' having no physical meaning in manipulation, and the 2026 benchmark reports consistent results across independently constructed copies with in-distribution and OOD protocols. But that cross-setup consistency is reported BY the originating group about its own copies — one benchmark, one group, no multi-vendor adoption. Modeled, and the promotion trigger is external rebuild, not more self-repo…

★ For the kids: Scientists made a robot test-kitchen that anyone can build a copy of, so everyone can check whether a robot brain really works.

Technical: Huang, Zhang, Tang & Xiang, VLA-REPLICA, arXiv:2605.20774 (May 2026): off-the-shelf build, in-distribution and OOD protocols, consistency across independently constructed setups reported by the originating group. Complements the TRI LBM blind randomized protocol (Science Robotics 2026).

Frozen prediction: By 2027-12-31, at least three institutions outside the originating group publish real-robot results on a shared replicable evaluation cell; if none do, cross-lab robot evaluation remains aspirational and no manipulation component may hold viable on cross-lab grounds.

early / contestedCamera-viewpoint generalization (the fixed-tripod dependency)embodied AI

Held open. The failure is reproduced across three groups; no fix reproduced.

★ For the kids: Move the camera a little and the robot forgets the chore; three teams are teaching it not to.

Technical: CamVLA arXiv 2607.05396; AnyCamVLA arXiv 2603.05868; arXiv 2608.06965.

Frozen prediction: By 2028-06-30, at least two major open-weight VLA releases document camera-pose conditioning or an equivalent view-invariance mechanism as part of the published training recipe.

reproduced / credibleContact-aided invariant state estimation (proprioceptive odometry)sensing

HELD viable — one of the few robot components that is genuinely reproduced across labs, open-sourced and shipped. Contact-aided invariant EKF on Lie groups fuses IMU, joint encoders and contact events with convergence properties ordinary quaternion EKFs lack, is peer-reviewed in IJRR, and is reused and extended across quadruped and humanoid stacks worldwide by groups unconnected to its origin. Remaining gap: drift when the rigid-contact assumption breaks on slippery or deformable ground.

★ For the kids: Robots use their balance sensor, their leg sensors and the feeling of their feet touching the ground to always know where they are.

Technical: Hartley, Ghaffari, Eustice & Grizzle (Michigan), IJRR 2020, 39(4):402-430 (RSS 2018 precursor, arXiv:1805.10410); extended by multi-sensor invariant filtering and smoothing (arXiv:2504.20615, 2025) and learned contact/leg-odometry representations (2026).

Frozen prediction: Through 2028, no learned end-to-end legged state estimator is shown by two independent labs to cut odometry drift by more than 30% versus a contact-aided invariant EKF baseline on the same hardware.

single-lab / plausibleFall recovery and get-up policies (the unattended-uptime prerequisite)locomotion

HELD modeled. A humanoid that cannot stand back up needs a human every time it falls, which destroys the economics of unattended deployment — the quiet gate under every lights-out pilot claim. Learned two-stage get-up policies now work on commodity human-sized hardware from supine and prone poses, on flat, deformable and slippery ground and on slopes including grass and snow. Held below viable because reproduction is mostly on one common platform, and no vendor publishes fall rate per operating hour, so the uptime effect is unmeasured.

★ For the kids: Robots are learning to pick themselves back up after they fall, instead of lying there waiting for a person.

Technical: He, Dong, Chen & Gupta (UIUC / Stanford), HumanUP, RSS 2025 (arXiv:2502.12152): trajectory discovery under minimal constraints then refinement into slow, smooth, torque-feasible motion; Unitree G1, tested on slopes, grass and snowfield. Extended by unified walk/run/recovery controllers (2026).

Frozen prediction: Through 2027, no commercial humanoid deployment report publishes falls per 1,000 operating hours alongside an unassisted-recovery success rate.

early / contestedElectrofluidic fiber muscles (pump-inside-the-muscle soft actuation)actuation

HELD open. A 2026 Science Robotics result puts the pump inside the muscle: millimetre-scale fibers with integrated electrostatic pumping at roughly skeletal-muscle power density (~50 W/kg), ~20% contraction strain, ~0.3 s response, untethered and silent, with bundles lifting ~4 kg, in a claimed continuously manufacturable fiber form. It is one lab, one paper, no cycle-life-under-load data, and ~50 W/kg sits one to two orders below the electric quasi-direct-drive joints humanoids actually ship. A single-lab actuator power-density result is not a basis to reprice an actuator supply chain.

★ For the kids: Scientists made thin robot muscle strings that carry their own tiny pumps, so they move quietly with no big machine.

Technical: Kilic Afsar, Cacucciolo, Pupillo, Vitucci, Babatain & Ishii (MIT Media Lab / Politecnico di Bari), 'Electrofluidic fiber muscles', Science Robotics, 9 April 2026 (doi 10.1126/scirobotics.ady6438). Architecturally distinct from HASEL (Acome et al., Science 2018) in embedding the pump per fiber.

Frozen prediction: Through 2029, no commercially shipped humanoid uses fiber-form electrofluidic or HASEL-class artificial muscle as the primary drive of a load-bearing joint; sustained power density of any shipped soft-muscle joint stays below 200 W/kg.

early / contestedSelf-healing actuator and skin materials (the wear-item hedge)actuation

HELD open. Thermoreversible Diels-Alder elastomers let a punctured pneumatic actuator recover most of its performance after a heat cycle, and the line has extended to multi-material reversible bonds and shape-memory-alloy-assisted damage closure. But healing needs hours at elevated temperature with the robot off duty, and nobody has published cycle-life for a healed actuator under load — which is the number that would make it a real answer to the dexterous-hand and e-skin wear-item gaps.

★ For the kids: Some robot muscles can heal their own cuts if you warm them up, a bit like a scab — but it takes hours.

Technical: Terryn, Brancart, Lefeber, Van Assche & Vanderborght (VUB), Science Robotics 2017, 2(9):eaan4268; multi-material reversible DA interfaces (2020); SMA-assisted damage closure, Scientific Reports 2023, 13:s41598-023-35943-6; 'Self-Healing and Damage Resilience for Soft Robotics', Frontiers in Robotics and AI 2017, 4:48.

Frozen prediction: By 2028-12-31, no shipping robot product lists a self-healing polymer as the wear-mitigation strategy for a load-bearing actuator or tactile skin.

single-lab / plausibleAnalytic grasp-quality metrics do not predict real success (the reproduced negative)evaluation-integrity

DEMOTED viable -> modeled by the corpus's OWN frozen independence rule, and the citation is garbled. The viable grade was awarded for a 'reproduced negative' resting on three legs: IROS 2017, RAS 121 (2019), and arXiv:1809.03276. Verified author lists: the IROS 2017 paper is Rubert, Kappler, MORALES, Schaal & Bohg; the RAS 2019 paper is Rubert, Kappler, Bohg & Morales; arXiv:1809.03276 is again Rubert et al. That is ONE lineage counted three times — Carlos Rubert's PhD work at Universitat Jaume I with Antonio Morales, plus Kappler/Bohg/Schaal at MPI-IS. Exactly the defect this corpus froze a …

★ For the kids: Scientists built a formula that scores how good a robot's grip looks, then found it doesn't match what really happens. But it was the same small group of people every time — so we can't call it settled yet.

Technical: VERIFIED: Rubert, Kappler, Morales, Schaal & Bohg, 'On the relevance of grasp metrics for predicting grasp success,' IEEE/RSJ IROS 2017 (IEEE Xplore doc 8202167; also listed at MPI-IS Autonomous Motion as 2017_iros_rkmsb). Rubert, Kappler, Bohg & Morales, Robotics and Autonomous Systems 121 (2019), doi S0921889019300247. Shared authors across all cited legs: Rubert (all), Morales (all), Kappler (two), Bohg (two). Consequence for downstream grades: the negative still CAPS dexterous-grasp-synthesis and cross-hand-embodiment-transfer at open — a modeled negative is sufficient to withhold promoti…

Frozen prediction: By 2028-12-31, no group sharing zero authors with the Rubert/Morales/Bohg lineage publishes a metric-versus-real-outcome study on >=2 grippers; the negative remains single-lineage. Separately unchanged: no SINGLE analytic metric reaches >=0.8 AUC for real-robot grasp success on a gripper absent from its fitting set.

early / contestedHuman-relative throughput (the unreported speed axis)evaluation

DEMOTED modeled -> open. The grade rested on exactly two sources and both fail the bar. (1) Epoch AI's 'Where Autonomy Works' — VERIFIED by direct fetch as Riviere & Denain, 10 Feb 2026, and verified as an explicitly SECONDARY analysis: the authors 'review the available evidence' from published videos, commercial deployments, papers and interviews, conduct no trials of their own, and acknowledge limited data. The 3-10x speed figures trace to vendor demonstrations (Figure's package sorting 'roughly four times slower'; Physical Intelligence key insertion 'about 5x'). Those are vendor demo numbe…

★ For the kids: Everyone says robots are much slower than people. The number comes from watching company demo videos, not from anyone with a stopwatch and a fair test.

Technical: VERIFIED: Riviere & Denain (Epoch AI), 'Where Autonomy Works: Evaluating Robot Capabilities in 2026,' published 2026-02-10; methodology confirmed as review of demonstrations and deployments, not original measurement; authors state 'in some cases, data on one or more of these dimensions is limited.' Arkhangelskiy, PhAIL, arXiv:2605.29710 (2026) — single author, Franka FR3, Human-Relative Throughput with bootstrap CIs anchored to same-fixture human teleoperation, best VLA ~7x slower per operation (RMST ratio), n unstated per condition. Promotion trigger: a second lab re-running the same-fixture…

Frozen prediction: By 2027-12-31, no second group publishes a same-fixture human-referenced throughput measurement using PhAIL's protocol or an equivalent, and no peer-reviewed real-robot generalist-manipulation paper reports median human-relative throughput >= 0.5 across >= 10 tasks.

reproduced / credibleLiDAR-inertial odometry (FAST-LIO family scope)sensing

HELD VIABLE, with two of its three cited legs STRUCK and the scope corrected — the component was internally incoherent as submitted. It named FAST-LIO2, LIO-SAM and KISS-ICP as the reproduced legs, then reported that its own third-party referee found LIO-SAM FAILING OUTRIGHT on the Bunker DVI data alongside DLIO, LIO-EKF, MAD-ICP, POINT-LIO and RESPLE. You cannot cite a system as a leg of a reproduced capability and simultaneously report that it collapsed under the only independent evaluation you cite. The corrected grade: viable attaches to the FAST-LIO family (FAST-LIO, FASTER-LIO, FAST-LIO…

★ For the kids: Robots find their way by sweeping a laser and feeling how they tip. Outsiders tested a dozen versions on new data: the well-known one passed, and several others just fell over — including two we had wrongly listed as proof it works.

Technical: VERIFIED extant: 'The benchmark of LiDAR odometry algorithms...', ISPRS Archives XLVIII-1/W6-2025, article 25 (2025), Bunker DVI dataset, ~20 algorithms evaluated; the referee is genuinely third-party. Xu, Cai, Zhou & Zhang, FAST-LIO2, IEEE T-RO 38(4):2053-2073 (2022, HKU MaRS). STRUCK as legs of the viable grade: Shan, Englot, Meyers, Wang, Ratti & Rus, LIO-SAM, IROS 2020 (reported as failing outright in the cited benchmark — it may not be counted as reproduction evidence in the same breath); Vizzo, Guadagnino, Mersch, Wiesmann, Behley & Stachniss, KISS-ICP, IEEE RA-L 8(2) (2023) — a fine sy…

Frozen prediction: Through 2028, no learned end-to-end LiDAR or LiDAR-inertial odometry is shown by two independent labs (neither authoring the method) to beat FAST-LIO2 by more than 20% absolute trajectory error on a public benchmark suite the authors did not create; classical filtering/registration frontends remain the shipped default.

reproduced / credibleMassively parallel precision positioner robots (astronomy fiber positioners)actuation

CONFIRMED VIABLE on independent verification, with one scoping correction the component omitted and which materially changes what it proves. Verified: DESI operates 5,000 robotic fiber positioners in 10 wedge-shaped petals, eccentric theta/phi kinematics, 10.4 mm pitch, fiber tips patrolling 12 mm disks, with a placement requirement of <=5 um RMS. Reproduced across genuinely independent observatories and teams (SDSS-V FPS, Subaru PFS, 4MOST, WEAVE), operating on sky for years, and validated by EXTERNAL methods (fiber dithering, focal-plane astrometric calibration) rather than by the positione…

★ For the kids: A telescope has five thousand tiny robot arms that each land a glass thread closer than a hair's width — but they get there by moving, checking, and correcting three times, not in one go.

Technical: VERIFIED: DESI focal plane — 5,000 positioners, 10 petals of 500, 36 deg per petal, ~6 mm nominal patrol radius with fiber tips patrolling 12 mm disks tangent to the focal surface, 10.4 mm pitch with overlapping patrol regions, one on-axis theta motor and one eccentric phi motor, requirement <=5 um RMS achieved iteratively (blind move to ~50 um, then two corrective moves). Primary: 'The Robotic Multiobject Focal Plane System of DESI', AJ (2023), doi 10.3847/1538-3881/ac9ab1. External validation: fiber dithering (arXiv:2403.05688); DESI focal-plane astrometric calibration (arXiv:2307.06238). I…

Frozen prediction: By 2028-12-31 no humanoid-hand vendor publishes per-actuator closed-loop placement repeatability measured against an EXTERNAL metrology reference (not the joint's own encoder), under any convergence protocol, at any stated tolerance.

single-lab / plausibleInstrument-grade joint metrology (the characterization discipline robotics skipped)sensing

CONFIRMED MODELED on independent verification of the source, which I fetched and read rather than inheriting from the pantry lead. arXiv:2607.22227, submitted 2026-07-24, title and author list confirmed exactly as cited (Funes Vecino, Mercant Rubio, Alvarez Uruena, Garcia Moreno, Carracedo Carballal, Ferro Rodriguez, Argelaguet Vilaseca, Piqueras Lopez), and the abstract confirms first results from the LPOA Engineering Model characterisation campaign on two rotary joints — shoulder and elbow — reporting angular resolution, discrete step tracking VALIDATED BY INDEPENDENT METROLOGY, rotation-ax…

★ For the kids: Before a telescope's little arm is trusted, engineers check it with a laser, a mirror and a gap sensor — three tools that can't be fooled by the arm's own ruler. Robot companies mostly just trust the ruler.

Technical: VERIFIED by direct fetch of arXiv:2607.22227 (title, all eight authors, submission date 2026-07-24, abstract). Reported quantities per the submitting lens's full-text read, retained but flagged as single-source: shoulder noise-limited angular resolution 4.34e-6 deg (~0.08 urad) against a 1.45e-6 deg floor; elbow 1.11e-4 deg (~1.9 urad); interferometric confirmation of 0.44 um steps with 0.22 um in-window noise; shoulder axis wobble dominated by one once-per-sweep harmonic (1839 urad combined) leaving 17.0 urad p-p / 4.9 urad RMS residual after removal; radial runout residual roundness 1.2 um …

Frozen prediction: By 2027-12-31, fewer than three peer-reviewed humanoid-robot papers report joint angular resolution, rotation-axis wobble OR bearing runout validated by an instrument independent of the joint's own encoder (autocollimator, interferometer or capacitive probe).

reproduced / crediblePassive-dynamic walking efficiency (the cost-of-transport road not taken)power

CONFIRMED VIABLE after applying the independence test the corpus froze this cycle (count PEOPLE and advisory lineages, not institution names). The Collins/Ruina/Tedrake/Wisse result is three PHYSICALLY DISTINCT machines built by three distinct PIs with distinct advisory lineages (Collins with Ruina at Cornell; Wisse in the Delft line; Tedrake at MIT), using genuinely different actuation modalities (solenoid ankle push-off vs pneumatic hip muscles), all measuring the same quantity by the same definition. Co-publication of independent replications is stronger evidence than sequential replicatio…

★ For the kids: In 2005 three different university teams each built a walking robot that used about as little energy as a person. One later walked 40 miles on one battery charge. Today's big humanoids use far more — and none of them will say how much.

Technical: Collins, Ruina, Tedrake & Wisse, 'Efficient bipedal robots based on passive-dynamic walkers', Science 307(5712):1082-1085 (2005). Origin: McGeer, 'Passive dynamic walking', IJRR 9(2):62-82 (1990). Endurance: Bhounsule, Cortell, Grewal, Hendriksen, Karssen, Paul & Ruina, IJRR 33(10):1305-1321 (2014). Definitions: c_et = energy used / (weight x distance); c_mt counts only positive actuator mechanical work. Verified figures: Cornell 12.7 kg at 0.44 m/s, 11 W total, c_et 0.2 / c_mt 0.06; Delft 7 kg at 0.4 m/s, c_et 1.3 / c_mt 0.1; human c_et ~0.2 / c_mt ~0.05. Note the Ranger follow-up shares the…

Frozen prediction: Through 2028-12-31, no commercial humanoid vendor publishes a measured total cost of transport under stated payload and speed; and the first independently published c_et for any commercially sold humanoid lands above 1.0.

reproduced / crediblePiezoelectric / ultrasonic motors (the reproduced actuator with zero-power holding)actuation

HELD VIABLE with a scoping guard added, because the grade is unusually easy to misread. The capability — precision rotary/linear actuation with self-locking and zero holding current — is reproduced at the strongest level available anywhere in this corpus: independent commercial manufacture by at least four unrelated firms on three continents (Shinsei, Canon EF USM since 1987, Nanomotion, Physik Instrumente), sustained for three decades, in safety- and metrology-critical duty (MRI-compatible mechanisms, satellite positioning, camera optics). Commercial reproduction by independent manufacturers…

★ For the kids: Some motors move by buzzing instead of spinning magnets. They're in camera lenses and hospital scanners, and they hold perfectly still using no power at all. They're also far too weak and wear out too fast for a robot leg.

Technical: Uchino, 'Piezoelectric ultrasonic motors: overview', Smart Materials and Structures 7(3):273 (1998); Sashida & Kenjo, 'An Introduction to Ultrasonic Motors' (Oxford, 1993); Naz & Xu, 'A Comprehensive Review of Piezoelectric Ultrasonic Motors', Micromachines 15(9):1170 (2024) — records self-locking with high holding torque and microsecond response, alongside the low-efficiency, stator-rotor friction-wear and 'not appropriate for continuous operation for long periods' limits; representative rotary torque ~0.94 Nm. Cross-links: static-holding-power-tax (the property humanoids want and lack), ele…

Frozen prediction: Through 2029-12-31, no commercially shipping humanoid, quadruped or mobile manipulator uses a piezoelectric or ultrasonic motor as the primary drive of a load-bearing joint.

early / contestedBiohybrid (lab-grown muscle) actuatorsactuation

HELD OPEN and the framing correction endorsed: this is an open research line, NOT an emerging actuator option, and it must not be cited to raise the grade of HASEL, twisted-coiled-polymer, dielectric-elastomer or any other soft-actuation component by association. The Science Robotics 2025 result is a genuine scale jump — an 18 cm hand driven by ten multiple-muscle-tissue actuators, where prior biohybrid devices were ~1 cm and single-joint. It is also millinewton-scale, decays after ~10 minutes of electrical stimulation with ~1 hour recovery, runs in culture medium, and comes from one lab with…

★ For the kids: Scientists grew real muscle in a dish, rolled it up like a sushi roll, and used it to pull a plastic hand's fingers. It works for about ten minutes, then it needs a nap.

Technical: Ren, Morimoto & Takeuchi (University of Tokyo / Waseda), 'Biohybrid hand actuated by multiple human muscle tissues', Science Robotics, 12 Feb 2025, doi 10.1126/scirobotics.adr5512. Ten MuMuTAs (two per finger), sushi-roll rolled cultured muscle strands cable-coupled to a 3D-printed multi-joint hand. Wireless bioelectronic control architectures for biohybrid systems are being explored (arXiv:2603.24959, 2026) but no independent group has reproduced a multi-joint biohybrid limb.

Frozen prediction: By 2029-12-31, no peer-reviewed paper demonstrates a biohybrid muscle actuator sustaining >=1 N for >=8 continuous hours without culture-medium exchange.

single-lab / plausibleLatch-mediated spring actuation (elastic power amplification)actuation

DEMOTED viable -> modeled. The proposed viable grade rested on 'four independent labs', but the four instances are Salto (series-elastic power modulation, 100 g), Hawkes work-multiplication (ratcheted rotary loading, 30 cm), RAMIEL (parallel-wire monopede) and a survey. These are four bespoke one-off devices with four different transmission topologies and four different measurements. No lab has reproduced another lab's jump height, agility figure or latch-cycle count, and there is no shared artifact or benchmark. That is multi-lab EXISTENCE, not multi-lab REPRODUCTION — precisely the distinct…

★ For the kids: Lots of teams built slingshot robots and they all jumped high. But every team built a different slingshot and measured it their own way, so nobody has checked anyone else's number.

Technical: Ilton et al. Science 360:eaao1082 (2018) is theory, not reproduction. Haldane et al. Sci. Robotics 1(1):eaag2048 (2016) and IROS 2017; Hawkes et al. Nature 604:657-661 (2022); RAMIEL arXiv:2311.04573 / 2403.11205. Each reports its own device against its own baseline. The deployment-relevant claim (duty-cycled elastic recycling at human scale) is unreproduced by the lens's own admission, and latch mass, latch wear and loading-time penalty all scale adversely. Modeled is the correct tier for a physically sound mechanism demonstrated repeatedly in bespoke hardware without cross-lab measurement a…

Frozen prediction: No commercially shipping general-purpose humanoid discloses a latch/clutch-mediated elastic power-amplification stage in a primary leg joint by 2027-12-31.

single-lab / plausibleMagnetorheological fluid clutch actuationactuation

HELD at modeled. Two genuinely independent groups (Western Ontario/Kermani, Sherbrooke/Plante) have built and bench-characterised working MR-clutch actuators, which clears the single-lab bar. Nothing above that is earned: every headline torque density (>100 Nm/kg) is author-reported and clutch-referenced rather than system-level, no group has reproduced another's figure, and multi-year fluid sedimentation, shear-gap thermal load and wear under a robot duty cycle are unpublished. The group's own centrifugal-pumping failure-mode work argues for holding, not raising.

single-lab / plausibleLiquid crystal elastomer (LCE) actuatorsactuation

HELD at modeled, with the scope stated explicitly: what is reproduced across many labs and a decade is the MATERIAL behaviour (muscle-comparable strain, stress and work density under thermal drive). What is not demonstrated anywhere is an LCE-driven joint carrying a kilogram-scale load at useful bandwidth for a useful number of cycles. Every 2025-2026 review names the same blocking limits — seconds-scale actuation and recovery, narrow operating-temperature window — so this is a material with a maturity, not a robot capability with a maturity.

Technical: Publication linkage retired: the photochromic hydrogel contact-lens paper (arXiv:2607.20770, verified 2026-07-22, UC San Diego) was attached to this component as supporting evidence. I fetched it. It is a real paper, but it is a photochromic dye-in-hydrogel system with no liquid crystal phase, no actuator and no robot — DMD grayscale lithography writes a dye gradient, and the aperture response is passive optical attenuation. It supports nothing about LCE actuation and has been struck from this component's evidence.

early / contestedActuator and transmission life qualificationactuation

HELD at open, correctly graded. Space mechanisms publish accelerated life tests with torque as the accelerating stress, material-pair screening and stated duty-cycle profiles; humanoid robotics publishes none of it. The single honest open-hardware data point is 60 hours (Berkeley Humanoid Lite, arXiv:2504.17249). No vendor L10 or MTBF against a torque spectrum exists in public, which makes every dollars-per-productive-hour claim in the sector unfalsifiable.

early / contestedAbsolute joint position encoderssensing

CONFLICT RESOLVED DOWNWARD: one lens proposed viable, another open. Open is correct. The viable proposal graded the generic industrial capability (optical/magnetic/inductive absolute encoders, decades-reproduced) and then carried humanoid-specific numbers that are entirely vendor-sourced: 30-40 joint modules per robot, two encoders per joint, +/-0.05 deg off-axis absolute accuracy, <0.01 deg direct-drive. None of those has an independent metrology source, and installed accuracy, thermal drift across the 20-80 C range these joints actually run at, and immunity to the joint's own motor field ar…

★ For the kids: Every robot joint needs a tiny ruler inside it. The rulers are real and they work — but all the numbers about how good they are once bolted into a robot come from the companies selling them.

Technical: The metrology to fix this has existed since 2003 and robotics has not adopted it: equal-division-averaged (EDA) self-calibration (Watanabe et al., Proc. SPIE 5190, 400-409) lets two previously uncalibrated angle detectors calibrate each other with no higher-order reference — which is exactly the configuration a joint with motor-side and output-side encoders already has installed. Output-side resolution also silently bounds every joint-torque estimate derived from deflection across a known stiffness, so this open grade propagates into the torque-sensing components.

early / contestedContactless inductive/resonant power transfer for mobile robotspower

HELD at open. Textbook physics, sparse and poor robot-relevant numbers: the one honestly measured mid-sized inspection-robot figure is 47% end-to-end, and coupling falls roughly with lateral offset squared over coil radius — the exact quantity autonomous docking fails to control. Losses land as heat inside a machine that already has a thermal budget problem.

reproduced / credibleCommodity 3D depth camerassensing

HELD at viable on CAPABILITY ONLY. Structured-light and active-IR stereo depth is used and reproduced across hundreds of independent labs and shipped products, which clears the bar. The concentration story attached to it does not, and has been struck: the '60% of AMRs / 80% of humanoids' attach share is a vendor press-release statistic with no disclosed denominator and no methodology, and it appears nowhere in this record as evidence. The corporate fact (RealSense spin-out completed Oct 2025, $50M raise) is an industry source establishing a corporate event, not a market measurement.

Technical: The load-bearing engineering point survives and is worth keeping: the depth ASIC, IR projector and calibration pipeline are a single-vendor bundle, so swapping vendors changes the noise model and policies trained on one depth camera do not transfer cleanly. That is a real coupling, independent of any share number.

reproduced / credibleVisual-inertial odometry and SLAMsensing

HELD at viable, and it is the cleanest viable grade in the whole set. Multiple open implementations from independent groups (ORB-SLAM3, OpenVINS, VINS-Fusion, Basalt, Kimera) evaluated by third parties on shared public datasets they did not create (EuRoC, TUM-VI) with millimetre-accurate motion-capture ground truth, plus author-independent comparison papers free to embarrass anyone. This is what the bar actually looks like, and it is carried here mainly as the reference against which every tactile and manipulation grade should be read.

single-lab / plausibleVisuo-tactile object pose and shape trackingtactile

HELD at modeled. NeuralFeels (Science Robotics 2024) plus independent MIT lineage (Tac2Pose, SimPLE) plus V-HOP (RSS 2025, Brown/UT Dallas) evaluating on the released FeelSight data clears the single-lab bar. It does not clear viable: V-HOP is BENCHMARK reproduction, not HARDWARE reproduction — no lab has rebuilt the hand and the tactile fingertips and recovered the effect — and the headline 94% figure is an under-occlusion conditional, not an average.

early / contestedThe marginal utility of touchsensing & tactile

Held open. Three single-lab positives on three different tasks/sensors/baselines (PKU, Darmstadt, Meta); no shared task at equal data budget.

★ For the kids: Three labs each say touch helped, on three different tests, so we cannot say how much touch is worth.

Technical: Bar: one shared physical task, pre-registered ablation, >= 50 trials/arm, >= 2 sites.

Frozen prediction: By 2027-06-30 a touch-vs-no-touch ablation on one shared task with >= 50 trials/arm is reported from >= 2 sites. p=0.25.

early / contestedTactile sensor metrology and unit-to-unit variancesensing & tactile

Held open. AnySkin (NYU 2024/ICRA 2025) reports zero-shot cross-instance transfer, single-lab, unreplicated.

★ For the kids: Two 'identical' fingertips feel different; one lab says it fixed that, nobody else has checked.

Technical: Bhirangi et al., arXiv 2409.08276.

Frozen prediction: By 2027-06-30 at least one lab outside NYU reports AnySkin-class cross-instance transfer retaining >= 80% success. p=0.45.

single-lab / plausibleWhisker and antenna tactile sensingtactile

HELD at modeled, and the entry-level grade-down is endorsed as the correct precedent. Northwestern, Sheffield/Bristol, Stanford and Purdue all build whiskers; each builds a bespoke filament with a bespoke transducer and its own inverse model, and no lab has replicated another's localisation figures. Fifteen years of multi-lab existence with zero multi-lab reproduction. This is the reasoning that also forces spring-latch-power-amplification down.

single-lab / plausibleThermal tactile sensing (material identification by heat flux)tactile

HELD at modeled. Multi-lab and multi-decade (BioTac shipped a thermal channel; Georgia Tech formalised recognition under varying initial conditions), so the physics is not in doubt. It stays at modeled because after ~15 years no manipulation policy, VLA or shipped robot uses a thermal touch channel as an input, and the measurement is confounded by contact area and the object's initial temperature — the two things a robot does not control.

single-lab / plausibleSoft-body and continuum proprioceptionsensing

HELD at modeled. Multiple independent groups (UCSD/Sant'Anna, MIT CSAIL, Khalifa) demonstrate the capability, each on its own body with its own sensor topology, with no shared benchmark. The blocking omission is uniform: no published proprioception accuracy after thousands of actuation cycles, in a system where the sensor and the structure are the same creeping viscoelastic material.

early / contestedSensorless learned external-force estimationsensing

DEMOTED modeled -> open. I fetched FACTR 2 (arXiv:2606.12406, submitted 2026-06-10, CMU) and confirmed it: NEXT estimates external joint torques with no dedicated force sensor from ~10 min of free-motion data, claims estimates 'comparable to dedicated joint-torque sensors', and reports >17% task-progress gain across five long-horizon tasks. The abstract states no trial counts. This is one lab, one paper, weeks old, and it is the branch carrying the entire economic implication. The classical branch (De Luca momentum residual, Haddadin survey) is textbook-solid but is a different, already-track…

★ For the kids: One team says a robot can guess how hard it is being pushed without a sensor, as well as a real sensor would. Nobody else has tried it yet, and they did not say how many times they tested.

single-lab / plausibleNon-invasive sEMG neuromotor interfacessensing & tactile

Held modeled. Kaifosh et al., Nature 645:702 (2025, Reality Labs): thousands of participants, open dataset — single company; cross-user generalization on independent hardware unreproduced.

★ For the kids: A wristband reads arm signals and types for strangers; one company has shown it.

Technical: EMGBench (2024) shows OOD generalization remains open for others.

single-lab / plausibleFoveated / gaze-driven active vision for manipulationperception

HELD at modeled. Two independent groups converge (Berkeley's Eye, Robot; Look, Focus, Act) plus 2026 follow-ons, which clears single-lab. It stops there: nobody outside the authoring labs has run either method head-to-head against a fixed multi-camera baseline, so the load-bearing claim — that moving the eye beats adding cameras — is untested by a disinterested party.

reproduced / credible6D pose estimation of unseen objectssensing

HELD at viable. The BOP Challenge has run an open leaderboard on shared datasets since 2017 with methods from dozens of independent groups, and since 2023 the unseen-object variant (CAD or reference-image onboarding at test time) works. Reproduction is structural rather than asserted. The honest caveat carried forward: unseen-object 2D DETECTION remains roughly 35% behind seen-object accuracy and is named by the organisers as the pipeline bottleneck, and it is an open question whether solved pose still raises end-to-end manipulation success once policies are learned end-to-end.

reproduced / credibleFleet-scale picking (suction and parallel-jaw singulation)manipulation

HELD at viable, scope narrowed to what is actually deployed. Deployed high-mix bin picking runs commercially at scale and is replicated by several independent vendors, and ARMBench (235K+ activities, 190K+ objects, with a failure taxonomy) is the only published fleet-scale manipulation corpus. The narrowing matters: the viable grade covers suction and parallel-jaw singulation, not dexterity, and ARMBench itself is one company's dataset — it is the deployment that is reproduced, not the dataset. Any dexterous-manipulation revenue model extrapolating from n=10 lab trials into this regime is ext…

early / contestedThe missing force channel in manipulation datamanipulation

HELD at open. The dominant collection pipeline (ALOHA-class leader-follower teleop, replicated at dozens of institutions) records joint positions and vision, not calibrated wrench; every force-recording rig is a one-off with bespoke calibration and no shared force-labelled corpus exists at scale. Both readings are unfavourable to dexterity claims: either force is necessary and the corpus is missing its most important channel, or force is unnecessary and the benchmarked tasks are too easy to evidence dexterity.

early / contestedThe sub-millimetre tolerance gap in learned assemblymanipulation

HELD at open, and this is one of the most decision-relevant numbers in the set. Academic manipulation benchmarks mate at 1-2 mm (FMB, chosen so parts print on commodity 3D printers); best sim-to-real assembly reaches 0.1-1.0 mm at 96.7% (IndustReal, single lab, per-part policies); the NIST Assembly Task Board runs at roughly 0.005-0.029 mm; the 2026 tactile frontier reaches 0.05 mm at 67%, a rate no production line accepts. Learned manipulation is 3-20x looser than the industrial reference at production-grade success. Fixtures and passive compliance, not learning, remain the load-bearing tech…

early / contestedGranular and fluid media manipulationmanipulation

HELD at open. Nine years after the first CNN granular-dynamics models, the frontier is a single-lab zero-shot diffusion policy graded by difficulty level over 465 trials, and the shared benchmark (TOTO) covers three materials. Pose estimation and grasp synthesis — the two workhorses of rigid manipulation — do not apply at all to a state that is effectively infinite-dimensional.

single-lab / plausibleHuman-robot object handovermanipulation

HELD at modeled. Mature review literature (T-RO 2021) plus reactive human-to-robot handover of arbitrary unseen objects (ICRA 2021) plus several independent follow-ons clears single-lab. No shared benchmark, no standard success metric, and lab-timed rather than deployment-measured results block anything higher — a benchmark vacuum in a skill every human-adjacent robot needs.

single-lab / plausibleAssistive shared-autonomy manipulationmanipulation

HELD at modeled. Multiple groups, real end users, out-of-lab studies and a genuine safety stake — stronger evidence than most manipulation research produces. Held below viable because the systems remain research prototypes or fixed-trajectory commercial devices, no cross-lab protocol exists, and the honest user-preference finding (users with mobility impairments do not uniformly prefer full autonomy) cuts against the full-autonomy target the humanoid field has set.

single-lab / plausibleHybrid force-position policy learninglearning

HELD at modeled. At least six independent groups in 2024-2026 find that emitting a compliance/stiffness or force target alongside position beats position-only imitation on contact-rich tasks, and they converge on the same structural finding — the vision-to-force modality handover point matters more than the force representation. What is reproduced is the direction; no single architecture's numbers are, because every paper uses different arms, different sensors and n in the 10-30 range.

reproduced / credibleVision-based agile aerial sim-to-reallocomotion

HELD at viable, and it is the strongest sim-to-real evidence in robotics — stronger than the legged-locomotion exemplar it should probably displace. Three institutions across two genuinely different method families (UZH model-free policy gradient; TU Delft model-based DreamerV3 from pixels; the FalconGym group) reach the same physical capability on different airframes: zero-shot transfer of RL flight policies to real quadrotors at aggressive speeds, onboard compute, no motion capture. Unlike the bespoke-device cases I demoted, the capability here is measured against an objective, externally c…

reproduced / credibleWearable exoskeleton assistance (the subfield with a referee)wearable-robotics

HELD at viable, on the narrow claim only: portable exoskeletons measurably reduce the metabolic cost of walking, reproduced across at least four independent labs using indirect calorimetry — a third-party-measurable physiological scalar with a physiological plausibility prior. The 2026 Matters Arising and rebuttal over the sim-trained magnitude is a feature, not a defect: it demonstrates the field can phrase a disagreement as a testable claim, which humanoid manipulation cannot. The viable grade covers the effect, NOT the contested experiment-free-simulation magnitude, which remains unreplica…

reproduced / credibleWarehouse AMR fleets (the only embodied autonomy proven at fleet scale)deployment

HELD at viable; the duplicate slug warehouse-amr-fleet-deployment is merged into this one. Peer-reviewed anchor (Kiva, AI Magazine 2008) plus independent replication by many vendors and many operators. Two negatives are the actual content: the stack is fiducial/SLAM navigation plus fleet scheduling plus constrained pick, so fleet-scale embodied autonomy was achieved WITHOUT the technologies this study tracks; and eighteen years from the founding paper to low-double-digit warehouse penetration forecasts is the adoption prior any steeper humanoid curve must beat. Scale figures (>1M units, billi…

single-lab / plausibleCross-embodiment visual navigation foundation modelsautonomy

HELD at modeled (three lenses agreed; no conflict to resolve). Open weights and code make this the default navigation baseline and third parties do run it, which is more than any manipulation VLA can claim. But every reported cross-embodiment transfer number traces to one research lineage, independent 2026 zero-shot field evaluations report substantial gaps rather than confirmations, and there is no third-party head-to-head against a tuned classical SLAM-plus-planner baseline on shared hardware — the comparison that would decide whether the foundation model is better or merely newer. Worth ca…

single-lab / plausiblePersistent open-vocabulary 3D scene graphs and spatial memoryperception

MERGED from three duplicate proposals (open-vocabulary-3d-scene-memory, spatial-semantic-scene-graphs, spatial-memory-scene-graphs) which were triple-counting the same Hydra/ConceptGraphs evidence across lenses. Modeled: the CONSTRUCTION step is genuinely reproduced — Hydra plus an independent monocular reimplementation, ConceptGraphs, Clio and five-plus groups with open code. Two things are not demonstrated and both are load-bearing: persistence (no published result maintains one scene graph on a deployed robot across weeks of a changing environment with relocalisation and node-consistency m…

single-lab / plausibleVLA quantization and distillationcompute

HELD at modeled. One lab has demonstrated the capability end-to-end (4-bit quantized VLA at 150.5 ms / ~6.6 Hz fully offline inside a ROS 2 loop), which is exactly a modeled-tier demonstration; the surrounding distillation and speculative-decoding papers largely lack real robots. Two disciplines enforced: the headline 93x / 5-6 ms figures come from vendor-adjacent non-peer-reviewed write-ups and are not evidence, and 6.6 Hz remains an order of magnitude below what dexterous contact-rich control needs. The one durable technical finding, replicated across several of these papers, is that the la…

early / contestedVLM reward models and task-progress estimationevaluation

HELD at open. Several 2025-2026 process-reward and progress models exist, but the benchmark that would tell us whether such a judge is trustworthy enough to close a training loop is itself brand new and unadjudicated. The structural hazard nobody tests: a reward model sharing a backbone with the policy it grades has correlated blind spots, which makes an autonomous data flywheel self-confirming rather than self-improving.

single-lab / plausibleRuntime safety filters for learned policies (CBF / shielding)safety

HELD at modeled. The mathematics is mature and reproduced across labs including on humanoid hardware, with several independent formulations converging. The gap blocking viable is institutional rather than mathematical: a CBF guarantee holds under an assumed model and disturbance bound, certification bodies assess systems rather than Lyapunov arguments, and no such filter has ever appeared inside a published conformity assessment. This gates humanoid safety certification, which means certification — not capability — is the binding constraint on deployment near people.

early / contestedCloud/fog offload of robot computationcompute

DEMOTED modeled -> open. Every cited artefact is one lab's line (FogROS2 plus its own extensions for server selection, latency-reliability and fault tolerance), so this does not clear even the multi-group bar, and the capability actually being graded — offload as a CONTROL-LOOP strategy — has no field evidence at all: no published tail-latency budget small enough for contact-rich control over commodity networks across a working day, and no vendor disclosure of what a robot does when the link drops. That last omission is also the disclosure that would separate an autonomous product from a remo…

single-lab / plausibleEgocentric wearable capture of human manipulationdata

HELD at modeled, with the two claims kept apart. The LOGISTICS claim — that headset-scale capture of egocentric video with per-frame 3D hand pose works at scale — is reproduced across groups (829 h / 194 tasks / 338K demonstrations with 25 joints per hand, plus smaller independent pipelines). The CAPABILITY claim — that this data improves real robots over an equal-budget robot-data baseline — has never been confirmed by an independent lab. The field is scaling the arm whose payoff is least tested, and the decisive experiment (hold robot-data budget fixed, vary the human-data source) is cheap …

early / contestedDifferentiable simulation and analytic policy gradientssimulation

HELD at open, and correctly graded as CONTESTED rather than merely early. The machinery is real and fast, but gradients through hard contact are biased when the contact normal moves and are exactly zero when bodies are not touching, so realistic solver stiffness and correct gradients pull against each other. A 2026 ICLR paper re-examined the field's headline result and found the proposed discontinuity-detection fix needs task-specific tuning and has low sample efficiency, with plain variance control often dominating it. No independent group has shown a real-world contact-rich win over an equa…

single-lab / plausibleSimulator physical-fidelity validationsimulation

HELD at modeled. Single-lab, but it is the only work measuring the simulator instead of the policy, in engineering units rather than success-rate percentages. The reproduced negative worth carrying: after per-system contact-parameter identification, inelastic impacts are captured reasonably by Drake/MuJoCo/Bullet and elasticity generally is not, and biped landing error is dominated by robot-model mismatch rather than solver choice. There is no adopted fidelity metric and no venue requires one.

single-lab / plausibleProcedural environment and task generationsimulation

HELD at modeled, grading the infrastructure and not the promise. The suites are genuinely reproduced as infrastructure — widely downloaded and used by other labs. The load-bearing claim, that synthetic environments substitute for real demonstrations on a real robot, has never been shown by a third party at adequate trial counts.

early / contestedRemote real-robot evaluation servicesevaluation

HELD at open, correctly. Verified real (42 simulated and 18 real tasks, cloud real-robot access, 30 policies, public leaderboard, submitted 2026-07-05). Open on three counts: it is weeks old; it is run by a 43-author consortium that also builds policies, which is the exact conflict structure that disqualifies a referee; and there is no second, independently operated service to check its rankings against. An evaluation monopoly is a governance problem before it is a scientific one.

single-lab / plausibleParallel-elastic actuation (springs carrying the gravity load)actuation & power

Held at modeled. Springs in parallel with the motor carry static/cyclic load so the motor supplies residual torque. Measured electrical-energy reductions exist in single-lab hardware (Sandia STEPPR; Michigan monoped series-vs-parallel comparison) and in vendor-built bipeds (Agility Cassie leaf springs), but no two independent groups have built and measured the same parallel-elastic joint with an ablation.

★ For the kids: Imagine a pogo stick inside a robot's leg. The spring holds the robot up so the motor does not have to push all the time, which saves battery.

Technical: Mazumdar et al., IEEE/ASME Trans. Mechatronics 2017 (STEPPR). Yesilevskiy, Xi & Remy, ICRA 2015 (monoped: parallel wins energy at high-frequency gaits, loses shock tolerance). Agility/OSU Cassie reports are vendor-adjacent and add no independent tier.

Frozen prediction: By 2028-08-21 at least two independent academic groups (not Agility, not Sandia) publish measured cost-of-transport reductions >=15% on a biped attributable to parallel elastic elements vs the same robot with springs removed. Probability 0.45.

single-lab / plausibleElectrohydraulic (HASEL-class) musculoskeletal limbsactuation & power

One peer-reviewed leg from one lineage (ETH Katzschmann / MPI-IS Keplinger, Buchner et al., Nature Communications 2024), tethered to an off-board kV supply. Strong single-lab result; no reproduction of a hopping/walking HASEL limb outside the lineage; no untethered limb-scale demonstration.

★ For the kids: Squishy oil-filled pouches that squeeze when zapped, like a muscle. One lab built a whole robot leg out of them that can hop.

Technical: Buchner et al., Nat. Commun. 2024; base science Acome et al., Science 2018; Kellaris et al., Sci. Robot. 2018; Rothemund et al., Adv. Mater. 2021. The energy comparison vs a geared motor is the authors' own, on their own rig.

Frozen prediction: By 2028-08-21 no lab outside the Keplinger/Katzschmann lineage publishes a peer-reviewed untethered legged robot whose leg joints are driven primarily by electrohydraulic muscles. Probability 0.7 the null holds.

early / contestedHumanoid runtime: vendor claims vs third-party measurementactuation & power

Held open. Anchors recorded: Unitree G1 421.2 Wh / ~2 h -> ~210 W implied; Walker S2 ~2 h / ~4 h. Neither third-party measured. Grading rule: claim 'met' if independent measurement at stated intensity reaches >= 80% of claimed runtime.

★ For the kids: We write down what companies claim so we can check when someone outside measures.

Technical: Distributor spec pages 2025-26 (G1: 9000 mAh, 13S, 46.8 V, 421.2 Wh); UBTech July 2025 press.

Frozen prediction: By 2027-09-01 an independent lab publishes a measured G1 continuous-walking runtime at >= 0.8 m/s and it is <= 96 min. p=0.60.

single-lab / plausibleArtificial-muscle figure-of-merit ledger (cross-review agreement)actuation & power

Two independent review groups tabulate specific power, stress, strain, efficiency and bandwidth across artificial-muscle classes and reach consistent rankings; no class wins efficiency plus cycle life together. The tabulations are reproduced; the per-class numbers are mostly single-paper best cases with no shared test protocol, so the ledger is modeled, not viable.

★ For the kids: Scientists made a scoreboard of all the different kinds of robot muscles. No single kind wins every column.

Technical: Mirvakili & Hunter, Adv. Mater. 2018; Zhang et al., IEEE T-RO 2019 ('Robotic Artificial Muscles: Current Progress and Future Perspectives'); Madden et al., IEEE J. Oceanic Eng. 2004 (metric framework).

reproduced / credibleVibrotactile texture & material classification (perception only)sensing & tactile

Offline closed-set texture/material classification from contact vibration and microgeometry is reproduced across labs and sensor families (USC BioTac Bayesian exploration; MIT GelSight texture recognition; Penn and Bielefeld accelerometer/piezo fingertips). Viable strictly as a PERCEPTION capability; no reproduced downstream manipulation gain, and cross-unit/cross-velocity generalization requires per-unit recalibration.

★ For the kids: Robots can tell sandpaper from silk by rubbing a fingertip on it; many labs have done this. Using that to do chores better is still unproven.

Technical: Fishel & Loeb, Frontiers in Neurorobotics 2012 (117 textures); Li & Adelson, CVPR 2013; spectral (FFT band) or learned features on 2-3 kHz pressure, 30-90 Hz images, or MEMS accelerometers.

Frozen prediction: By 2027-08-21 no peer-reviewed study shows a texture/material classifier transferring zero-shot across two tactile sensor families AND two sensor units with <10 pp accuracy loss on a >=20-class set.

single-lab / plausibleBioTac multimodal fingertip (reproduced sensor, dead supply chain)sensing & tactile

DEMOTED from viable. The SynTouch BioTac was the most-reproduced multimodal fingertip of the 2010s (USC, Penn, Bielefeld, Stanford, Shadow integrations). But the component's claim is a supply fact, not a current capability: the hardware can no longer be bought, so the capability cannot be reproduced today and per-unit calibration was always required. A historical multi-lab record on unobtainable hardware is modeled, not viable.

★ For the kids: One robot fingertip was so good that lots of labs used it for years; then the company stopped making it, so newer robots can't have it.

Technical: Wettels, Santos, Johansson & Loeb, Advanced Robotics 2008. Fluid-filled elastomer over rigid core; 19-electrode impedance, DC/AC pressure, thermistor+heater. Failure modes: unit variance, fluid leakage, wear.

Frozen prediction: By 2027-08-21 no commercial fingertip with all four BioTac channels is purchasable with a >=3-independent-group peer-reviewed track record.

single-lab / plausibleDistributed capacitive humanoid skin (iCub class)sensing & tactile

Longest-running whole-body tactile skin (IIT; ~4,000 capacitive taxels), shipped on dozens of iCub copies and used across European labs. Reproduced by distribution from one lab, not by independent re-implementation, so modeled.

★ For the kids: One lab built a robot with touch-sensitive skin over its whole body and sent copies to many other labs, but nobody else has built their own version.

Technical: Maiolino et al., IEEE Sensors Journal 2013; Schmitz et al., IEEE T-RO 2011. Triangular 12-taxel modules, CDC per module, ~10 mm resolution, ~0.5-1 N threshold, periodic baseline re-zeroing for drift.

Frozen prediction: By 2027-08-21 no second humanoid platform from a different lab or company publishes a peer-reviewed whole-body skin (>1,000 taxels, >50% coverage) with measured taxel density and a drift/recalibration protocol.

early / contestedTactile sensing bandwidth vs control-rate gap (camera-rate touch)sensing & tactile

Vision-based tactile sensors run at 30-90 Hz with tens of ms latency; slip/impact transients occur in 1-10 ms. Whether camera-rate touch suffices for reactive manipulation or only for perception-before-action is untested head-to-head.

★ For the kids: Robot 'eyes in the fingertip' look only about 30-60 times a second, but things slip in a hundredth of a second.

Technical: Taylor, Dong & Rodriguez, GelSlim 3.0, ICRA 2022; slip-onset windows ~10-30 ms (Veiga et al. 2015 and related). No published ablation holds the sensor fixed and varies only rate/latency.

Frozen prediction: By 2027-08-21 no peer-reviewed same-hardware ablation shows tactile sampling <100 Hz matching >=500 Hz on a slip-arrest task with >=100 trials.

single-lab / plausiblePiezoresistive fabric / FSR pressure arrays (the cheap, hysteretic skin)sensing & tactile

Reproduced as a sensing modality in many labs (MIT scalable tactile glove, e-textile skins) but hysteresis, creep and drift keep them out of force-controlled manipulation; the literature uses them for pattern/grasp recognition, not force servoing.

★ For the kids: A very cheap pressure-sensing fabric can feel where you touch it, but it 'remembers' old squeezes so it can't tell exactly how hard you push.

Technical: Sundaram et al., Nature 2019 (548 sensors, ~$10 BOM). Hysteresis 10-25% FS; row-column scanning with crosstalk suppression; non-trivial temperature coefficient.

Frozen prediction: By 2027-08-21 no piezoresistive-fabric array appears in a peer-reviewed closed-loop grip-force servoing result with force error <10% over >=1,000 cycles.

single-lab / plausibleForce estimation from vision-based tactile sensors (the calibration question)sensing & tactile

Single-unit calibrated force estimation is good; cross-unit and post-wear calibration transfer is poorly characterized, and most policy papers consume raw images and sidestep force. Strong single-lab method, weak multi-lab metrology.

★ For the kids: The fingertip camera can guess how hard it's pressing, but every fingertip needs its own tuning and gets less accurate as the rubber wears.

Technical: Yuan, Dong & Adelson, Sensors 2017; Si & Yuan, Taxim, RA-L 2022. Gel modulus drifts with wear; ~5-10% normal-force error claimed, larger in shear/torsion.

Frozen prediction: By 2027-08-21 no study reports cross-unit (>=5 units) force-calibration transfer error for any vision-based tactile sensor on a standardized load protocol.

single-lab / plausibleTactile datasets & pretraining corporasensing & tactile

Held modeled. Meta FAIR Sparsh lineage plus cross-lab corpora (UniTouch, TVL); every downstream gain is self-evaluated; TacBench written by the evaluated party.

★ For the kids: Big collections of touch pictures help robots learn, but only their makers have graded them.

Technical: Higuera et al. CoRL 2024 (arXiv 2410.24090); arXiv 2506.14754; arXiv 2505.11420.

Frozen prediction: By 2027-08-21 no public tactile corpus exists with >=3 sensor families, >=100k samples and per-sample ground-truth contact force, released by >=3 institutions.

reproduced / credibleFixed-axis in-hand rotation of simple objects via sim-to-real (the reproduced sub-capability)manipulation

NEW slug carved out so the promotion is scoped honestly. Sim-to-real rotation of simple objects (cubes, cylinders, ~50-100 mm) about a fixed axis with a multi-finger hand is reproduced across >=5 independent labs on different hands: OpenAI (2019, Shadow), MIT (Chen et al., CoRL 2021; Sci. Robot. 2023), NVIDIA DeXtreme (ICRA 2023, Allegro), Berkeley/Meta (Qi et al., CoRL 2022), UCSD (Yin et al., RSS 2023; Yuan et al., ICRA 2024, Allegro/LEAP). Shared recipe, not hand-specific. Viable for this sub-capability only; trial counts remain 10-50/object with no shared cross-lab protocol.

★ For the kids: Several different robot labs can now spin a block around in a robot hand without dropping it.

Technical: Isaac Gym teacher-student distillation, domain randomization on mass/friction/scale, proprioception or touch as the student input. Drop rate rises with object aspect ratio; continuous rotation overheats torque-limited fingers.

reproduced / credibleMyoelectric prosthetic hands (dexterity with a clinical referee)hands-and-actuation

Multi-vendor commercial hands (Ottobock, Össur, Psyonic) evaluated with standardized clinical instruments; pattern-recognition sEMG control reproduced across labs; osseointegrated neuromusculoskeletal prostheses in chronic home use. The one branch of dexterous-hand engineering with a referee and a multi-year durability regime, built with 5-6 actuators.

★ For the kids: Robot hands that real people wear every day already exist; they are simpler than the flashy robot hands but they are tested properly.

Technical: Farina et al., Nat. Biomed. Eng. 2017; Ortiz-Catalan et al., Sci. Transl. Med. 2014, NEJM 2020. Abandonment rates (Biddiss & Chau 2007) ~20-40% remain the honest deployment metric.

Frozen prediction: By 2028-12-31 no humanoid hand vendor publishes a durability figure (cycles to failure or months of daily use, with method) meeting the bar already met by commercial prosthetic hands.

reproduced / credibleTeleoperated surgical robotics at fleet scaledeployment

The only embodied fine-manipulation business at fleet scale (roughly 10,000 installed da Vinci systems, >2.5M procedures/yr per Intuitive's 2024 10-K; adoption audited in Sheetz, Claflin & Dimick, JAMA Netw Open 2020), and it is 100% teleoperated. Sets the reliability/cost reference a humanoid teleop pilot must beat.

★ For the kids: The most successful robot hands in the world are moved by a human surgeon in the same room; the robot does not decide anything itself.

Technical: Cable-driven wristed instruments with ~10-use limits. Autonomous soft-tissue manipulation remains single-lab porcine work (STAR, Saeidi et al., Sci. Robot. 2022).

Frozen prediction: Through 2028-12-31 no FDA-cleared surgical robot performs an autonomous soft-tissue suturing step on humans.

early / contestedAgricultural harvesting robots (forty years of demos; one lineage bending the apple curve) — MERGED from two lensesfield robotics

SKEPTIC: two lenses extended the same slug; merged, lens 5's wording correction adopted. Apples: MSU/USDA lineage ~80% at 6-7.5 s over two seasons in commercial orchards (arXiv 2506.05714, 2606.14089), beating the Bac 2014 mean (66%, 33 s) — but one lineage; Monash's independent 62.8% at 9.18 s sits at the 2014 mean. Strawberries: 84.3% overall on 281 fruit in one greenhouse (Bashir/Zahid, arXiv 2605.23863). OrchardBench baseline harvests ~1/8 of reachable fruit. Not reproduced, not soft fruit at field scale; open.

★ For the kids: One team's apple robot finally beat the old average two years running; a record only counts once another team matches it.

Technical: Bac et al., J. Field Robotics 31(6) 2014; Au et al., CEA 213:108164 (2023); HortiBot IROS 2024 (n=24); Williams & Polydoros arXiv 2505.08458; OrchardBench arXiv 2607.06337.

Frozen prediction: By 2027-12-31 no lab outside the MSU/USDA lineage reports >= 75% per-attempt apple success at <= 8 s per fruit over >= 1,000 cycles in a commercial orchard. p=0.65.

single-lab / plausibleQuadruped autonomous inspection missionsdeployment

DEMOTED viable -> modeled (upheld). Shipped on multiple vendors' platforms (Spot Autowalk, ANYmal) but no third-party mission-success or MTBF figure; the reproduced quantity is the product, not the capability.

★ For the kids: Robot dogs walk inspection routes for real customers, but nobody outside has published how often they finish the route.

Technical: Re-promote on an independent operator publishing mission-success over >=1,000 missions.

Frozen prediction: By 2027-12-31 neither vendor publishes a third-party-audited unattended-mission completion rate with denominator.

early / contestedEmbodied-reasoning VLM leaderboards (the 2026 'embodied brain' race)manipulation & dexterity / reasoning front-end

Held open. Embodied-R1.5, RxBrain, Hy-Embodied-VLM, RynnBrain 1.1 are leaderboard claims plus small self-run demos; benchmark-to-real transfer uncorrelated. Pantry leads chased to arXiv and carry zero grade weight.

★ For the kids: Scoring well on a quiz about the physical world is not doing the chores.

Technical: arXiv 2606.11324; 2607.14187; 2607.12894; 2607.17977.

Frozen prediction: By 2027-09-05 no study shows rank correlation >= 0.6 between an embodied-VLM suite score and third-party real-robot success across >= 5 models. p=0.70.

single-lab / plausibleUnified cross-embodiment action spaces with embodiment maskingmanipulation & dexterity / policy architecture

Held modeled. Two labs, two masking designs, self-evaluated, no shared evaluation.

★ For the kids: One brain drives three bodies by hiding the parts that do not fit; it works a little on the robots its makers tested.

Technical: RynnBrain-VLA arXiv 2607.17977; unified hand action space arXiv 2607.03570.

Frozen prediction: By 2027-06-30 a masked cross-embodiment VLA is evaluated by an independent referee on an embodiment absent from training and scores > 50% on >= 3 tasks. p=0.30.

single-lab / plausibleJoint language + goal-image subgoal planningembodied AI

RxBrain (arXiv 2607.14187, verified) is a second industrial instance of the architecture; the ablation isolating the goal-image channel on a real robot does not exist. Modeled.

★ For the kids: The brain draws a picture of the next step before doing it; nobody has proved the picture helps.

Technical: Missing ablation: language-only vs language+imagined vs language+oracle goal image, same policy, >=50 trials/task.

Frozen prediction: A goal-image-vs-language-only real-robot ablation with >=50 trials/task and CIs appears by 2027-06-30 (P=0.5).

single-lab / plausibleTwisted string actuators (TSA)actuation-power

Motor-twisted string bundle as a gram-scale, near-zero-backlash transmission. Reproduced as a MECHANISM across independent labs for >15 years (Bremen ICRA 2010; Bologna T-RO 2013; KAIST; 2025 Measurement review; Lee et al. AIS 2025 lumbar-assist wearable). Not reproduced as a deployed limb capability: nonlinear ratio under load, string fatigue is the wear item, no commercial limb joint uses one. Modeled is the ceiling until a shipping joint exists.

★ For the kids: Twist a shoelace and it gets shorter and pulls harder. Robots can use that as a tiny gearbox, but the lace wears out.

Technical: Ratio swings ~2.1x over 0.1-1.5 kg load (arXiv 2410.12097); cycle life bounded by string creep/fibre fatigue.

Frozen prediction: No humanoid vendor ships a TSA in a leg or arm joint (hands excluded) by 2027-06-30 (P=0.85).

single-lab / plausibleCapstan / cable-drive reducersactuation-power

Rope-on-drum low-backlash reducer. Evidence: open-source maker stands (Capstan-Drive, GitHub 2023-24), a capstan quadruped kit, and one academic arm (D3-ARM, arXiv 2502.12963). Multiple independent BUILDS but zero long-duration wear data and no commercial load-bearing joint. Skeptic note: maker builds are demonstrations, not referee-graded reproductions; modeled is generous and must not rise on more builds without wear data.

★ For the kids: Instead of gears, a rope around wheels turns a fast small spin into a slow strong one.

Technical: Practical ratio ~8-15:1 (drum geometry); competes with planetary QDD, not strain-wave.

Frozen prediction: No commercial humanoid or cobot ships a capstan/cable reducer in a load-bearing joint by 2027-06-30 (P=0.85).

early / contestedElectrostatic (magnet-free) motorsactuation-power

Only industrial data is vendor-authored (C-Motive white paper, PACK EXPO 2025 demos); academic film actuators (arXiv 2511.08005) are bench devices. No robot-joint demonstration, no third-party dynamometer data. Open.

★ For the kids: Tiny electric charges pull the motor round instead of magnets. Not strong enough for a robot leg yet.

Technical: High torque at low speed with no gearbox, but poor volumetric torque density and kV-class drive electronics.

Frozen prediction: No robot joint or gripper product ships with an electrostatic motor by 2028-12-31 (P=0.9).

early / contestedHuman-equivalence actuation scoring (HLAS / HEE)actuation-power

Sunbeam, arXiv 2511.06796 (Nov 2025, verified): kinematic atlas + per-joint Human-Equivalence Envelopes + aggregate score incl. thermal sustainability, with measurement protocols. Single author, zero adopters, zero measured commercial joints. Open.

★ For the kids: A scorecard checking whether a robot joint is really as strong and tireless as a human joint.

Technical: Score inputs: workspace coverage, HEE coverage, torque-mode bandwidth, efficiency, thermal sustainability; all from bench protocols rather than datasheet peaks.

Frozen prediction: By 2027-12-31 no group other than the proposing author publishes an HLAS/HEE measurement of a commercial humanoid joint (P=0.7).

single-lab / plausibleMarker-based optical tactile sensing (TacTip class)sensing-tactile

One lineage (Bristol, ~10 years, spin-off) plus independent variants (DenseTact, GelTip) but no cross-lab reproduction of the same sensor with unit-to-unit variance. Modeled.

★ For the kids: A soft fingertip with dots inside; a tiny camera watches the dots wobble.

Technical: Marker displacement fields at 30-90 Hz; pose/contact inference learned per sensor; transfer between physical units unresolved.

Frozen prediction: By 2027-06-30 at least one non-Bristol lab reports a TacTip-class replication with per-unit force calibration error across >=3 units (P=0.40).

single-lab / plausibleStretchable optical waveguide proprioception and touchsensing-tactile

Clean mechanism (Zhao et al. Sci. Robot. 2016; Bai et al. Science 2020) but essentially one lab (Shepherd, Cornell); no closed-loop manipulation result with a referee; no long-cycle drift data. Modeled.

★ For the kids: A bendy rubber rope of light: squeeze it and less light gets through.

Technical: Intensity-loss and multi-core dyed lightguides; elastomer hysteresis/creep bound repeatability.

Frozen prediction: By 2028-01-01 no multi-lab benchmark uses lightguide sensing in a closed-loop manipulation task with n>=50 trials (P=0.80).

early / contestedTriboelectric (TENG) tactile sensorssensing-tactile

Thousands of device papers; zero cross-lab robot-manipulation reproduction, no closed-loop control, no unit variance. Publication volume is not evidence. Open.

★ For the kids: Rubbing makes a little static spark that notices a touch, but cannot tell how hard you keep pressing.

Technical: AC transient on touch/release; static force needs another modality; humidity-dependent.

Frozen prediction: By 2028-06-30 zero peer-reviewed closed-loop grasp results with TENG taxels from >=2 independent labs at n>=50 (P=0.85).

single-lab / plausibleIontronic flexible pressure sensingsensing-tactile

EDL-capacitance device class reproduced across several materials groups (UC Davis 2015 onward) and small-scale commercialised (Tacterion); no cross-lab robot manipulation study. Modeled as a device, not a robot capability.

★ For the kids: A squishy salty-jelly pad whose electric fullness changes when pressed.

Technical: nF-range capacitance eases readout; gel drying, temperature dependence and hysteresis open.

Frozen prediction: By 2027-12-31 a shipping gripper/hand from a vendor outside the originating labs uses iontronic taxels with a public datasheet (P=0.30).

early / contestedHuman tactile performance as the referee spec for robot touchsensing-tactile

DEMOTED modeled -> open by skeptic. The human psychophysics numbers (Johansson & Vallbo 1979; Johansson & Flanagan 2009) are reproduced, but the COMPONENT is the mapping from those numbers to a robot spec, which nobody has adopted or validated. Same evidence shape as HLAS (a proposed referee with zero adopters), which is graded open; consistency requires open here.

★ For the kids: Before saying a robot hand feels well, compare it to your own hand — but nobody has agreed to use that comparison yet.

Technical: Referee fields: afferent density, slip-to-correction latency (~60-80 ms), spatial acuity (~2-3 mm), dynamic range. Robot fingertips beat human local resolution and lose on coverage and loop latency.

Frozen prediction: By 2027-12-31 fewer than 10% of tactile-manipulation papers at ICRA/RSS/CoRL report closed-loop touch-to-actuation latency in ms (P=0.75).

single-lab / plausibleIntrinsic contact sensing (localizing contact from joint torque / F-T)sensing-tactile

Classical theory (Bicchi 1993) plus later single-lab implementations (Manuelli & Tedrake 2016; Haddadin et al. survey 2017). Localization accuracy never benchmarked on a shared protocol across labs. Modeled.

★ For the kids: Even without skin, a robot can guess where it was poked from which joints got pushed.

Technical: Single-contact rigid-link assumptions dominate; cm-level accuracy, single lab each.

Frozen prediction: By 2028-01-01 no two independent labs report contact-localization error on the same public protocol (P=0.70).

single-lab / plausibleTactile hardness and compliance estimationsensing-tactile

Demonstrated in several labs (MIT GelSight ICRA 2017; TacTip/DIGIT follow-ups) on private object sets; no shared benchmark; no cross-unit transfer without retraining. Modeled.

★ For the kids: Press a fingertip on something and watch how much it flattens.

Technical: Contact-area growth vs indentation; learned regressors give few-Shore-unit error on held-out lab objects.

Frozen prediction: By 2027-12-31 no multi-lab study reports hardness error on a shared published object set (P=0.65).

early / contestedMarginal utility of finger countmanipulation & dexterity

Negative-space question: no peer-reviewed study holds policy class, data budget and task suite fixed and varies only the end-effector. Forty years of hand design argue for simplicity (Bicchi 2000; Piazza et al. 2019). Open.

★ For the kids: Nobody has fairly tested whether five fingers do the job better than a two-finger clamp.

Technical: Reproduced learned-manipulation results are parallel-jaw/suction; multi-finger results are constrained (fixed-axis rotation).

Frozen prediction: By 2027-08-25 no peer-reviewed paper shows a five-finger hand beating a parallel-jaw gripper by >20 pp on a shared household suite with identical policy and demo budget, reproduced in two labs.

single-lab / plausible3D point-cloud policy representations (DP3 class)manipulation & dexterity

Strong single-origin result (Ze et al. RSS 2024) with many re-users; no independent controlled 2D-vs-3D replication on a shared physical benchmark. Modeled.

★ For the kids: The robot sees a cloud of 3D dots instead of a flat photo; other labs still need to double-check.

Technical: Reported ~+24 pp over image DP; inherits depth-camera transparent/specular blind spots.

Frozen prediction: By 2027-06-30 two author-independent labs publish a controlled 2D-vs-3D comparison; the 3D advantage shrinks below 10 pp under viewpoint/lighting randomization.

single-lab / plausibleSE(3)/SE(2)-equivariant manipulation policiesmanipulation & dexterity

Two groups on different formulations (Northeastern CoRL 2024; KAIST CVPR 2024), sim-heavy, no shared real benchmark. Modeled.

★ For the kids: If the robot knows turning the table does not change the task, it needs far fewer lessons.

Technical: Group-equivariant networks or score fields; 5-20 demo regimes reported to match 100-demo baselines.

single-lab / plausibleBilateral (force-feedback) teleoperation as a demonstration sourcemanipulation & dexterity

Control theory mature (Lawrence 1993; Hokayem & Spong 2006); as a learning-data source it is single-lab (Keio, RA-L 2020) while all reproduced low-cost pipelines are unilateral. Modeled.

★ For the kids: When the operator can feel what the robot touches, the lesson includes how hard to push.

Technical: Four-channel bilateral trades transparency vs stability under latency.

Frozen prediction: By 2027-12-31 no cross-lab reproduction shows bilateral demos beating unilateral by >10 pp on contact-rich insertion at equal demo count.

single-lab / plausibleHuman-in-the-loop intervention learning (HG-DAgger / Sirius class)manipulation & dexterity

Same mechanism in three labs (Michigan 2019; Berkeley 2021; UT Austin 2023) each on its own tasks with n~20; no shared benchmark. Modeled, not viable.

★ For the kids: A person grabs the controls when the robot slips, and the robot learns from those moments.

Technical: Operator takeover segments up-weighted in training; monotone improvement over deployment rounds reported.

reproduced / credibleKinesthetic lead-through teachingmanipulation & dexterity

Reproduced across every major cobot vendor and decades of PbD literature; fleet-scale industrial adoption. Records positions only. Viable with a narrow ceiling.

★ For the kids: You move the robot's arm through the job once; it replays the motion.

Technical: Needs gravity compensation; captures paths, not forces or hand poses.

early / contestedBenchmark-score to real-robot-success transfer (missing correlation)evaluation

No published study correlates embodied-VLM benchmark rank with third-party real-robot rank across >=5 models. Only in-paper bridge is RynnBrain 1.1's own VLA comparison. Open.

★ For the kids: A spelling-test score does not tell you who can bake a cake.

Technical: Required: >=6 open brains, identical VLA head and demos, distributed pairwise referee; report Spearman.

Frozen prediction: By 2027-08-25 no study reports benchmark-vs-real Spearman over >=6 embodied brains evaluated by non-authors (P=0.7).

single-lab / plausibleCompositional generalization from atomic skills (ATOM-Bench)embodied AI

ATOM-Bench (arXiv 2606.16826, verified: 30 atomic + 24 held-out compositional tasks, 3,000 demos, 2,700 rollouts, 5 policies) is the first adequately powered physical measurement of the recombination wall. Single lab (BAAI FlagEval); not reproduced on a second site. Modeled.

★ For the kids: A robot learns two tricks separately, but asking for both together in a new order often breaks it.

Technical: Metrics separate motor-execution, instruction-grounding and compositional-reuse failure.

Frozen prediction: By 2027-08-26 no non-BAAI lab reports >50% mean success on held-out compositional tasks while >80% on atomic (P=0.65).

single-lab / plausibleHybrid variable-reluctance (HVR) force actuators (magnet-biased reluctance, near-zero holding power) — MERGED with reluctance-actuator-deformable-mirrorsactuation & power

SKEPTIC: two lenses filed the same subject under two slugs (hybrid-variable-reluctance-actuators, reluctance-actuator-deformable-mirrors); merged here, grade modeled held. One vendor (TNO), one telescope on-sky (IRTF-ASM-1, 36 actuators, closed loop since April 2024, 'no hardware issues' per the July 2026 progress report), academic partners UH/UCSC. The ~75x efficiency-vs-voice-coil figure is developer/partner stated (Bowens-Rubin et al. 2021, UCSC/TNO) with no third-party actuator metrology. Robotics relevance is an analogy until study hvr-fine-positioning-rival-test runs. Single-vendor + si…

★ For the kids: A magnet does the holding and a small coil only nudges; a Hawaii telescope has used 36 of these since 2024. One company makes them, so we say 'promising', not 'proven'.

Technical: Kuiper et al., SPIE 13097 (2024) / arXiv:2407.11289; Lee et al., SPIE 2024, arXiv:2407.06444; Chun, Lai, Hinz et al., SPIE 13100 (2024); Lee et al., arXiv:2607.04385 and arXiv:2608.00373 (2026); Bowens-Rubin, Dillon, Hinz & Kuiper, arXiv:2110.01693 (2021). Pantry lead 7544427e3b66 chased to these primary sources by three lenses; lens 4 could not find it in the rotated feed.

Frozen prediction: By 2028-12-31 an HVR actuator array is reported in hardware in a second mirror or robotic fine-positioning stage not built by TNO. p=0.35.

early / contestedRobot battery-pack safety certification (UL 2271 / IEC 62133-2 / UN 38.3)actuation & power

Held open. The standards exist; no humanoid vendor spec sheet surveyed names a pack certification (Unitree G1 421.2 Wh / 13S / 46.8 V spec names only a proprietary BMS). Watch-item, census-driven.

★ For the kids: Big batteries need a safety stamp; we could not find a walking robot that says it has one.

Technical: UN 38.3; IEC 62133-2:2017; UL 2271; ISO 13482:2014; UL 3300 Outline (2021). Study pack-certification-census is the instrument.

Frozen prediction: By 2027-12-31 at least one top-10 humanoid vendor publicly names UL 2271 or IEC 62133-2 for its pack. p=0.55.

early / contestedHotel load: the non-actuation power budget (compute + sensing + drive quiescent)actuation & power

Held open. Bounds only from vendor figures (Jetson AGX Orin 15-60 W envelope; G1 ~210 W implied mean; Walker S2 2 h walking / 4 h standing). Peer-reviewed component breakdowns exist only for a small servo humanoid (IEEE 2015) and a welcome robot (IEEE 2025). No shunt-measured breakdown of a commercial-class humanoid.

★ For the kids: A robot burns electricity just thinking and looking; nobody has measured that bill on the robots in the videos.

Technical: Seok et al. TMech 2015; Bledt et al. IROS 2018; Katz et al. ICRA 2019; IEEE Xplore 7244843 (2015); IEEE Xplore 11100596 (2025). Bar to modeled: one published three-state shunt measurement (study hotel-load-audit-protocol).

Frozen prediction: By 2027-09-01 a paper/preprint reports non-actuation load >= 25% of mean total power for a commercial-class humanoid walking <= 1 m/s. p=0.60.

single-lab / plausibleEnergy-aware locomotion policy learning (cost of transport as the reward)actuation & power

Held modeled. Method reproduced across labs (Berkeley/CMU CoRL 2021 on A1; NTUA Laelaps II 2022/2025; ETH 2025), but each reports its own baseline on its own platform; hardware-measured CoT against a matched non-energy-aware policy is rarely reported. Ubiquitous but ungraded.

★ For the kids: Give a robot points for saving battery and it learns thriftier walking; the labs grade themselves so we do not know how much it saves.

Technical: Fu, Kumar, Malik, Pathak, CoRL 2021; Laelaps II, Springer LNNS 2022 (978-3-031-15226-9_21); ETH arXiv:2509.10128; arXiv:2509.01765.

Frozen prediction: By 2028-06-30 two independent labs report hardware-measured CoT reduction >= 20% vs a matched baseline on the same platform. p=0.50.

single-lab / plausibleCarbon-nanotube yarn and sheath-run electrochemical artificial musclesactuation & power

Held modeled. Material physics reproduced beyond UT Dallas (Nano Research 2022; Adv. Fiber Mater. 2026), no robot joint driven beyond bench demos; cycle life under load and electrolyte packaging unreported.

★ For the kids: Twisted threads that shrink with a little electricity; several labs make them, none has moved a robot arm for long.

Technical: Lima et al., Science 338:928 (2012); Mu et al., Science 365:150 (2019); Nano Research 15 (2022); Adv. Fiber Mater. (2026) doi 10.1007/s42765-026-00705-2.

Frozen prediction: By 2028-12-31 no CNT-yarn/sheath-run electrochemical muscle drives an untethered robot joint for >= 1e5 cycles at >= 2% stroke in a lab independent of UT Dallas. p=0.75.

single-lab / plausibleRobot olfaction & gas-source localizationsensing & tactile

Held modeled. 25-year multi-lab literature (Örebro, TUAT, Swansea, IBEC) reproduces controlled-environment source localization; field trials single-site; no shipped product performs autonomous source-seeking.

★ For the kids: Robots can sniff toward a gas leak in the lab; none you can buy does it on its own.

Technical: Ishida et al., IEEE Sensors J. 2012; Francis et al., J. Field Robotics 39(8) 2022; T-RO 2024 doi 10.1109/TRO.2024.3426368.

Frozen prediction: By 2027-09-05 no commercially sold ground or aerial robot documents autonomous gas-source localization as a product function. p=0.80.

early / contestedTactile sensor unit cost (self-reported BOMs vs one verified list price)sensing & tactile

Held open. One verified list price (GelSight Mini $499, 2022); all 'low-cost' figures are author-reported BOMs without yield, labor or gel replacement; Digit 360 has no public price.

★ For the kids: Nobody outside the makers has checked what a robot fingertip really costs to keep working.

Technical: Per-taxel-per-year cost (sensor + consumable + recalibration labor) is the number that matters; no paper reports it.

Frozen prediction: By 2027-09-05 a vendor other than GelSight Inc. lists a camera-behind-elastomer fingertip under $200/unit. p=0.40.

single-lab / plausibleFusion in-vessel remote handling (JET bilateral teleoperated maintenance) — DEMOTED viable -> modeledmanipulation & dexterity / deployed teleoperation

SKEPTIC DEMOTION. Full in-vessel divertor/first-wall replacement under a campaign-restart referee has been done at ONE facility (JET). The cross-facility legs offered (ITER prototypes at DTP2 Tampere, Sellafield glovebox robotics) are test rigs or a different task class, not reproductions of fusion in-vessel campaigns; ITER has not operated. The 30,000 h figure comes from the operator's own staff (Buckingham & Loving are RACE/UKAEA) and could not be verified from the Nature Physics article in this pass. The broader class — nuclear master-slave teleoperation in hot cells — is reproduced worldw…

★ For the kids: Inside one fusion machine, humans drove robot arms to swap heavy parts for decades. Only one machine has done it, so it is 'promising', not 'proven everywhere'.

Technical: Buckingham & Loving, Nature Physics 12:391 (2016) — commentary by the operator; Rolfe et al., Fusion Eng. Des. 1999. Bilateral Mascot master-slave on 8 m booms. Re-promotion: a second fusion facility (JT-60SA, ITER) completes an in-vessel remote-handling campaign, or a third-party audit of JET's hours ledger.

Frozen prediction: By 2027-09-05 no humanoid vendor publishes a third-party-verified operating-hours figure for a single fleet exceeding 30,000 h. p=0.85.

reproduced / credibleSelf-driving-lab labware manipulation (chemistry-outcome referee, with a documented referee failure)manipulation & dexterity / structured-environment autonomy

Viable HELD, scope narrowed: structured labware handling across multi-day unattended runs is reproduced at LBNL (A-Lab, Nature 2023) and Liverpool (Nature 2020, 2024), with Aalborg and a Japanese group (IJIRA 2025) reproducing the mobile-manipulator hardware pattern. The viable grade covers manipulation only; closed-loop autonomous discovery is single-lab and lives under mobile-manipulator-lab-automation (modeled). A-Lab's novelty claims were challenged (PRX Energy 2024) and corrected (Nature, Jan 2026): grade the hand and the claim separately.

★ For the kids: Robots run chemistry for days and a real test grades the powder; once the robot did fine but the scientists' claim had to be corrected.

Technical: Szymanski et al., Nature 624 (2023, corrected 2026); Leeman et al., PRX Energy 3:011002 (2024); Burger et al., Nature 583 (2020); Dai et al., Nature 635 (2024).

Frozen prediction: By 2027-12-31 at least one further Nature/Science-family self-driving-lab synthesis paper carries a correction or peer-reviewed critique of its novelty or phase-identification claims. p=0.45.

reproduced / credibleSemiconductor wafer-handling robots under SEMI E10 (the reliability referee that already exists)manipulation & dexterity / reliability metrology

Viable HELD for the narrow claim: a named reliability standard (SEMI E10, since 1986) exists and is used contractually across multiple vendors and fabs. No vendor MTBF figure is quoted or endorsed; the lens verified none. The claim is about the standard, not a number.

★ For the kids: Chip-factory robots must report how often they break under one shared rulebook; humanoids have no rulebook.

Technical: SEMI E10 state model and MTBF/MCBF/MTTR with uncertainty; SEMI S28 robot safety.

Frozen prediction: By 2027-09-05 no humanoid vendor publishes MTBF or MCBF for a fleet of >= 50 units under SEMI E10 or another named standard. p=0.85.

single-lab / plausibleOn-orbit robotic servicing manipulation (RSGS/MRV era)manipulation & dexterity / space

Held modeled. SKEPTIC verified: RSGS payload on SpaceLogistics' MRV launched 2026-07-21 (DARPA news), ~10-month electric-propulsion transfer, service work expected 2027. One flight, no on-orbit dexterous operation yet; MEV-1/2 docked but did not manipulate.

★ For the kids: A robot with arms just launched to fix satellites; it has not fixed one yet.

Technical: Twin NRL-developed dexterous arms with tool changers on a GEO vehicle.

Frozen prediction: By 2027-12-31 RSGS/MRV completes at least one publicly confirmed robotic-arm contact operation on a client spacecraft in GEO. p=0.50.

early / contestedPolicy inference energy per completed task (joules-per-success)embodied AI / onboard compute x power

Held open. No paper reports absolute joules per successful episode on a named board and named robot; EcoVLA reports relative efficiency only.

★ For the kids: Robots have a battery for moving and one for thinking; nobody has written down the thinking bill per chore.

Technical: EcoVLA arXiv 2608.15502; Zhou et al. arXiv 2604.24447; Jetson-PI arXiv 2607.12659.

Frozen prediction: By 2027-06-30 no more than two papers report absolute joules per successful episode for a VLA on a named onboard board AND a named physical robot. p=0.70.

early / contestedLearned-policy robustness to joint-level physical faults (actuator wear as distribution shift)embodied AI x actuator durability

Held open. Jo et al. 2026 (single lab, sim/real split unstated) and UZH quadrotor adaptation (real, but a control policy not a VLA). Zero physical-humanoid rows.

★ For the kids: A robot that learned with a new elbow may fail when the elbow gets stiff; nobody has checked.

Technical: arXiv 2606.10501; arXiv 2606.27353.

Frozen prediction: By 2027-06-30 no study re-evaluates a VLA on the same physical robot after >= 90 days or >= 500 h with per-week success tracked. p=0.70.

early / contestedVLA weight-access tiers (open weights vs trusted-tester gating)embodied AI / reproduction substrate

Held open; census-type domain. Only tier A/B releases can ever reach viable in this ledger.

★ For the kids: Some robot brains are given away so others can check them; some are locked. Only the first kind can be proven.

Technical: Gemini Robotics On-Device 2 model card 2026-07-30 (trusted testers only); GR00T N1.6, openpi, Embodied-R1.5 open.

Frozen prediction: Gemini Robotics On-Device weights remain non-downloadable to the public through 2027-03-31. p=0.75.

single-lab / plausibleRegulator-mandated field incident reporting (the referee autonomy got and manipulation never did)adversarial skeptic

Held modeled. The mechanism is reproduced across three regulators (CA DMV, NHTSA SGO 2021-01, FDA MAUDE) and independently analysed (DSN 2018; PLoS ONE 2016; Sci. Data 2021); nothing transfers it to workplace robots; the referee is gameable via self-defined disengagements.

★ For the kids: Self-driving cars and surgical robots must report every mishap; humanoids do not, so nobody can count.

Technical: Banerjee et al. DSN 2018; Alemzadeh et al. PLoS ONE 2016; Sinha et al. Sci. Data 8:298 (2021); ISO 10218:2025 has no reporting leg.

Frozen prediction: Through 2027-12-31 no US federal or state regulator issues a standing incident-reporting order covering humanoid or mobile-manipulation robots in workplaces. p=0.85.

early / contestedPost-demo venture mortality (the survivorship denominator behind every showreel)adversarial skeptic

Held open. Verified cases (Rethink, Anki, Jibo, Kuri) are trade-press post-mortems; no denominator exists. Use as a prior: a demo moves a maturity grade by nothing.

★ For the kids: Many robot companies made amazing videos and then closed; ask how often videos like it came from survivors.

Technical: The Robot Report 2018; Hoffman, IEEE Spectrum 2019-05-01.

Frozen prediction: By 2028-06-30 at least two humanoid startups that raised >= $50M and published a demo video by 2025-12-31 cease operations, are acquired, or publicly exit humanoids. p=0.70 (base-rate guess, stated as such).

single-lab / plausibleScenario-based validation campaigns (enumerated conditions x repetitions as the referee)adversarial skeptic

Held modeled. RoboVAST (arXiv 2607.06248): 5,480 configurations x 20 repetitions, simulation-only, navigation-only, one lineage; practice reproduced in automotive SOTIF.

★ For the kids: Instead of one nice video, test thousands of situations twenty times each and count what breaks.

Technical: arXiv 2607.06248; 2604.25772; 2605.29973.

Frozen prediction: Through 2027-12-31 no humanoid or mobile-manipulation vendor publishes a scenario campaign with >= 1,000 configurations and >= 10 repetitions each with failure classes enumerated. p=0.80.

reproduced / credibleRobotic milking (fleet-scale contact manipulation of a live deformable target)manipulation

Viable HELD: ~50,000 units from at least three vendors (Lely, DeLaval, GEA), thirty years deployed, and peer-reviewed per-attempt failure rates (7.6% of 35,291 milkings, Bach & Busto 2005). This is the strongest viable in the set — reproduced across vendors, farms and countries with an outcome referee. Open sub-question: current attachment-failure rate on 3D-camera systems.

★ For the kids: For 30 years robots have milked cows by themselves, and farmers count every miss.

Technical: Jacobs & Siegford, J. Dairy Sci. 95:2227 (2012); Rodenburg, J. Dairy Sci. 100:7729 (2017); Bach & Busto, J. Dairy Res. 72:101 (2005); Cogato et al., Animals 2021; Lage et al., Animals 2024.

Frozen prediction: Through 2028-12-31 no commercially sold automatic milking system lists a learned end-to-end policy as its teat-cup attachment controller. p=0.85 (skeptic-assigned; proposing lens froze none).

reproduced / credibleCable-driven parallel robots (large-workspace precision by winches)actuation

Viable HELD: reproduced across entertainment (Skycam), industry (Fraunhofer IPAnema), science (FAST 30 t feed cabin) and academia (LIRMM CoGiRo). Sub-mm figures are metrology-assisted and must be reported with the sensing stack.

★ For the kids: Robots hanging from motor-pulled ropes can move a camera or a 30-ton cabin very precisely.

Technical: Pott, Springer 2018; Guillory, Gouttefarde et al., Precision Engineering 97 (2026); Jiang et al., Engineering 28:21 (2023); Yao et al., RAA 20:68 (2020).

Frozen prediction: Through 2027-12-31 no published CDPR result reports <= 100 um absolute positioning over a >= 10 m workspace without external optical metrology closing the loop. p=0.80 (skeptic-assigned; proposing lens froze none).

single-lab / plausibleSoft continuum manipulator arms (tentacle-class)manipulation

Held modeled. Cosserat-rod modelling reproduced across labs (ETH 2022, UIUC 2026); real-task control single-lab; no benchmark.

★ For the kids: Rubber tentacle arms filled with air; computers now predict how they bend.

Technical: Fischer et al., Adv. Intell. Syst. 2022 (arXiv 2201.02151); Kim et al., Adv. Intell. Syst. 2026 (arXiv 2507.10121).

Frozen prediction: Through 2028-12-31 no peer-reviewed result reports a soft continuum arm as sole manipulator completing a NIST task board or YCB protocol with a stated success rate over >= 50 trials. p=0.80 (skeptic-assigned; proposing lens froze none).

single-lab / plausibleMobile manipulators in self-driving laboratories (closed-loop leg)embodied AI

Held modeled. Hardware pattern reproduced (Liverpool, Aalborg, Japan); closed-loop autonomous discovery remains one lab (Cooper group). Sibling of self-driving-lab-manipulation (viable for labware handling only).

★ For the kids: Wheeled robots with an arm roll around a chemistry lab at night; only one lab lets the robot decide what to try next.

Technical: Dai et al., Nature 635 (2024); Sulaiman et al., IJIRA 2025 (arXiv 2510.19081); Sasaki et al., IJIRA 2025 (arXiv 2506.11384).

Frozen prediction: By 2027-12-31 at least one group outside Liverpool publishes a >= 48 h unattended closed-loop mobile-manipulator chemistry campaign in a peer-reviewed venue. p=0.40 (skeptic-assigned; proposing lens froze none).

Studies

single-lab / plausibleWhich humanoid capabilities are actually reproduced, not just demoed?

Q: For each robot skill, is there an independently reproduced result, or only a single choreographed demo?

Legged locomotion clears both bars (viable). VLA foundation models and electrohydraulic muscles clear one (modeled). General dexterous in-hand manipulation and all-day power clear neither (open).

Verdict: The deployment gap is not uniform: mobility is largely solved, but hands, generalist policies, and endurance are the real gates. Humanoid full-body demos borrow the credibility of solved locomotion to imply unsolved manipulation.

early / contestedRanking the picks-and-shovels bottleneck

Q: Of actuators, sensors, compute, and power, which component most limits mass humanoid deployment today?

Compute and actuators are the most mature (viable-ish); tactile sensing is close (viable but integration-limited); dexterous control software and power/endurance are the binding constraints (open).

Verdict: The next capability inflection most likely comes from manipulation-capable foundation models plus tactile integration, not from a new actuator; endurance is a parallel hard gate.

single-lab / plausibleDo any artificial muscles beat electromagnetic drives for humanoid joints yet?

Q: Across HASEL, TCP, DEA, and pneumatic McKibben muscles, does any soft/artificial actuator clear the deployment bar (efficiency, bandwidth, force density, cycle life, driver simplicity) that electromagnetic QDD + strain-wave/roller-screw already clear?

Every artificial-muscle line reproduces in labs but fails at least one deployment gate: TCP (~1-3% efficiency, thermal bandwidth), DEA (kV drive, breakdown), HASEL (kV drive, cycle life), McKibben (compressor mass, control). Electromagnetic + precision-reducer/roller-screw remains the only reproduced deployment-grade path.

Verdict: The 'artificial muscle will replace motors' narrative is over-tiered: no artificial muscle is viable for a humanoid joint. They stay modeled/open; the shipping stack is still electromagnetic.

early / contestedWhat gates humanoid endurance — battery energy or actuator heat?

Q: Is untethered humanoid work-time limited first by pack energy density (Wh/kg) or by actuator thermal derating under sustained torque?

Commercial Li-ion cells sit ~250-300 Wh/kg and have plateaued (Schmuch et al., Nature Energy 2018); solid-state is not yet in robots (Janek & Zeier, Nature Energy 2016). Continuous-torque thermal limits are well-modeled, but field data to say which binds first is missing.

Verdict: Endurance is a two-headed OPEN gate. Battery-swap logistics may win before either ceiling is lifted; solid-state remains a watch-item, not a solution.

single-lab / plausibleWhich tactile modality is actually reproduced, and on what axis?

Q: Across vision-based (retrographic), magnetic-taxel, and large-area e-skin, which capability is reproduced across labs, and on which axis (spatial resolution / force resolution / durability-replaceability / coverage area / integration cost)?

Vision-based leads on spatial resolution and reproduction (viable as a principle). Magnetic-taxel leads on thinness, 3-axis force and replaceability (modeled). Large-area e-skin leads on potential coverage but is unsolved at robot scale (open). No single modality dominates every axis.

Verdict: 'Tactile sensing' is not one bottleneck — it is a Pareto front. The deployment question is which axis a given task needs, and different sensing suppliers can own different axes.

early / contestedIs tactile sensing a reproduced manipulation advantage, or an assumed one?

Q: When tactile sensing is added to a manipulation policy, is the improvement reproduced across labs and sensors, or is it lab- and sensor-specific?

Local contact events — slip detection and contact-rich insertion — show fairly robust, reproduced touch benefit. General dexterous in-hand manipulation benefit from touch is demonstrated (Qi et al., CoRL 2023) but not yet reproduced across sensors/labs at scale. Touch foundation models (Sparsh) show gains only on their originating-lab benchmark.

Verdict: Touch pays off reproducibly for LOCAL contact events (slip, insertion) but its contribution to GENERAL dexterity is still modeled/open. Do not let a solved slip-detection demo imply solved dexterity.

single-lab / plausibleIn-hand reorientation: what is reproduced vs what is still open

Q: Has 'dexterous in-hand manipulation' actually been reproduced across labs, or is each result a single-lab showpiece?

The narrow claim clears the reproduction bar: OpenAI (2019) -> NVIDIA DeXtreme (ICRA 2023) -> MIT Visual Dexterity (Science Robotics 2023) -> Berkeley/Meta Hora/RotateIt (CoRL 2022/23). The general claim clears neither reproduction nor out-of-lab robustness.

Verdict: Grade the DOMAIN open, but explicitly credit the reproduced narrow slice. Humanoid showreels imply the general capability by borrowing the credibility of the narrow one.

early / contestedWhere the manipulation stack actually binds: hands, touch, data, or generalization

Q: Of hand hardware, tactile sensing, teleop data, and policy generalization, which most limits deployable dexterous manipulation today?

Hand hardware is de-bottlenecked (LEAP-class, reproduced) and tactile sensing is close (GelSight/DIGIT commercial, integration-limited). The binding constraints are (a) GENERALIZATION of learned policies beyond the demonstrated distribution and (b) DATA — imitation learning is teleop-data-bound.

Verdict: The next manipulation inflection comes from data-efficiency + generalization (UMI, cross-embodiment pretraining), not a new hand or actuator. Endurance/power is a parallel gate owned by other domains.

single-lab / plausibleDoes adding touch actually help in-hand manipulation, or is it decoration?

Q: Is vision-plus-touch a reproduced improvement over vision-only for tracking and manipulating objects in-hand?

NeuralFeels (Meta, Science Robotics 2024) shows touch refines and disambiguates visual pose/shape estimates precisely when the object is occluded by the hand, and releases the FeelSight benchmark. Single-lab but rigorous; grade modeled pending independent replication.

Verdict: Tactile is not decoration for in-hand work: it earns its place under occlusion. But the result is one lab; treat 'touch is necessary for dexterity' as modeled, not settled, until reproduced.

single-lab / plausibleWhat in embodied AI is reproduced vs a single-lab flagship demo?

Q: Within VLA / robot foundation models / sim-to-real, which results are independently reproduced and which rest on one lab's flagship showreel?

Reproduced: narrow imitation (diffusion/ACT), GPU-parallel sim-to-real for locomotion, and open-weight cross-embodiment VLA transfer (OpenVLA/Octo). Single-lab / not reproduced: the 2025 generalist humanoid flagships (Gemini Robotics, pi-0.5, GR00T N1) and end-to-end humanoid autonomy showreels.

Verdict: The reproduced embodied-AI capability is NARROW — per-task imitation, sim-to-real locomotion, and transfer of learned representations. 'Generalist embodied intelligence' is modeled at best.

early / contestedIs embodied AI gated by data or by compute?

Q: Does VLA capability keep scaling with data/compute the way LLMs did, or does the physical world impose a data ceiling language models never faced?

Sim/compute scales cheaply via GPU simulation but only helps where a simulator is faithful (locomotion, rigid contact). Real-world manipulation reliability is teleoperation-data-gated and expensive; representation pretraining only partially substitutes for real robot data.

Verdict: The binding input for GENERALIST manipulation is real robot demonstration data, not compute. That reprices teleoperation-data collection as the true near-term bottleneck. Open until a reproduced scaling law for real-world manipulation exists.

single-lab / plausibleHow much headline humanoid capability is actually autonomous?

Q: For each viral humanoid manipulation demo, is the behavior autonomous, teleoperated, human-in-the-loop, or edited/sped-up — and is that provenance disclosed?

Teleoperation as a data-collection/control method is genuinely reproduced (Mobile ALOHA, UMI, robomimic). Unsupervised autonomous execution of varied manipulation on hardware is largely undisclosed or sped-up in commercial reels; where disclosed, intervention rates are high. Locomotion remains the only body-scale capability that clears an autonomy bar across labs.

Verdict: The demo-theater tell is a missing intervention-rate number plus edited playback speed. Until a third-party autonomous-fraction audit exists, 'the robot did this' should be read as 'a human helped the robot do this.'

single-lab / plausibleThe honest humanoid deployment ledger: narrow pilots vs general labor

Q: What is actually deployed for pay, on what task breadth, at what audited reliability and cost — versus what is merely demonstrated?

Narrow, structured tasks (tote/bin moving) are in single-vendor paid pilots (Agility Digit/GXO, 2024) — modeled, not reproduced across vendors or economically audited. General dexterous labor is open: no audited varied-task autonomous deployment exists.

Verdict: Deployment is real but narrow and single-vendor; the industry's implied jump from tote-moving to general household/factory labor is unsupported by reproduced evidence.

single-lab / plausibleIs the manipulation bottleneck hardware or demonstration data?

Q: Where did recent real-robot manipulation capability gains actually come from — new actuators/sensors, or cheaper data plus better imitation learning?

The reproduced gains rode on cheap demonstration data (ALOHA/UMI, reproduced across many labs) and generative imitation policies (Diffusion Policy, reproduced as a method) on adequate existing hardware — not on a new actuator or sensor breakthrough.

Verdict: For the next manipulation inflection, research and capital leverage points at demonstration-data infrastructure and generative policies, not at a new joint actuator.

reproduced / credibleWhat actually made sim-to-real locomotion 'solved'?

Q: Which single ingredient turned legged locomotion from a demo into a reproduced, deployment-grade capability?

The common substrate is GPU-parallel simulation plus domain randomization plus learned actuator modeling, reproduced independently across ETH, NVIDIA, and many downstream labs. Locomotion clears both the 'reproduced' and 'survives outside the lab' bars.

Verdict: The compute/sim layer is the reproduced picks-and-shovels behind 'solved' locomotion. The same substrate is now pointed at manipulation but has NOT yet delivered reproduced dexterity — so locomotion success should not be borrowed to imply manipulation is solved.

single-lab / plausibleWhich endurance path wins: Wh/kg, swap, or fast charge?
single-lab / plausibleArtificial-muscle scorecard: does any soft actuator clear the humanoid joint bar?
single-lab / plausibleWhich tactile modalities are reproduced?
early / contestedThe unpriced consumable: do soft tactile skins survive industrial duty cycles?
single-lab / plausibleIs in-hand reorientation reproduced, or still one lab's trick?
single-lab / plausibleDemo provenance: how much 2024-2026 humanoid dexterity is teleoperated or unverified?
single-lab / plausibleWhat has actually reproduced in VLA-land (2023-2026)
reproduced / credibleSim-to-real humanoid whole-body control: a reproduction census
early / contestedDemo-provenance audit: teleop backstops behind 2025's humanoid moments
early / contestedDemo-provenance forensics protocol
single-lab / plausibleReproduction census of flagship VLA claims
early / contestedIntervention-rate ledger (freezing the vacuum)
single-lab / plausibleWheels-plus-legs vs pure bipeds: where is the verified deployment evidence?
single-lab / plausibleThe humanoid cost collapse: what does $6k actually buy?
early / contestedIs a $2 microphone the poor man's tactile skin?
single-lab / plausibleThe actuation stack is commoditizing from the bottom
single-lab / plausibleEndurance triage: chemistry vs thermal headroom vs logistics
early / contestedDoes adding touch produce reproduced gains?
single-lab / plausibleIs in-hand reorientation reproduced? Lab-by-lab scorecard
single-lab / plausibleSim-inflation audit: grasp synthesis demotes to open
early / contestedThe wear-item gap: nobody publishes hand MTBF
single-lab / plausibleVLA reproduction gap: self-reported vs third-party
early / contestedDemo-provenance ledger for humanoid claims
early / contestedReplication watch: real-robot RL post-training
single-lab / plausibleDo vendor production forecasts count as deployment evidence? (No.)
early / contestedFleet security: the unpriced deployment gate
single-lab / plausibleLoco-manipulation: reproduced direction, unreproduced results
single-lab / plausibleDo road-not-taken actuation lines threaten the incumbents?
early / contestedIs learning-from-watching reproduced?
reproduced / credibleHow often does an independence claim survive checking the actual author list?

Q: This corpus froze a rule one lens ago: independence counts PEOPLE and advisory lineages, never institution names, logos or country flags. How many of the submitted independence claims survive when the rule is applied to the verified author lists rather than to the citation lines?

Three failures found in one pass, all in the same direction. (1) Colosseum V2 was described as 'a successor built by a different lead team'; Singh and Thomason are on both V1 and V2. (2) The grasp-metric 'reproduced negative' was counted as three legs; Rubert and Morales are on all three, Kappler and Bohg on two. (3) Sparsh-skin was already caught by the submitting lens as ReSkin lineage via Hellebrekers — that one the corpus found itself. Two additional near-misses were correctly self-declared rather than hidden: magnetorheological actuators (Sherbrooke/Exonetik is one lineage, and Plante is…

Verdict: The rule works, and the failure mode is entirely predictable: independence is asserted from the citation line (institution names, 'a different lead team') rather than checked against the author list. Every failure took under two minutes to find. The corrective is procedural, not intellectual — any sentence in this ledger claiming a second leg must name the disjoint author sets, or it does not count. Applied consistently, that single discipline would have caught all three before submission.

reproduced / credibleEvery proposed referee in this batch grades its own homework

Q: The corpus's central complaint is that robotics has no referees. Several 2025-26 instruments claim to be that referee. Who has validated the referees?

Not one. World-model evaluators: RoboWorld's r=0.989 and GigaWorld's WMBench findings are self-measured by their builders against real numbers already in hand. Colosseum: V1's R^2=0.614 and V2's 'strong correlations' both come from the same author lineage, and they disagree with each other. VLA-REPLICA's reproducibility rests on setups built inside the authoring project. RoboArena and RoboChallenge are two genuine outside referees who have never been compared to each other. AutoEval replaces the human oracle with a learned success classifier, importing the vlm-reward-success-detection error s…

Verdict: The referee layer has the same defect as the capability layer one level up, and the corpus was in the middle of committing it again by promoting world-model-policy-evaluation on self-scored correlations. The rule that fixes it is already written down in this batch and simply needs enforcing everywhere: no proxy number supports a maturity unless rank preservation was measured by someone who did not build the proxy. On that rule the qualifying count across every instrument here is zero, which is why sim-proxy-evaluation-validity and world-model-policy-evaluation both sit at open and third-party…

reproduced / credibleWhat actually cleared the bar this cycle, and what those cases have in common

Q: Stripping out every claim that failed the audit, which components genuinely satisfy demonstrated-AND-reproduced-across-independent-groups — and is there a pattern?

A short list with a striking shape. (1) Astronomy fiber positioners: 5,000 units on DESI plus independent builds at SDSS-V, Subaru PFS, 4MOST and WEAVE, externally validated by fiber dithering and astrometric calibration rather than by their own encoders. (2) Piezoelectric ultrasonic motors: independent commercial manufacture by four unrelated firms across three decades under warranty. (3) Passive-dynamic walking efficiency: three physically distinct machines, three PIs, different actuation, one agreed metric, unchallenged for 21 years. (4) Impedance and hybrid force/position control: forty y…

Verdict: Reproduction in robotics is concentrated where someone outside the lab was obliged to check — an observatory acceptance review, a warranty, a safety case, or an adversary trying to break the claim. Where the only checker is the claimant, nothing clears the bar regardless of how much activity there is. That reframes the corpus's recurring 'missing referee' complaint from an evaluation-infrastructure problem into an incentive problem: astronomy publishes bearing runout next to angular resolution because an instrument review board requires it, and the HARMONI LPOA-EM paper is the proof that the …

early / contested
early / contested
early / contested
single-lab / plausible
single-lab / plausible
early / contested
early / contested
early / contested
early / contested
early / contested
early / contested
early / contested
early / contested
early / contested
early / contested
early / contested
early / contested
early / contested
early / contested
early / contested
early / contested
early / contested
early / contested
early / contested
early / contested
early / contested
single-lab / plausible
early / contested
early / contested
early / contested
single-lab / plausible

Both dual-arm papers share Zhu, Lammers, Lu and Li — one lineage; Monash 62.8%/9.18 s at the 2014 mean; HortiBot n=24.

Verdict: Single-lineage advance; component stays open.

single-lab / plausible

battery-swap-endurance-architecture demoted; SKEPTIC extension this pass: wheeled-legged-hybrid-locomotion and third-party-policy-evaluation reversed on the same product-vs-capability rule; fusion-remote-handling demoted on single-facility.

early / contested

Every independent fleet-scale analysis rests on a regulator-mandated record; workplace robots have design standards but no reporting mandate.

Verdict: The vacuum is structural; SKP-P6 frozen.

reproduced / credible

~50,000-unit fleet; 7.6% attachment failure over 35,291 milkings (2005); ~2.5 failed/incomplete milkings per robot-day; no humanoid paper cites AMS.

Verdict: Milking robots are the reproduced existence proof; the loner agricultural-harvesting-robots is the same industry's failure case.

single-lab / plausible

Sub-mm figures are metrology-assisted; platform class reproduced across three institutions.

Verdict: Report precision with the sensing stack.

single-lab / plausible

Chased to real sources: RxBrain (2607.14187), RynnBrain 1.1 (2607.17977), Embodied-R1.5 (2606.11324), TacEvo (2606.30109), IRTF-ASM-1 (2407.06444/2407.11289/2607.04385/2608.00373), soft-arm papers (2201.02151, 2507.10121), Garrabé QD (2608.30983). Embodied-brain leads are evaluated parties' own abstract prose: zero weight. c41881cd2e0d untraceable; 666f09e66364 ENGINE-origin; 7c27af43b808 prior kill; MEDit-Bench, micro_biorobot_agent, photochromic-iris, 5d5e8caf2d4b out of lens or untraced.

Verdict: Twelve leads chased to real sources; none raised a grade; the 7c27af43b808 listing is struck from usedPantryLeads.

Frozen research claims 50 · self-assessed crew research (hash-frozen, accumulating) — not the externally graded ledger

Key literature

Hydraulically amplified self-healing electrostatic actuators with muscle-like performanceScience (Acome, Keplinger et al., Univ. of Colorado Boulder) 2018

Founding paper of HASEL artificial muscles, the leading soft-actuator alternative to electric/hydraulic drives; a bottleneck-repricing candidate if drive voltage and cost fall.

GelSight: High-Resolution Robot Tactile Sensors for Estimating Geometry and ForceSensors (Yuan, Dong & Adelson, MIT) 2017

Reference design for camera-based tactile sensing, a per-fingertip component every dexterous humanoid hand must buy; now commercial (GelSight Mini) and productized as Meta DIGIT.

General In-Hand Object Rotation with Vision and TouchCoRL (Qi, Malik et al., UC Berkeley / Meta AI) 2023

State-of-the-art on the hardest, least-solved robot skill; a marker for grading manipulation hype honestly as still 'open.'

Learning robust perceptive locomotion for quadrupedal robots in the wildScience Robotics (Miki, Hutter et al., ETH Zurich) 2022

The clearest example of a reproduced, deployment-grade robot capability; sets the 'viable' bar that manipulation and endurance have not met.

RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic ControlGoogle DeepMind (Brohan et al.) 2023

Launched the VLA paradigm now central to humanoid 'brains'; the compute-and-data picks-and-shovels story (GPUs, robot-data pipelines) rides on this line of work.

π0: A Vision-Language-Action Flow Model for General Robot ControlPhysical Intelligence (Black, Levine et al.) 2024

Leading generalist robot policy from the best-funded robot-foundation-model startup; a bellwether for whether VLA reliability crosses from 'modeled' to 'viable.'

Strain wave gearing (US Patent 2,906,143)C. W. Musser (United Shoe Machinery Corp.) 1959

Founding patent of the strain-wave 'harmonic' drive — the precision-reducer bottleneck every geared robot joint depends on and a concentrated picks-and-shovels supply chain.

On the Kinematic Error in Harmonic Drive GearsASME J. Mechanical Design (Ghorbel, Gandhi & Alpeter, Rice Univ.) 2001

Independent, reproduced characterization of strain-wave error grounding the reducer's viable-but-imperfect grade.

Kinematics of Roller Migration in the Planetary Roller Screw MechanismASME J. Mechanical Design (Jones & Velinsky, UC Davis) 2012

Reference kinematic model for the roller screw now used in humanoid linear actuators; grounds the modeled grade.

Artificial Muscles from Fishing Line and Sewing ThreadScience (Haines et al., Baughman group) 2014

Founding TCP-muscle paper; widely reproduced but low-efficiency, stays modeled.

High-Speed Electrically Actuated Elastomers with Strain Greater Than 100%Science (Pelrine, Kornbluh, Pei & Joseph, SRI International) 2000

Founding DEA paper; 25-year reproduced soft-actuator line still gated by kV drive.

Design Principles for Energy-Efficient Legged Locomotion (MIT Cheetah)IEEE/ASME TMech (Seok et al., MIT) 2015

Regenerative drivetrain energetics baseline for supercap-buffering hypothesis.

Performance and cost of materials for lithium-based rechargeable automotive batteriesNature Energy (Schmuch et al., MEET/Munster) 2018

Documents the ~250-300 Wh/kg practical cell plateau capping untethered humanoid endurance.

Series Elastic ActuatorsIEEE/RSJ IROS (Pratt & Williamson, MIT Leg Lab) 1995

Founding force-control-via-elasticity paper reproduced across Baxter, Valkyrie, exoskeletons.

Sparsh: Self-supervised touch representations for vision-based tactile sensingCoRL (Higuera et al., FAIR/UW/CMU) 2024

First serious 'touch foundation model'; bellwether for whether tactile pretraining crosses from single-lab metrics to reproduced cross-sensor value.

AnySkin: Plug-and-play Skin Sensing for Robotic TouchBhirangi et al. (NYU / Meta) 2024

Makes magnetic tactile skin swappable, low-cost and calibration-reusable — the BOM-repricing candidate for the tactile-sensor slot.

ReSkin: versatile, replaceable, lasting tactile skinsCoRL (Bhirangi et al., CMU / Meta AI) 2021

Founding reproduced result of the magnetic-taxel skin line; separates wear surface from electronics.

Digit 360: an artificial fingertip for omnidirectional, high-resolution touchMeta FAIR / GelSight (open-sourced) 2024

Pushes vision-based tactile to a fingertip with >8M taxels and ~1 mN resolution; marker for retrographic fingertips becoming standard.

Event-based Vision: A SurveyIEEE TPAMI (Gallego et al.) 2022

Consolidates neuromorphic vision; sets the honest line that the event SENSOR is reproduced/commercial while a reproduced robotics advantage is still emerging.

Skin electronics from scalable fabrication of an intrinsically stretchable transistor arrayNature (Wang et al., Bao lab, Stanford) 2018

State-of-the-art large-area stretchable e-skin; materials science real, robot-scale deployment open.

DeXtreme: Transfer of Agile In-hand Manipulation from Simulation to RealityNVIDIA Robotics (Handa et al.) 2023

Key REPRODUCTION anchor: re-demonstrated OpenAI-style cube reorientation on cheaper hardware.

Visual Dexterity: In-Hand Reorientation of Novel and Complex Object ShapesMIT (Chen, Tippur, Wu, Kumar, Adelson, Agrawal) 2023

Strongest single-lab general-shape reorientation result (downward-facing hand, <$5k rig); a bar the field must reproduce.

Diffusion Policy: Visuomotor Policy Learning via Action DiffusionColumbia / TRI / MIT (Chi et al.) 2023

The reproduced workhorse of learned manipulation; grounds the DATA + generalization bottleneck read.

Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT / ALOHA)Stanford (Zhao, Kumar, Levine, Finn) 2023

Democratized bimanual teleoperation + imitation; ALOHA rebuilt across labs, part of the reproduced imitation workhorse.

LEAP Hand: Low-Cost, Efficient, and Anthropomorphic Hand for Robot LearningCMU (Shaw, Agarwal, Pathak) 2023

De-bottlenecked dexterous-hand HARDWARE (~$2k, ~1/8 Allegro); shifts the constraint onto control/durability.

Dex-Net 2.0: Deep Learning to Plan Robust Grasps with Synthetic Point Clouds and Analytic Grasp MetricsUC Berkeley AUTOLAB (Mahler, Goldberg et al.) 2017

Origin of the reproduced-and-DEPLOYED grasping capability; the honest contrast that pick-and-place is solved while dexterity is open.

Neural feels with neural fields: Visuo-tactile perception for in-hand manipulationMeta AI / FAIR (Suresh et al.) 2024

Quantifies that touch disambiguates vision under hand-occlusion; released FeelSight benchmark.

IndustReal: Transferring Contact-Rich Assembly Tasks from Simulation to RealityNVIDIA / USC (Tang, Narang, Fox et al.) 2023

Industrial edge of manipulation on the NIST Assembly Task Board; single-lab, not reproduced across labs — anchors the contact-rich sim-to-real open grade.

OpenVLA: An Open-Source Vision-Language-Action ModelStanford / GDM / TRI / Berkeley / MIT (Kim et al.) 2024

Open-weight VLA making cross-embodiment transfer reproducible across labs — the reason robot-foundation-models rate modeled not open.

Octo: An Open-Source Generalist Robot PolicyOcto Model Team, UC Berkeley / Stanford (RSS) 2024

Second independent open-weight generalist policy on Open X-Embodiment; its reproducibility earns the field partial credit for cross-lab transfer.

Isaac Gym: High-Performance GPU-Based Physics Simulation for Robot LearningNVIDIA (Makoviychuk et al., NeurIPS Datasets) 2021

GPU-parallel simulation substrate underlying the reproduced sim-to-real recipe; the compute picks-and-shovels story.

DayDreamer: World Models for Physical Robot LearningUC Berkeley (Wu, Escontrela, Hafner, Abbeel, Goldberg, CoRL) 2022

Clearest real-robot learned-world-model demonstration; marks world-model learning as still open for contact-rich manipulation.

Gemini Robotics: Bringing AI into the Physical WorldGoogle DeepMind 2025

A 2025 flagship generalist VLA — impressive but single-lab and not independently reproduced; the bellwether the field must not mistake for viable generalist embodied intelligence.

Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body TeleoperationCoRL (Fu, Zhao & Finn, Stanford) 2024

Reference low-cost whole-body TELEOPERATION rig; central to the demo-provenance case as the data-collection method later re-presented as autonomy.

Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots (UMI)RSS (Chi et al., Stanford/Columbia/TRI) 2024

Handheld data capture decoupling demonstration from the robot; the clearest picks-and-shovels case that data hardware is the binding manipulation input.

DextrAH-G: Pixels-to-Action Dexterous Arm-Hand Grasping with Geometric FabricsCoRL (Lum et al., NVIDIA / Univ. Washington) 2024

Recent sim-to-real dexterous grasping; single-lab and grasp-focused, keeping the sim-to-real manipulation tier at open.

What Matters in Learning from Offline Human Demonstrations for Robot Manipulation (robomimic)CoRL (Mandlekar et al., Stanford/UT Austin) 2021

Foundational reproducibility study showing manipulation results are highly sensitive to data quality and eval protocol.

Agility Robotics Digit — commercial RaaS deployment at GXO LogisticsIndustry disclosure (Agility Robotics / GXO Logistics) 2024

Clearest real paid humanoid deployment — but single-vendor, narrow-task, no audited economics; anchors the narrow-vs-general deployment ledger.

Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement LearningCoRL (Rudin, Hoeller, Reist, Hutter, ETH Zurich) 2021

Showed massively parallel GPU sim trains robust locomotion in minutes, reproduced independently of NVIDIA.

Learning quadrupedal locomotion over challenging terrainScience Robotics (Lee, Hwangbo, Hutter et al., ETH Zurich) 2020

Independently reproduced link in the ETH sim-to-real locomotion lineage; strengthens the viable grade for perceptive legged locomotion.

Mini Cheetah: A Platform for Pushing the Limits of Dynamic Quadruped ControlICRA (Katz, Di Carlo & Kim, MIT) 2019

Open modular QDD design that seeded the commodity-actuator ecosystem; anchors electric-qdd-actuators (viable, director-merged).

Retrofitted Liquid Cooling for High-Power ActuatorsActuators MDPI (Paine & Sentis, UT Austin) 2015

Measured ~2x continuous-current gain; basis of liquid-cooled-actuators (modeled).

Design principles of a human mimetic humanoid (Kengoro)Science Robotics (Asano, Okada & Inaba, JSK) 2017

Single-lab musculoskeletal flagship; anchors musculoskeletal-humanoids at open (skeptic-demoted).

A Structural Battery and its Multifunctional PerformanceAESR (Asp et al., Chalmers) 2021

Coupon-scale ~24 Wh/kg; anchors structural-batteries-robotics (open).

HyQReal torque-controlled hydraulic quadrupedIIT (Semini et al.) 2019

Non-BD reproduction of hydraulic legged actuation; supports compact-hydraulic-actuation (viable).

Amprius 500 Wh/kg silicon-nanowire cell, third-party verificationAmprius Technologies 2023

Cell chemistry ahead of robot integration; pins high-energy-density-cells (modeled).

Variable impedance actuators: A reviewRAS (Vanderborght et al.) 2013

Multi-lab VSA survey; anchors variable-stiffness-actuators (modeled).

The DLR Hand Arm SystemICRA (Grebenstein et al., DLR) 2011

Peak tendon-drive + VSA evidence; roads not taken by deployment-grade humanoids.

Gecko-inspired adhesives grasp large objects in microgravityScience Robotics (Jiang et al., Stanford/JPL) 2017

Strongest adhesion-gripper result; modeled grade.

Versatile soft grippers with intrinsic electroadhesionAdvanced Materials (Shintake et al., EPFL) 2016

Independent adhesion-physics reproduction leg.

Lightweight low-power electroadhesive clutch and springICRA (Diller, Majidi & Collins, CMU) 2016

Founding zero-power torque-holding result; electrostatic-clutch-actuation (modeled).

DextrES: thin electrostatic brake hapticsACM UIST (Hinchet et al., EPFL/ETH) 2018

Second-lab reproduction of the electrostatic-clutch principle.

A Survey of Wide Bandgap Power Semiconductor DevicesIEEE TPEL (Millán et al.) 2014

Mature device physics; robotics adoption evidence is what's missing (wide-bandgap open).

Robot Collisions: Detection, Isolation, IdentificationIEEE T-RO (Haddadin, De Luca & Albu-Schäffer) 2017

Reproduced momentum-observer safety science underlying joint-torque collision detection.

Digitizing Touch with an Artificial Multimodal Fingertip (DIGIT 360)Meta FAIR + GelSight (arXiv:2411.02479) 2024

Multimodal-fingertip ceiling + free-hardware reproduction engine (modeled).

Embedding high-resolution touch across robotic hands (F-TAC)Nature Machine Intelligence 7:889-900 (PKU/QMUL) 2025

Whole-hand coverage milestone; strongest single-lab touch-payoff evidence.

Sparsh: Self-supervised touch representationsCoRL (Higuera et al., Meta FAIR) 2024

Flagship touch-foundation-model result; origin of TacBench.

AnySkin: Plug-and-Play Skin SensingICRA (Bhirangi, Pinto et al., NYU) 2025

Replaceability breakthrough for magnetic skins; same lineage as ReSkin (skeptic lineage-count correction).

uSkin: A soft skin with distributed 3-axis force sensingIEEE RA-L (Tomo et al., Waseda) 2018

Second truly independent magnetic-taxel lineage; commercialized by XELA, deployed on iCub — the leg that carries the viable promotion.

Evetac: Event-based Optical Tactile SensorIEEE T-RO (Funk et al., TU Darmstadt) 2024

1 kHz event tactile at ~1000x data reduction; hardware leg of neuromorphic-event-tactile (modeled, hardware-only).

3D-ViTac: Visuo-Tactile Fine-Grained ManipulationCoRL (Huang et al., Columbia/UIUC) 2024

Low-cost dense-taxel evidence that touch helps imitation policies (origin-lab).

SonicSense: In-Hand Acoustic Vibration PerceptionCoRL (Duke) 2024

Anchor for acoustic-contact-sensing (open).

TacO: Benchmarking Tactile Sensors for Object ManipulationarXiv:2605.21976 2026

First benchmark aimed at the tactile referee gap; origin-lab-run so far.

Visual dexterity: In-hand reorientation of novel shapesScience Robotics 8(84) (Chen et al., MIT) 2023

Strongest single in-hand result; one leg of the multi-lab rotation-primitive case.

Rotating without Seeing: In-hand Dexterity through TouchRSS (Qi et al., Berkeley/Meta) 2023

Independent lab + modality for the rotation primitive; bounds the claim (rotation, not goal-pose).

TaskGrasp: Data and Semantic Knowledge for Task-Oriented GraspingCoRL (Murali et al., CMU) 2020

Reference dataset separating stable from functional grasping.

Differentiable Physics and Stable Modes for Tool-Use PlanningRSS Best Paper (Toussaint et al., MIT/Stuttgart) 2018

Founding modern robot-tool-use treatment; still ahead of deployed capability.

Integrated Task and Motion PlanningAnnu. Rev. Control Robot. Auton. Syst. (Garrett et al., MIT) 2021

Canonical TAMP survey; shared formulation multiple labs build on.

Deep Whole-Body Control: Unified Policy for Manipulation and LocomotionCoRL (Fu, Cheng & Pathak, CMU) 2022

Founding learned loco-manipulation paper; first of three independent lineages.

Pedipulate: Manipulation with a Quadruped's LegICRA (Arm et al., ETH Zurich) 2024

Second independent loco-manipulation lineage.

UMI on Legs: Manipulation-Centric Whole-body ControllersCoRL (Ha et al., Stanford/Columbia) 2024

Third independent loco-manipulation lineage; interface-level recipe.

IndustReal: Sim-to-Real Contact-Rich AssemblyRSS (Tang et al., NVIDIA/USC) 2023

Strongest contact-rich sim2real evidence; single-lab-centered, defines the promotion trigger.

SpeedFolding: Bimanual Garment FoldingIROS (Avigal et al., Berkeley/KIT) 2022

Deformable-manipulation high-water mark; garment-class-specific.

vla-eval: Unified VLA Evaluation HarnessarXiv:2603.13966 2026

Reproduced published scores across six codebases; documented eval pitfalls.

VLA-REPLICA: Low-Cost Reproducible Real-World VLA BenchmarkarXiv:2605.20774 2026

Cross-site-consistent physical rigs; first credible path to cross-lab real-robot VLA scores.

RoboArena: Distributed Real-World EvaluationCoRL 2025

Multi-institution blind pairwise policy evaluation.

pi*0.6 / RECAPPhysical Intelligence (company report) 2025

Vendor-reported RL-on-deployment gains — cited as claim, not evidence, per skeptic demotion of vla-rl-posttraining.

RLinf: RL Infrastructure for Embodied AI (incl. WoVR)open-source 2026

Lowers reproduction barrier for EMB-P2.

HIL-SERL: Human-in-the-Loop RL for ManipulationUC Berkeley (arXiv) 2024

Sample-efficiency existence proof; adjacent to, not qualifying evidence for, VLA fleet post-training.

OpenVLA + OFT fine-tuning studyCoRL 2024 / RSS 2025 2024

Community-reproduced fine-tuning recipe; anchors open-vla-finetune-transfer (viable, scoped).

pi0 / pi0.5Physical Intelligence 2025

Strongest single-lab open-world generalization claim; ungraded by third parties.

GR00T N1: Open Foundation Model for HumanoidsNVIDIA (arXiv:2503.14734) 2025

Open dual-system humanoid VLA with human-video co-training and latent-action codes; company-evaluated.

Gemini Robotics (+ On-Device)Google DeepMind technical report 2025

Company-evaluated VLA tier; cited as claim.

LAPA: Latent Action Pretraining from VideosICLR (Ye et al., KAIST/UW/NVIDIA) 2025

Peer-reviewed leg of latent-action-pretraining (modeled); insufficient alone for human-video promotion (skeptic resolution).

Genie: Generative Interactive EnvironmentsICML best paper (DeepMind) 2024

Latent-action world models at scale — the underlying mechanism.

V-JEPA 2Meta FAIR (arXiv) 2025

Video world model with zero-shot arm results (single lab, company report).

DreamerV3: Mastering Diverse Control Tasks through World ModelsNature (Hafner et al.) 2025

Peer-reviewed, widely reproduced world-model RL (in sim).

Cosmos World Foundation Model PlatformNVIDIA (arXiv) 2025

Neural-simulator substrate; policy-evaluation validity unproven.

Open X-Embodiment / RT-XICRA best paper 2024

22 embodiments, 21 institutions; pooled-data gains.

Data Scaling Laws in Imitation Learning (Lin et al.)ICLR 2025

Diversity-dominated power laws; environment count dominates demos-per-environment.

RialTo: Real-to-Sim-to-Real Robust ManipulationRSS (MIT CSAIL) 2024

~67pp robustness gain from scene digitization.

ACDC: Digital Cousins for Robust Policy LearningCoRL (Stanford) 2024

Zero-shot sim-to-real from auto-generated scene variants.

SafeVLA: Safety Alignment via Constrained LearningarXiv (PKU) 2025

Early single-lab VLA safety work; anchors vla-safety-guardrails (open).

Helix: VLA for Generalist Humanoid ControlFigure AI technical report (uncontrolled) 2025

Cited as claim, not evidence — showreel class, inadmissible for maturity.

SmolVLAHugging Face (arXiv) 2025

Community-scale replication substrate for open-VLA fine-tuning.

SIMPLER: Evaluating Real-World Policies in SimulationCoRL 2024

The rank-correlation bar neural-sim evaluators must meet.

What Matters in Learning from Offline Human Demonstrations (robomimic)CoRL (Mandlekar et al.) 2021

Demonstrator proficiency dominates offline-IL outcomes; anchors demonstration-data-quality.

Data Quality in Imitation LearningNeurIPS (Belkhale, Gupta & Sadigh, Stanford) 2023

Formalism: action divergence and transition diversity govern IL outcomes.

Re-Mix: Optimizing Data Mixtures for Large-Scale ILCoRL (Hejna et al., Stanford) 2024

Mixture reweighting at fixed volume changes downstream success — curation is a lever.

Unitree Go1 remote-tunnel + UniPwn BLE root disclosuresIndependent security research (Makris & Finisterre; UniPwn authors) 2025

Concrete evidence base for robot-fleet-cybersecurity (open).

On the relevance of grasp metrics for predicting grasp successIEEE/RSJ IROS 2017 (C. Rubert, D. Kappler, A. Morales, S. Schaal & J. Bohg; Universitat Jaume I / MPI-IS) 2017

VERIFICATION NOTE, not a new find. Author list checked and found to overlap with both other cited 'legs' of the reproduced negative, and to differ from the author list given in the submitted component (which misattributed it to Rubert, Leon, Morales & Sancho-Bru — a different 2018 paper). Basis for demoting grasp-metric-predictive-validity viable -> modeled.

Predicting grasp success in the real world — a study of quality metrics and human assessmentRobotics and Autonomous Systems 121 (C. Rubert, D. Kappler, J. Bohg & A. Morales) 2019

Second leg of the grasp-metric negative — verified to share Rubert, Kappler, Bohg and Morales with the 2017 paper. One lineage, not independent reproduction.

Colosseum V2: Benchmarking Generalization for Vision Language Action ModelsarXiv:2605.27759 (J. Morgan, P. Vijay, H. Oh, J. Song, A. Arora, A. Du, G. S. Sukhatme, J. Thomason & I. Singh; USC) 2026

VERIFICATION NOTE. Author list confirmed to include Jesse Thomason and Ishika Singh, both co-authors of THE COLOSSEUM (RSS 2024) — refuting the submitted claim that V2 is 'a successor built by a different lead team'. 28 tasks, 13 categories, two morphologies on ManiSkill; evaluates ACT and pi0.5; asserts 'strong correlations between simulation and real-world metrics', which conflicts with V1's own R^2=0.614.

THE COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic ManipulationRSS 2024 (W. Pumacay, I. Singh, J. Duan, R. Krishna, J. Thomason & D. Fox; UW / AI2 / NVIDIA / USC) 2024

The V1 instrument, genuinely adopted by outside groups, whose 30-50% per-axis and >=75% combined degradation finding survives. Its moderate R^2=0.614 sim-real correlation is the number V2 implicitly contradicts from within the same lineage.

Where Autonomy Works: Evaluating Robot Capabilities in 2026Epoch AI (Y. Riviere & J.-S. Denain) 2026

VERIFICATION NOTE. Fetched and confirmed: published 2026-02-10, explicitly a secondary analysis reviewing demonstrations, deployments, videos, papers and interviews with no original trials; the 3-10x speed figures trace to vendor demonstrations (Figure ~4x, Physical Intelligence ~5x); authors acknowledge limited data. Basis for demoting human-relative-throughput to open and for striking these figures as grade evidence. The audit function itself remains valuable and is retained under independent-capability-auditing at open.

HARMONI at ELT: Design & Test on the current status of Low Order Wavefront Subsystem Pick Off Arm Engineering Model (LPOA-EM)arXiv:2607.22227 (I. Funes Vecino, G. Mercant Rubio, A. Alvarez Uruena, L. Garcia Moreno, G. J. Carracedo Carballal, I. M. Ferro Rodriguez, H. Argelaguet Vilaseca & J. Piqueras Lopez) 2026

VERIFIED BY DIRECT FETCH — title, all eight authors, submission date 2026-07-24 and abstract confirmed. The abstract states first results from the LPOA Engineering Model characterisation campaign on two rotary joints (shoulder and elbow) including angular resolution, discrete step tracking VALIDATED BY INDEPENDENT METROLOGY, rotation-axis wobble, and radial and axial runout of the bearing assembly. Chased from pantry leads claim:kraken:547be5482ecf and claim:kraken:82abbe548b5d to the real source; the leads pointed, this paper is the evidence. Carries instrument-grade-joint-metrology at model…

The Robotic Multiobject Focal Plane System of the Dark Energy Spectroscopic Instrument (DESI)The Astronomical Journal (DESI Collaboration), doi 10.3847/1538-3881/ac9ab1 2023

VERIFICATION NOTE. Confirms 5,000 positioners in 10 wedge petals of 500 (36 deg each), eccentric theta/phi kinematics, 10.4 mm pitch with overlapping patrol regions, fiber tips patrolling 12 mm disks, requirement <=5 um RMS — AND that the requirement is met ITERATIVELY (blind move to ~50 um, then two corrective moves), which is the scoping correction applied to robotic-fiber-positioner-arrays.

The benchmark of LiDAR odometry algorithms utilised for a mobile platformISPRS Archives XLVIII-1/W6-2025, article 25 (Bunker DVI dataset) 2025

VERIFIED EXTANT — a genuine third-party evaluation of ~20 LiDAR odometry systems on data its authors did not create, and the reason lidar-inertial-odometry holds at viable despite resting on a single HKU MaRS lineage. It is also the source that reports LIO-SAM failing outright, which is why LIO-SAM and KISS-ICP are struck as legs of that grade.

Flatness Preserves Instruction Following in Vision-Language-Action ModelsCMU (H. Zhang & Y. Bisk), arXiv:2606.23641 2026

The mechanism-plus-fix leg that lifts vla-language-grounding-shortcut to a confirmed viable: names the failure 'instruction blindness', attributes it to limited-data fine-tuning producing sharp high-curvature minima that destroy the pretrained VL representation, and reports >60% instruction-following improvement from sharpness-aware minimization alone with no extra data or architecture change. Zero author overlap with LangGap, Naver or RoboSemanticBench.

Efficient bipedal robots based on passive-dynamic walkersScience 307(5712):1082-1085 (S. H. Collins, A. Ruina, R. Tedrake & M. Wisse; Cornell / MIT / TU Delft) 2005

Confirmed viable after applying the people-not-institutions independence test: three physically distinct machines, three distinct PIs and advisory lineages, different actuation modalities, one agreed metric. Co-publication of independent replications under a shared protocol is stronger evidence than sequential replication, not weaker.

A Comprehensive Review of Piezoelectric Ultrasonic Motors: Classifications, Characterization, Fabrication, Applications, and Future ChallengesMicromachines 15(9):1170 (S. Naz & T. B. Xu) 2024

Supplies both halves of the piezo grade and therefore the scoping guard: the enabling property (self-locking with high holding torque, no standing current) and the permanent disqualifiers for a load-bearing joint (two-stage conversion with low efficiency, stator-rotor friction wear, 'not appropriate for continuous operation for long periods', representative rotary torque ~0.94 Nm).

A Careful Examination of Large Behavior Models for Multitask Dexterous ManipulationToyota Research Institute (TRI LBM Team), arXiv:2507.05331 2025

Establishes the binary success oracle's own noise floor — QA on ~27% of ~2,700 rollouts gave 2.31% success-label and 6.25% rubric disagreement between human graders. Load-bearing for the self-scored-referee-problem study, because it shows that automating a noisy oracle (AutoEval's learned success classifier) scales the noise rather than removing it.

Parallel elastic elements improve energy efficiency on the STEPPR bipedal walking robot 2017
A comparison of series and parallel elasticity in a monoped hopper 2015
Electrohydraulic musculoskeletal robotic leg for agile, adaptive, yet energy-efficient locomotion 2024
Peano-HASEL actuators 2018
Design principles for energy-efficient legged locomotion and implementation on the MIT Cheetah robot 2015
Proprioceptive actuator design in the MIT Cheetah 2017
Design of high torque and high speed leg module for high power humanoid 2010
An overview on principles for energy efficient robot locomotion 2018
Artificial Muscles: Mechanisms, Applications, and Challenges 2018
Robotic Artificial Muscles: Current Progress and Future Perspectives 2019
Bayesian exploration for intelligent identification of textures 2012
Sensing and recognizing surface textures using a GelSight sensor 2013
Biomimetic tactile sensor array 2008
A flexible and robust large scale capacitive tactile system for robots 2013
Methods and technologies for the implementation of large-scale robot tactile sensors 2011
GelSlim 3.0 2022
GelSight: High-resolution robot tactile sensors for estimating geometry and force 2017
Taxim 2022
Learning the signatures of the human grasp using a scalable tactile glove 2019
DIGIT 2020
Sparsh 2024
Touch and Go 2022
More Than a Feeling: Learning to Grasp and Regrasp using Vision and Touch 2018
See to Touch 2023
In-Hand Object Rotation via Rapid Motor Adaptation 2022
DeXtreme 2023
Visual dexterity: In-hand reorientation of novel and complex object shapes 2023
Rotating without Seeing 2023
Robot Synesthesia 2024
Learning to Walk in Minutes Using Massively Parallel Deep RL 2021
RMA: Rapid Motor Adaptation for Legged Robots 2021
The extraction of neural information from the surface EMG for the control of upper-limb prostheses 2017
Self-contained neuromusculoskeletal arm prostheses 2020
Trends in the Adoption of Robotic Surgery for Common Surgical Procedures 2020
Autonomous robotic laparoscopic surgery for intestinal anastomosis 2022
Harvesting Robots for High-value Crops: State-of-the-art Review and Challenges Ahead 2014
Development of a sweet pepper harvesting robot 2020
ANYmal in the Field 2021
RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model 2026
Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models 2026
RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination 2026
Hy-Embodied-VLM-1.0: Efficient Physical-World Agents 2026
CrossFormer: Scaling Cross-Embodied Learning 2024
HPT: Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained Transformers 2024
Open X-Embodiment: Robotic Learning Datasets and RT-X Models 2024
THE COLOSSEUM 2024
RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies 2025
RB2: Robotic Manipulation Benchmarking with a Twist 2021
SuSIE 2024
OpenVLA 2024
Human-Level Actuation for Humanoids 2025
Twisted string actuators: Comprehensive review 2025
A Novel Twisted-Winching String Actuator 2024
D3-ARM: Fully Decoupled Cable-driven Robotic Arm 2025
A Two-Layer Electrostatic Film Actuator with Integrated Brake 2025
Evaluation of Electrostatic Motor Efficiency 2025
Amprius SiCore / SA88 announcements 2025
The TacTip Family 2018
Optoelectronically innervated soft prosthetic hand 2016
Iontronic microdroplet array for tactile sensing 2015
Tactile sensibility in the human hand 1979
Contact sensing from force measurements 1993
Shape-independent hardness estimation with GelSight 2017
Improved GelSight for geometry and slip 2017
Stabilizing novel objects by learning to predict tactile slip 2015
More Than a Feeling 2018
A Century of Robotic Hands 2019
3D Diffusion Policy 2024
Equivariant Diffusion Policy 2024
Diffusion-EDFs 2024
Imitation Learning Based on Bilateral Control 2020
HG-DAgger 2019
Robot Learning on the Job (Sirius) 2023
Robot Programming by Demonstration 2008
ATOM-Bench 2026
RoboWorld: Fast and Reliable Neural Simulators for Generalist Robot Policy Evaluation 2026
FTP-1: A Generalist Foundation Tactile Policy 2026
Teleopit: A Full-Embodiment Humanoid Teleoperation System 2026
pi0.5: a Vision-Language-Action Model with Open-World Generalization 2025

Research artifacts: literature reads, models, and predictions about robotics component science. Evidence-tiered by REPRODUCED capability; single scripted demos are not treated as solved. Findings feed the wealth engine (Parallax) + the telescope dream as candidate analogies and signal seeds, not claims of fact. · as of 2026-09-05