An evaluation in simulation and on a physical robot found that one shared interface could support three instruction modes, with gesture-based modes often performing better when wording, surfaces or objects changed.
An arXiv preprint proposes a 3D-grounded test for robot video models. Cosmos led composite scores, but detailed checks found weaker object localization and trajectory accuracy, underscoring the gap between plausible footage and executable behavior.
An arXiv preprint tests a multi-critic training method for robots that push and transport objects and open a dishwasher through contact. It reports 94.1% simulation success and 69.0% success in 58 trials on four unseen objects, while the dishwasher test is qualitative.
In a randomized experiment, people working with a non-human-shaped robot disclosed more when its small talk contained fewer personal details, while teamwork and coordination ratings were also higher.
A four-level robotic bin-picking system cleared all 30 experimental bins, while the study found that bin clearance and individual grasp success told different stories.
A simulated comparison found higher success and shorter routes for an LLM-guided UAV than for a conventional lawn-mower scan, but the evaluation did not test unknown environments.
A preprint describes MaCoPlanner, a system that turns equipment manuals into structured knowledge, checks robot plans before action, and rejects unresolved candidates. It reported higher success on the hardest benchmark tasks than the listed baselines, but the physical evaluation was small and limited to a no-load simulator.
A robotics preprint reports higher task success for a visual-track system in simulation and physical manipulation, along with better future-video prediction scores.
A preprint audit of 12 physical-AI benchmarks found positive relationships across every pair, with especially strong agreement for two pairs of tests. The analysis covered 51 models and suggests that benchmark averages can repeat the same strengths, while a compact suite may preserve useful model separation.
A two-specimen laboratory study found that a distributed measure of vegetation stiffness stayed stable across robot push heights, while a lumped stiffness varied.
FlashVLA, a streaming method for vision-language-action policies, cut reported inference time while maintaining or improving task success across selected simulated benchmarks and three physical-arm tests.
A one-path benchmark reported the lowest observed latency for ROS2 Connect among the tested systems, alongside higher stability and improved scalability under concurrent load.
A methods study reports that SOPO-CD was faster than stated comparison methods in tested placement benchmarks, while its design leaves open questions about search breadth and real-world testing.
Aerial tracks from 20 intersections were used to train planners and motion predictors for a new city in offline tests, with real-world driving left untested.
A laboratory study of one soft pneumatic actuator reported path tracking, user-guided motion, obstacle clearance and local stability checks for a structured learned controller.
A preprint reports a GPU simulator that is extremely fast on a simple rigid-body task, while deformable and surgical tests expose limits in scaling, validation and what the results can show.
A tendon-driven prototype responded to preset voice commands and repeated finger movements, but its open-loop grip was sensitive to object surfaces and thickness.
A preprint reports a robot controller that used variation among sampled future actions to adapt stiffness and damping during collaborative transport. The proposed condition recorded a 0.95 average success rate across the tested directions, compared with 0.83 for a fixed-stiffness control and 0.69 for deterministic FACTR.
A preprint reports that RAEM, a LiDAR-equipped quadruped robot system, completed multi-floor exploration in simulation and reached the fifth floor of a real stairwell. The work also reports lower GPU map-construction times, while real-world repetition counts were not provided.
LAC is a whole-body humanoid controller designed to adjust responses to pushes and twists separately. The reported tests covered simulated stiffness sweeps, real-robot contact responses and VR-controlled tasks including object carrying and door opening.
In simulation, the reported controller reached a target pose even though it was outside the initial obstacle-free box, while measured iteration times stayed below the sampling period.
A preprint presents a heat-transfer model for industrial automated tape laying that incorporates finite radiation geometry and mixed convection. It matched a reported temperature series with a 4.02°C RMSE, but independent validation remains outstanding.
Composites Part A: Applied Science and Manufacturing (2026)3 min read
A constrained multibody model suggests that a welding umbilical can add non-negligible modeled joint torques despite its much lower mass. The result comes from a deterministic planar simulation with prescribed robot motion, while experimental validation remains future work.
20th International Symposium on Advances in Robot Kinematics, Jun 2026, Barcelone, Spain3 min read
A preprint reports higher aggregate scores for confidence-guided episode selection in an EVAC robot world-model benchmark, while the Semantics result remains uncertain.
A retrieval-based robot-learning approach reported higher average success than comparison systems on held-out LIBERO and UR5e manipulation tasks, while the paper said absolute out-of-distribution success remained low.
A simulated social-navigation system uses explicit memory to retain rare, high-cost failures and respond to unfamiliar social behavior, according to an arXiv preprint.
A robot-navigation system that combines learned visual waypoints with geometry-aware local control reported higher success and path-efficiency scores than a matched baseline in simulated and office tests.
The planner used explicit rules for grasping, placing, sliding and lifting and recorded higher success and faster planning than tested alternatives in a simulated kitchen-table benchmark.
GaussianDream++ reported 98.6% average success on LIBERO, 87.8% on LIBERO-Plus Overall and 52.5% pooled success in a physical evaluation, compared with 29.2% for reproduced π0.5.
A robotics preprint reports broader feasible sampling in a single-joint diagnostic, task-specific success rates under different tuning settings, and reported transfer to physical UR5e manipulators.
A GNSS positioning study reported lower combined error scores for a self-supervised route in ID and Heavy OOD tests, while supervised-only training led in Slight OOD.
LM-X recorded higher mean success than an action-only backbone in a pretraining gate and than GR00T N1.7 in simulation and real-world comparisons, while its RTG and variance traces showed temporal patterns that were not calibrated as failure detectors.
A preprint describes a robot-control method that updates touch data during execution and reports higher benchmark success than several comparison policies in the tested tasks.
AGRO-Nav, a graph-based orchard planner, stayed closest to row centers in the reported simulation and real-orchard comparisons, while the study evaluated only static global planning at one commercial orchard site.
A generalized robot-planning method reported the lowest mean path costs on most tested manipulation problems and the highest average route-class coverage in 2D map tests.
A robotics preprint reports that MA-VLA scored higher than Pi0 on listed in-domain benchmarks and produced nonzero results when known atomic actions were recombined into specified multi-arm collaboration patterns, but the evidence is limited to selected simulations and one dual-arm platform.
DESCENT reported lower trajectory errors than Amelia-TF in critical-agent and random-agent tests, while using fewer parameters but showing higher reported latency.