Simulated Tasks Reveal a 4.8-Fold Gap in Agent Recovery Costs
A preprint reports that, in the Habitat-Sim simulation, Reflexion and CoPAL had similar success rates but sharply different recovery costs.
1381–1400
A preprint reports that, in the Habitat-Sim simulation, Reflexion and CoPAL had similar success rates but sharply different recovery costs.
Preprint: In one stochastic mine-pump model, the method reached the highest-reward design after 130 simulations, while Bayesian optimisation outperformed a genetic algorithm on reward summaries.
Preprint: The method had the lowest error in a baseline simulation and was also applied to schooling data.
Preprint: An agent-led process was associated with higher preference rates on held-out Chinese K–12 tutoring tasks, with human checks favoring later answers.
Preprint states the bound for −∞ ≤ p ≤ 1/d, with least multiplicative constant C = 1; for m ≥ 2, its power is optimal when C remains 1.
Preprint: PARTAB uses four stages to map a table question to an answer.
Preprint: MnemoDyn showed strong reconstruction and prediction results in tests, but the evaluation did not establish clinical usefulness.
A mathematical preprint reports a best early stopping scale, followed by covariance loss and, at far longer times, collapse of the reverse process onto the training set.
Preprint: Luce reports leading scores on several 3D-generation benchmarks, while its authors flag limits with tiny details and complex materials.
A mathematical proof study describes a Bakry–Émery framework and derives conditional Poincaré and modified logarithmic Sobolev inequalities.
Preprint tests a controller using encoder position and motor effort, with high grasp success in firmer simulations but a sharp grasp–damage trade-off in the softest cases.
Preprint: A qualitative review of deepagents, pi and dsh finds five recurring design elements, but none offers an independently verifiable record.
The report compares 140,200 pre-SCOUT tool-schema tokens with 1,300 production tokens in one implementation and reports an 84.8% first-result hit rate across 45 evaluable benchmark queries.
An arXiv preprint reports lower attack success on three question-answering benchmarks, but its guarantee depends on assumptions about the retrieved documents.
Preprint: A physics-informed system reported lower 3D pose errors on benchmarks and in an indoor zero-shot test.
Preprint: LT-MKT combines cognitive-load modeling with cross-domain knowledge transfer in an AI learning model.
A preprint reports a theoretical result linking the equality to open Richardson varieties and a type-A polytope criterion.
A preprint reports a multi-week experiment on Kuaishou's production traffic covering more than 40 million users, with TAGR outperforming the DLRM baseline on revenue and engagement rates.
Preprint: Training with paired visual and privileged physical trajectories was associated with lower forecast drift and higher average control success in simulated robot tasks.
Preprint examines social-choice mechanisms for AI systems affecting many people and shows that tighter welfare floors can carry rising average-welfare costs.