A preprint examining French and Egyptian Arabic movie dialogue finds that six AI systems align more closely with humans on visible social cues than on subtle power relationships, while multimodal results are limited by incomplete coverage.
The model with the highest pooled scores retained minute-level overnight patterns, but the small hospital sample and limited calibration do not support individual care use.
The paper presents CrabOS as an operating-system approach to human-AI handoffs, using shared text objects and common capability controls, but reports architecture and case studies rather than measured evidence of better results.
A text-only reconstruction of Reactome preserved its large-scale network pattern, but the study measured graph similarity rather than biological correctness.
A benchmark study found the largest routing gap when user-profile details changed which skill best matched a task, while profile-insensitive cases stayed similar.
A preprint reports stronger image-quality results from Flow Matching than diffusion on sparse-view chest CT, with FlowDPS leading the reconstruction comparison.
A diagnose-then-correct system called DoCtOR reported higher success rates on HotPotQA, ChartQAPro and Mind2Web. Its attribution model was less accurate on held-out failures, and open-ended, multilingual and larger-team settings were not tested.
A buffered sample-selection method called RECAST reported lower blood-pressure estimation error than TTC and no adaptation on PulseDB and MC-MED. The results came from simulated streaming benchmarks, with extra but sub-second latency and no prospective clinical evaluation reported.
A methods preprint compares three real-valued position encodings, finding similar classification averages and an exact, faster algebraic shift for its Sinusoid variant.
GRACE, a gradient-guided method for selecting compact forget and retain sets, reported higher forget-sample retrieval accuracy than RASLIK and the highest model utility in six of eight algorithm-model combinations, while forget-quality scores varied little across selectors.
A randomized experiment found that list position still shaped which hotels AI agents inspected, but much less than in a human benchmark; effects on final booking varied by model and were no longer statistically significant at the highest tested reasoning effort.
A preprint reports higher separation between marked content and instructions in frozen-model tests, with utility nearly unchanged and copied passages remaining highly similar.
A computational preprint reports sharply lower code and math scores in targeted MLP-neuron suppression tests, while several non-target benchmarks retained much higher scores.
ProgRouter reported strong task and citation results while tracking operating-cost budgets, GPU energy use, and execution time across four benchmark streams.
A study of 536 failed web-task runs found that whole-trajectory reruns often produced different outcomes without showing that the recorded failure mechanism had been corrected.
A computational evaluation reports that HarnessLens improved average held-out performance by 7.6% to 13.6% across three agent harnesses and four benchmarks, while using a smaller configured interaction budget than the baselines.
A preprint reports that KLOD can update targeted facts inside language models while keeping edit success near 100 percent in benchmark tests, although broader generalization and preservation of unrelated predictions remain in tension.
A training method for autoregressive spiking language models improved reported benchmark averages over matched knowledge-distillation checkpoints and stayed non-collapsed in a 10-seed stress test. The work remains an early computational evaluation, with no hardware measurements or evidence about human language performance.
A preprint describes an offline model-based reinforcement-learning framework for cost-sensitive incentive allocation. It reports higher net profit in offline and online comparisons, while diagnostics show weaker evaluation under policy drift.
A score learned from ICU trajectories distinguished survivors from non-survivors across four baseline-severity groups, yet its patient-level behavior varied between two hospital systems.
A preprint reports that MAIL, an automated literature-based system for generating chemistry hypotheses, posted the strongest listed overlap scores on the 51-paper TOMATO-Chem benchmark. The result shows reference-idea recovery, not demonstrated chemical discovery.
A live edge testbed evaluation recorded zero wrongful actuation and complete benign success for the full Edge Skillguard policy, with the outcome measured at the software handoff to device adapters.
A preprint reports higher MO-IKE results than comparison methods in several frozen-model benchmarks. Results vary by benchmark, with retention remaining weaker than other measures in key tests.
A computational preprint tested one internal direction across five language models, reporting a wide range of tool-call rates, transfer to held-out tools and a live-search cost-accuracy trade-off.
A seven-design evaluation found that DEVICES was more accurate than two one-shot prompting baselines and used an 8.6-fold smaller input context, but it tested only documentation-level compatibility.
A retrospective registry evaluation found that a hierarchical model combining lobe-aware CT analysis with selected EHR groups reached a mean AUC of 0.8750, slightly above imaging-only REN, but the difference was not statistically significant.
MICCAI Machine Learning in Medical Imaging (MLMI 2026)5 min read
A modeling study of five signalized intersections found that pooled training produced lower 10-second trajectory errors than site-specific models at every site. Performance was less consistent at shorter horizons, and fine-tuning helped some held-out sites but not all.
A benchmark of 13 frontier AI models found partial success on prescribed computational-biology workflows, with performance lower on the deepest analyses and largest raw-data tasks.
A preprint introduces a benchmark that tests financial language models on distinct review operations and on decisions about when to request more evidence.
An arXiv preprint reports a gap between retaining required values and placing them correctly in complex JSON and tables, then tests a reward aimed at structural placement.
An AI coding benchmark linked stripped-down task specifications and thinking effort to different costs, while a low-cost probe was linked to better estimates for a held-out task.
A controlled 2D benchmark found that multimodal agents made most reconstruction gains early, often lost accuracy later, and remained well below a data-policy reference.
A preprint describes an open simulator for beyond-visual-range air combat that supports mixed aircraft, high-throughput testing and standard multi-agent interfaces. Its reported evidence comes from simulated benchmarks, aggregate policy-transfer results and a limited module-level missile case, not operational testing.
A computational preprint compared reinforcement-learning designs for language-model auditors. One calibrated pairwise checkpoint scored above an untrained baseline on the reported benchmark measures, but results varied by reward strategy and a concerningness-focused run paired a higher production-discovery score with sharply lower false-positive calibration.
PonsRAG links character and plot evidence through a separate bridge layer and led the study's reported single-step and multi-step multiple-choice comparisons.
An arXiv preprint reports that CaSKG recorded higher benchmark scores and fewer mean environment steps than Graph-of-Skills across six language-model backbones.