AI agents report five math results judged novel in test
Preprint says an open-world system beat AlphaEvolve on three other problems in a 14-problem evaluation.
1521–1540
Preprint says an open-world system beat AlphaEvolve on three other problems in a 14-problem evaluation.
An arXiv preprint describes DriftAD, which adapts text descriptions to local visual features and reports results on MVTec-AD and VisA.
Preprint: GAP-Prompt reports results across three image benchmarks, while task-querying accuracy ranges from 44.30% to 55.75%.
Preprint: OrbitNet posted the lowest mean absolute error on all six unseen constellations, although its other error scores varied.
Preprint: G2I aggregates local graph clauses into a global intervention policy and reports higher AUCC than CF and CF2 across intervention datasets.
Preprint analysis reports a 3.09-point gap on one reasoning condition despite high semantic similarity, with differences also appearing in tokenization, training signals and rewards.
Preprint: A computational study found that embeddings, structural methods and eligibility rules were associated with different evidence and wording in AI-generated Big Five candidate forms.
An arXiv preprint reports higher model-based scores on difficult shape descriptions, but results vary by backbone and benchmark.
A preprint reports results from a 4,144-video benchmark for 12 simulated care classes, with temporal models near 94% top-1 accuracy versus 23.17% for a single-frame comparator.
The lightweight system estimates runtime, throughput, cost and bottlenecks from model and hardware characteristics, with a reported median error below 15%.
Preprint BotScan reports more live-server finds in selected IPv4 tests, but says encryption and unreliable responses remain obstacles.
Preprint: Potential-modulated iSCAT produced different signals from connected and isolated patterned ITO regions.
An arXiv preprint reports higher scores for a decentralized LLM-agent market with subcontracting, but finds agents' cost estimates remain unreliable.
The method was applied to 6,126 American Time Use Survey responses, where ICL favored a two-component contaminated fit.
A version 1 arXiv preprint develops theorem-based results for Lagrange, Hermite and best uniform approximation in an abstract mathematical framework.
This arXiv preprint constructs a meromorphic function on the complex plane that omits 0, 1 and infinity in the upper half-plane but lies outside the Nevanlinna class.
Preprint: In two numerical PDE testbeds, NEMO's reported operator-evaluation times were 0.012 seconds versus 541 seconds in one test and 0.006 versus 5.6 seconds in the other.
Preprint: A task-routed system reported a 62.9 average across five COIN tasks and higher throughput than an 8B baseline in one profile.
A version 1 arXiv preprint dated 25 August 2026 reports conditional existence results and a variation-of-constants formula for abstract equations with time-dependent infinite delay.
A formal mechanism-design analysis says some efficient rules that work with self-interested agents cannot be implemented when opponents are required to maximize one another's payoffs.