Branch-Based AI Method Scores Higher on Math and Search Tests
An arXiv preprint in its first version reports higher averages than ARPO on five math and five search benchmarks at two model scales.
1421–1440
An arXiv preprint in its first version reports higher averages than ARPO on five math and five search benchmarks at two model scales.
Preprint tests SenseShift on stories and reviews, with stronger automated sentiment adherence but more mixed human judgments.
Preprint: QisMC handled selected large quantum programs, but its speed depended sharply on circuit structure and numerical precision.
Preprint: In three simulated industrial-load cases, two reduced formulations used fewer listed variables and had lower reported error scores than SAL and, where available, OVB.
An arXiv preprint reports that CPPC was fastest in tested systems of 1,000 or more variables, while spectral methods led in other regimes.
The proof-based study bounds Borel-subalgebra indices and extends zero-index results to nilradicals, their ideals and dual abelian-ideal modules.
An arXiv preprint reports strong benchmark agreement from a model that combines collaborative and adversarial paths, while fresh human-rating tests remain needed.
Preprint: The first reported laboratory measurement differs from the ground-state spectrum and suggests possible markers for naphthalene’s triplet state.
Preprint: A theorem and simulations report uniform angle guarantees across the full Grover-angle range, with different depth windows producing different error and shot requirements.
Preprint: An offline-learning version of ant colony optimization reported higher observation benefit across 14 simulated scenarios, while operational applicability remains untested.
A preprint under review reports that a small probing update can forecast which internal parts matter after targeted tuning, with the strongest pattern on controlled benchmarks.
An arXiv preprint found that directing one of two verification checks to a critical provenance path sharply improved decisions in controlled model tests.
This arXiv Preprint applies a class-level model to Project STAR data, with the reported contrast depending on graph and normalization choices.
A preprint reports higher results for IAPO on several service-agent benchmarks, while multi-turn function-calling scores were comparable.
A preprint reports higher held-out likelihood when a low-rank model captures correlated uncertainty across outputs.
A preprint reports stronger factor alignment for X-MULTI, while some sensor and viewpoint predictions remained unreliable.
Preprint: In simulated speech scenes, ADEPS led on signal-distortion scores across the tested arrays, while other measures were mixed.
A formal analysis traces how Fu and Nijhoff’s U-system connects KP to two-component KP, AKNS and ASDYM, while leaving BKP and CKP open.
Preprint: The method uses predicted output entropy as a continuous difficulty signal to adjust the reasoning-path budget for each question.
Preprint researchers report that INCEPT transfers across signal, brain-state and brain-health EEG tasks, while the benchmark does not test patient outcomes.