A new preprint reports a method that builds lower and upper disturbance-reserve bounds, reaches exactness after finite expansions and delivers benchmark speedups over warm-started full solves.
An arXiv preprint presents a unified Rényi framework for composite binary testing, combining finite-sample bounds with asymptotic results that separate two Type II error regimes.
CF-YOLO, a YOLOv11-based detector built to refine context and fine detail, scored higher than YOLOv11n on CTDD but delivered mixed results on external NEU-DET data.
Benchmark tests on three RVV 1.0 RISC-V CPU configurations found that juFFTe's gains varied by hardware, with a reported threefold average speedup on the SG2044 multi-core comparison, while AMD Zen 5 led across the full range.
A theoretical preprint reports that fixed-order Khatri-Rao sketches can achieve a near-linear subspace-embedding dimension in the subspace size, improving the stated dependence on that size over earlier bounds. The result is a proof under sub-Gaussian assumptions, not an experimental performance claim.
A tiny two-recording test found that a parameter-free color extrapolator beat copying the last frame and every trained comparator on held-out copper in both directions. The air-to-chamber advantage was individually separated from zero, while the chamber-to-air interval included zero.
An evaluation in simulation and on a physical robot found that one shared interface could support three instruction modes, with gesture-based modes often performing better when wording, surfaces or objects changed.
H-Scale uses calibration activations to choose hardware-valid NVFP4 scales, with higher average scores reported across Qwen3 and LLaMA tests but a narrow evaluation scope.
Gibbs-family methods led the aggregate scores in Traffic Hourly, Electricity Hourly and Solar Weekly, but classical baselines remained on top in some M4 frequency and disagreement groups.
VICT uses a task's terminal verifier to trace credit through long action sequences instead of using it only as a final pass-or-fail signal. In ALFWorld and WebShop tests, it beat GRPO and achieved higher validation AUC over the same 300 updates, while the authors caution that verifier-defined links do not establish causal necessity.
An arXiv preprint proposes a 3D-grounded test for robot video models. Cosmos led composite scores, but detailed checks found weaker object localization and trajectory accuracy, underscoring the gap between plausible footage and executable behavior.
A theoretical result forces the number of cells in a bent partition to be a prime power and narrows its exponent using the dimension of the underlying finite space.
A video AI method called Token-Budget Distillation retained strong benchmark scores after aggressive visual-token compression, although its training still depended on the uncompressed teacher model.
An anatomy-aware system called CheXtriev reported stronger case-retrieval scores than global and local comparison methods on selected chest radiographs. The gains were especially notable for several lower-prevalence findings.
An arXiv preprint tests a multi-critic training method for robots that push and transport objects and open a dishwasher through contact. It reports 94.1% simulation success and 69.0% success in 58 trials on four unseen objects, while the dishwasher test is qualitative.
A preprint examining French and Egyptian Arabic movie dialogue finds that six AI systems align more closely with humans on visible social cues than on subtle power relationships, while multimodal results are limited by incomplete coverage.
Across 120 slots per condition, first post-edit re-verification appeared in 78.3% of cadence-guided slots and 26.7% of cadence-omitted slots. The same descriptive comparison showed fewer cadence violations and more bounded final successes with the guidance.
A mathematical study links the worst-case rank needed to approximate normalized attention to support geometry, while a fixed BERT-base calibration reports lower effective dimensions in some tested cells.
The model with the highest pooled scores retained minute-level overnight patterns, but the small hospital sample and limited calibration do not support individual care use.
A new preprint benchmarks language models on questions aligned with CFA Levels I to III and FRM Parts I to II. Leading systems exceeded 97% on Easy items but fell sharply on Hard cases, while the gated result was 0.39 percentage points higher on held-out questions.
In a randomized experiment, people working with a non-human-shaped robot disclosed more when its small talk contained fewer personal details, while teamwork and coordination ratings were also higher.
An arXiv preprint reports an LLM-assisted workflow for assigning application tasks to a processor or FPGA, with reported speedups up to 92.53 times and repeated partition choices in the tested configurations.
EfficientNetB0 correctly classified 97.36% of 303 held-out mango images, with eight errors, in a study that also deployed the model through a public web app. The result is an initial within-dataset benchmark, not evidence that the tool will deliver the same performance across regions or improve agricultural decisions.
Journal of Bangladesh Academy of Sciences, vol. 50, Supplement 1, p. 114, 20265 min read
A redesigned estimator separates measurement-dependent and grid-dependent work, cutting reported runtime and memory demands while retaining similar simulated error performance.
The paper presents CrabOS as an operating-system approach to human-AI handoffs, using shared text objects and common capability controls, but reports architecture and case studies rather than measured evidence of better results.
A four-level robotic bin-picking system cleared all 30 experimental bins, while the study found that bin clearance and individual grasp success told different stories.
A preprint reports that dynamic sensor selection and time-frequency allocation outperformed equal resource sharing in a simulated cloud radar network when communication capacity was tight.
A text-only reconstruction of Reactome preserved its large-scale network pattern, but the study measured graph similarity rather than biological correctness.
A mathematical framework for heterogeneous, constrained agents reached consensus in reported simulations and had the lowest listed cost in comparisons with 50 and 100 robots.
The fixed-round PTD model reported sharp speed gains and higher scores across several video benchmarks, while leaving disjoint events and multiple matching targets largely unexplored.
A simulated waveform was associated with better OOK detection in the reported tests, but its effect on existing OFDM data varied with channel complexity and signal shape.
Ouakouak, B.E., Zegrar, S.E. and Arslan, H., 2025. CP-Aware OFDM-Based OOK Signaling. IEEE Wireless Communications Letters, 15, pp.935-9394 min read
A preprint describes one OCR system trained across 13 Indic scripts; its reported overall character error rate was 6.9%, compared with 8.6% for monolingual models.
A theoretical and simulation study suggests that private mean estimation can outperform non-private estimation over time when leakage-participation feedback is strong enough.
Tests across nine systems found a sharp drop in numeric precision as prompts asked for more objects. Layout, composition and appearance also mattered, but high-count results were harder to validate.
A preprint by Junmin An and Jon-Lark Kim gives exact formulas for the shortest self-orthogonal and LCD embeddings of linear codes over the ring Fq + uFq.
D-TAIA combined parameter-efficient language-model adaptation, domain-aware pre-training and retrieval, then matched or improved on LLM and RNN baselines across four event logs.
A benchmark study found the largest routing gap when user-profile details changed which skill best matched a task, while profile-insensitive cases stayed similar.