Text Models Spot Evasion Better Than They Read Vocal Confidence
Models performed strongly on written signs of evasion but struggled more with vocal unconfidence, and speaker-level calibration helped only modestly.
Scientific publication4 min read
Research field
Artificial intelligence, computing, robotics, energy, and engineering.
Models performed strongly on written signs of evasion but struggled more with vocal unconfidence, and speaker-level calibration helped only modestly.
Scientific publication4 min read
A systematic review finds that most deep-learning brain MRI reconstruction studies did not compare image-fidelity scores with radiologist assessments on the same data, leaving key safety questions unanswered.
Scientific publication4 min read
A preprint introduces egRUE, a method that combines uncertainty scoring with feature-level explanations; its reported gains came from benchmark tests and a small BloodMNIST expert study.
Scientific publication5 min read
A new benchmark finds that language models often handle the meaning of Chinese internet neologisms better than the sounds, characters and source forms used to create them.
Scientific publication5 min read
A new preprint reports a method that builds lower and upper disturbance-reserve bounds, reaches exactness after finite expansions and delivers benchmark speedups over warm-started full solves.
Scientific publication4 min read
An arXiv preprint presents a unified Rényi framework for composite binary testing, combining finite-sample bounds with asymptotic results that separate two Type II error regimes.
Scientific publication4 min read
CF-YOLO, a YOLOv11-based detector built to refine context and fine detail, scored higher than YOLOv11n on CTDD but delivered mixed results on external NEU-DET data.
Scientific publication4 min read
Benchmark tests on three RVV 1.0 RISC-V CPU configurations found that juFFTe's gains varied by hardware, with a reported threefold average speedup on the SG2044 multi-core comparison, while AMD Zen 5 led across the full range.
Scientific publication4 min read
The system sends a short token prefix and generates the untransmitted tail at the receiver, trading communication overhead for computation.
Scientific publication5 min read
A theoretical preprint reports that fixed-order Khatri-Rao sketches can achieve a near-linear subspace-embedding dimension in the subspace size, improving the stated dependence on that size over earlier bounds. The result is a proof under sub-Gaussian assumptions, not an experimental performance claim.
Scientific publication4 min read
A tiny two-recording test found that a parameter-free color extrapolator beat copying the last frame and every trained comparator on held-out copper in both directions. The air-to-chamber advantage was individually separated from zero, while the chamber-to-air interval included zero.
Scientific publication5 min read
An evaluation in simulation and on a physical robot found that one shared interface could support three instruction modes, with gesture-based modes often performing better when wording, surfaces or objects changed.
Scientific publication4 min read
H-Scale uses calibration activations to choose hardware-valid NVFP4 scales, with higher average scores reported across Qwen3 and LLaMA tests but a narrow evaluation scope.
Scientific publication4 min read
Gibbs-family methods led the aggregate scores in Traffic Hourly, Electricity Hourly and Solar Weekly, but classical baselines remained on top in some M4 frequency and disagreement groups.
Scientific publication4 min read
VICT uses a task's terminal verifier to trace credit through long action sequences instead of using it only as a final pass-or-fail signal. In ALFWorld and WebShop tests, it beat GRPO and achieved higher validation AUC over the same 300 updates, while the authors caution that verifier-defined links do not establish causal necessity.
Scientific publication4 min read
An arXiv preprint proposes a 3D-grounded test for robot video models. Cosmos led composite scores, but detailed checks found weaker object localization and trajectory accuracy, underscoring the gap between plausible footage and executable behavior.
Scientific publication5 min read
A theoretical result forces the number of cells in a bent partition to be a prime power and narrows its exponent using the dimension of the underlying finite space.
Scientific publication4 min read
A video AI method called Token-Budget Distillation retained strong benchmark scores after aggressive visual-token compression, although its training still depended on the uncompressed teacher model.
Scientific publication5 min read
An anatomy-aware system called CheXtriev reported stronger case-retrieval scores than global and local comparison methods on selected chest radiographs. The gains were especially notable for several lower-prevalence findings.
Scientific publication4 min read
An arXiv preprint tests a multi-critic training method for robots that push and transport objects and open a dishwasher through contact. It reports 94.1% simulation success and 69.0% success in 58 trials on four unseen objects, while the dishwasher test is qualitative.
Scientific publication4 min read
A preprint examining French and Egyptian Arabic movie dialogue finds that six AI systems align more closely with humans on visible social cues than on subtle power relationships, while multimodal results are limited by incomplete coverage.
Scientific publication4 min read
Across 120 slots per condition, first post-edit re-verification appeared in 78.3% of cadence-guided slots and 26.7% of cadence-omitted slots. The same descriptive comparison showed fewer cadence violations and more bounded final successes with the guidance.
Scientific publication4 min read
A mathematical study links the worst-case rank needed to approximate normalized attention to support geometry, while a fixed BERT-base calibration reports lower effective dimensions in some tested cells.
Scientific publication5 min read
The model with the highest pooled scores retained minute-level overnight patterns, but the small hospital sample and limited calibration do not support individual care use.
Scientific publication4 min read
A new preprint benchmarks language models on questions aligned with CFA Levels I to III and FRM Parts I to II. Leading systems exceeded 97% on Easy items but fell sharply on Hard cases, while the gated result was 0.39 percentage points higher on held-out questions.
Scientific publication5 min read
In a randomized experiment, people working with a non-human-shaped robot disclosed more when its small talk contained fewer personal details, while teamwork and coordination ratings were also higher.
Scientific publication4 min read
An arXiv preprint reports an LLM-assisted workflow for assigning application tasks to a processor or FPGA, with reported speedups up to 92.53 times and repeated partition choices in the tested configurations.
Scientific publication4 min read
EfficientNetB0 correctly classified 97.36% of 303 held-out mango images, with eight errors, in a study that also deployed the model through a public web app. The result is an initial within-dataset benchmark, not evidence that the tool will deliver the same performance across regions or improve agricultural decisions.
Journal of Bangladesh Academy of Sciences, vol. 50, Supplement 1, p. 114, 20265 min read
A redesigned estimator separates measurement-dependent and grid-dependent work, cutting reported runtime and memory demands while retaining similar simulated error performance.
Scientific publication5 min read
The paper presents CrabOS as an operating-system approach to human-AI handoffs, using shared text objects and common capability controls, but reports architecture and case studies rather than measured evidence of better results.
Scientific publication4 min read
A four-level robotic bin-picking system cleared all 30 experimental bins, while the study found that bin clearance and individual grasp success told different stories.
Scientific publication4 min read
A preprint reports that dynamic sensor selection and time-frequency allocation outperformed equal resource sharing in a simulated cloud radar network when communication capacity was tight.
Scientific publication4 min read
A text-only reconstruction of Reactome preserved its large-scale network pattern, but the study measured graph similarity rather than biological correctness.
Scientific publication3 min read
A mathematical framework for heterogeneous, constrained agents reached consensus in reported simulations and had the lowest listed cost in comparisons with 50 and 100 robots.
Scientific publication4 min read
The fixed-round PTD model reported sharp speed gains and higher scores across several video benchmarks, while leaving disjoint events and multiple matching targets largely unexplored.
Scientific publication4 min read
A preprint describes one OCR system trained across 13 Indic scripts; its reported overall character error rate was 6.9%, compared with 8.6% for monolingual models.
Scientific publication4 min read
A theoretical and simulation study suggests that private mean estimation can outperform non-private estimation over time when leakage-participation feedback is strong enough.
Scientific publication5 min read
A preprint presents URIUM as an open, adaptable compiler-design course built around a small language, a staged compiler and several assembly backends.
Scientific publication3 min read
Tests across nine systems found a sharp drop in numeric precision as prompts asked for more objects. Layout, composition and appearance also mattered, but high-count results were harder to validate.
Scientific publication5 min read
A preprint by Junmin An and Jon-Lark Kim gives exact formulas for the shortest self-orthogonal and LCD embeddings of linear codes over the ring Fq + uFq.
Scientific publication5 min read