A preprint introduces egRUE, a method that combines uncertainty scoring with feature-level explanations; its reported gains came from benchmark tests and a small BloodMNIST expert study.
Gibbs-family methods led the aggregate scores in Traffic Hourly, Electricity Hourly and Solar Weekly, but classical baselines remained on top in some M4 frequency and disagreement groups.
VICT uses a task's terminal verifier to trace credit through long action sequences instead of using it only as a final pass-or-fail signal. In ALFWorld and WebShop tests, it beat GRPO and achieved higher validation AUC over the same 300 updates, while the authors caution that verifier-defined links do not establish causal necessity.
A mathematical study links the worst-case rank needed to approximate normalized attention to support geometry, while a fixed BERT-base calibration reports lower effective dimensions in some tested cells.
A theoretical and simulation study suggests that private mean estimation can outperform non-private estimation over time when leakage-participation feedback is strong enough.
D-TAIA combined parameter-efficient language-model adaptation, domain-aware pre-training and retrieval, then matched or improved on LLM and RNN baselines across four event logs.
Spectral and frequency-based features gave the clearest signal for detecting respiratory events in ballistocardiography recordings, with nonlinear classifiers performing strongly when each patient was held out for testing. The results point toward compact systems, but remain limited to one hospital-based sensor setup and cohort.
A preprint reports higher average accuracy, lower variability and better average ranks for a residual-guided neural-network procedure tested against RVFL, ELM and BLS baselines on 71 UCI datasets.
A computational preprint reports that TransMod had the lowest listed forecasting errors in New York City and Chicago and remained competitive when ride-hailing data were limited.
A mathematical preprint reports a proof of Colombo's determinant conjecture for even-dimensional difference-power matrices with pairwise-distinct real nodes. It also gives an exact rank formula and determinant sign rule, while the formal verification described covers the odd-exponent derivation.
A computational study reported zero mean final error for Mycelial Search on one benchmark function at two tested dimensions and the lowest reported mean on another at the higher dimension, while other tests showed uneven performance.
A preprint describes a mixed-precision approach that assigns bit widths layer by layer using sensitivity estimates and reports quality gains near three-bit compression, while AWQ remains faster in the tested speed comparison.
A preprint proposes a five-stage framework for systematic trading that charges model complexity and search effort against an effective-sample budget. Its one-index test found no robust position, and the authors say the work does not establish profitability.
A new computational study reports that deeper lookahead can make Whittle-index approximations more accurate for partially observable restless bandits, with a measured runtime cost.
An analytic example shows that an exact FullCP region can be star-shaped without being convex, while a two-dimensional simulation found robust nonconvexity witnesses often but with small normalized gaps.
PRQ-KMeans reported higher retrieval scores than RQ-KMeans in industrial search and across four public recommendation benchmarks, alongside higher codebook-use measures.
A preprint presents OPDVR, a gated on-policy distillation method that combines teacher guidance with verifiable reward. It reports higher average accuracy than sampled-token OPD in same- and cross-architecture tests, alongside a GRPD extension that also outperformed GRPO on the reported benchmarks.
A preprint's KENDO method replaced MCMC hyperparameter sampling with a weighted kernel ensemble and produced leading benchmark results with less per-iteration computation.
A preprint describes a model that selectively shares information among disease-specific mobility outcomes. It beat eight comparison methods on held-out prediction scores, but exploratory disability classifications were not validated for prospective clinical use.
A computational case study found that more varied first tokens in refusal completions were associated with higher stable rank and smaller refusal changes after a target-set-adapted single-vector ablation. The pattern appeared in frozen-model analyses and controlled fine-tuning, but the study did not establish transfer to unseen prompts.
An experiment found that per-entity LoRA adapters could answer held-out single-answer questions without source text, while the tested ways to select the correct adapter from a query failed.
A methods study reported faster runs in two digital-evolution workloads, but also showed why runtime tails, hardware anomalies and migration bias need monitoring.
A new policy for language-model agents uses task information and signs of trouble during a task to decide when to make a permanent switch to a stronger model. The preprint reports better success-cost results than several routing strategies, while showing a less decisive comparison on a small custom benchmark.
The analysis gives HFHRMC explicit convergence and iteration bounds for nonconvex sampling targets, but its certified advantage is confined to a lower-order finite-accuracy term.
A preprint evaluates the Frame Kernel Method for multiscale PDE surrogate modeling, reporting the lowest error on three of four tests in each benchmark set while a proposed doubled convergence rate remains a conjecture.
A memory-augmented model scored higher than a Transformer across three neural decoding benchmarks when training data were limited, while results varied with sequence context and model size.
A four-model test found that decode-phase GPU energy rose 16.98% to 17.92% for two MHA models as context grew from 128 to 1,800 tokens, compared with about 3.32% to 3.62% for two GQA-based models. Increasing batch size from one to eight was associated with more than 80% lower energy per generated token in nearly all configurations.
A computational study reports that GRAPE, a local Bayesian-optimization method, used fewer attack queries and finished with the lowest prompt regret across four embedding dimensions.
An evaluation of Flower Hub found task-dependent results across five federated-learning benchmarks and reported that the same application could run in simulation and deployment without source-code changes.
A preprint combines representation learning with conditional generation and reports stronger agreement between related outputs in a low-dimensional synthetic benchmark. A shuffled-condition test was consistent with condition use, while mode balance remained imperfect and transfer to natural data was not established.
A nine-player modeling study evaluated PART, a multimodal framework that combines several tennis data streams to estimate wellness, injury risk, physical capability and playing style. Its reported scores apply to this dataset, with no external validation population reported.
2025 IEEE 5th International Conference on Human-Machine Systems (ICHMS), pp. 28-34, 20254 min read
A study of time-series forecasting tests when text-based context adds information beyond the latest observations and finds sharply different results across datasets.
A new arXiv preprint reports that attack transfer was associated with differences between attacker and victim models and their client data. In simulations on CIFAR-10 and SVHN, its proposed defense generally scored higher against four transfer-based attacks than Federated Adversarial Training, while clean accuracy remained competitive.
The matched benchmark results favored different planners on different discrete tasks, while BFN-RL remained competitive in the reported continuous-control comparison.
A proof-based arXiv analysis finds matching worst-case bounds for alternating regret in the expert problem and general online convex optimization. Its broader lower bound uses a constructed compact convex domain, so it does not establish the same rate for every fixed domain.
A variational-learning method recovered interaction and environmental forces from simulated trajectories, with its semi-parametric version often estimating environmental forces more accurately. It also selected mechanisms and preserved collective behavior in larger simulated populations.
A 44-sample computational benchmark reported higher PDE-term recovery scores from compact physical measurements than from raw field slices, but it used only noise-free simulations.