An arXiv preprint presents a unified Rényi framework for composite binary testing, combining finite-sample bounds with asymptotic results that separate two Type II error regimes.
Benchmark tests on three RVV 1.0 RISC-V CPU configurations found that juFFTe's gains varied by hardware, with a reported threefold average speedup on the SG2044 multi-core comparison, while AMD Zen 5 led across the full range.
A theoretical preprint reports that fixed-order Khatri-Rao sketches can achieve a near-linear subspace-embedding dimension in the subspace size, improving the stated dependence on that size over earlier bounds. The result is a proof under sub-Gaussian assumptions, not an experimental performance claim.
An evaluation in simulation and on a physical robot found that one shared interface could support three instruction modes, with gesture-based modes often performing better when wording, surfaces or objects changed.
An arXiv preprint proposes a 3D-grounded test for robot video models. Cosmos led composite scores, but detailed checks found weaker object localization and trajectory accuracy, underscoring the gap between plausible footage and executable behavior.
A theoretical result forces the number of cells in a bent partition to be a prime power and narrows its exponent using the dimension of the underlying finite space.
An arXiv preprint tests a multi-critic training method for robots that push and transport objects and open a dishwasher through contact. It reports 94.1% simulation success and 69.0% success in 58 trials on four unseen objects, while the dishwasher test is qualitative.
Across 120 slots per condition, first post-edit re-verification appeared in 78.3% of cadence-guided slots and 26.7% of cadence-omitted slots. The same descriptive comparison showed fewer cadence violations and more bounded final successes with the guidance.
In a randomized experiment, people working with a non-human-shaped robot disclosed more when its small talk contained fewer personal details, while teamwork and coordination ratings were also higher.
An arXiv preprint reports an LLM-assisted workflow for assigning application tasks to a processor or FPGA, with reported speedups up to 92.53 times and repeated partition choices in the tested configurations.
A four-level robotic bin-picking system cleared all 30 experimental bins, while the study found that bin clearance and individual grasp success told different stories.
A preprint by Junmin An and Jon-Lark Kim gives exact formulas for the shortest self-orthogonal and LCD embeddings of linear codes over the ring Fq + uFq.
A simulated comparison found higher success and shorter routes for an LLM-guided UAV than for a conventional lawn-mower scan, but the evaluation did not test unknown environments.
A preprint describes MaCoPlanner, a system that turns equipment manuals into structured knowledge, checks robot plans before action, and rejects unresolved candidates. It reported higher success on the hardest benchmark tasks than the listed baselines, but the physical evaluation was small and limited to a no-load simulator.
A theoretical preprint proposes algorithms for flexible graph connectivity across several nested tiers. Its main randomized method succeeds with probability at least one third for a fixed number of tiers, while a related multi-graph problem receives a factor-two approximation.
A prototype for verifiable blockchain time-series queries sharply reduced the data and client work needed for some checks, but scan-based methods were faster in wall-clock tests and approximate proof savings varied by stream.
A preprint study found that failures across LLM defenses were positively correlated, challenging the idea that stacking independent safeguards automatically multiplies protection.
The preprint maps how false channel-state information could alter power allocation, user ranking, decoding, fairness and secrecy in power-domain NOMA. It offers a qualitative threat taxonomy and impact analysis, but reports no empirical sample or dataset and does not quantify the resulting harm.
A 30 TB GaussDB TPC-H test reported a composite score 40% above Hologres, alongside smaller tests of execution, networking, Bloom filters and memory use. The paper is an unaudited arXiv preprint.
The CyberFactory framework links data construction, trajectory synthesis and model training, and OpenAegis recorded the highest reported CyberGym score among the compared models.
A case study of a 5 km Destination Earth simulation compares operational, active-only and life-cycle accounting, showing how each boundary changes the reported energy and carbon totals.
A robotics preprint reports higher task success for a visual-track system in simulation and physical manipulation, along with better future-video prediction scores.
A model-based system for wireless sensor networks produced semantically enriched measurements, while most micro-ontology generation took place within 500 microseconds. Compression depended on the serialization format, and the reported engineering results came from an illustrative robot-based setup.
2020 46th Euromicro Conference on Software Engineering and Advanced Applications (SEAA), Portoroz, Slovenia, 2020, pp. 561-5684 min read
A theoretical analysis finds that generalized covering codes with product-form centers reach the ordinary sphere-covering rate, while linear codes over prime-power alphabets reach the same asymptotic benchmark.
In simulations and controlled hardware experiments, MagPie completed all reported first rendezvous, while variable energy and imposed clock-loss stress reduced delivery in modeled networks.
A 20-person comparison reported higher task-success odds, shorter completion times, lower perceived workload and stronger expert ratings with TailorCoPilot than with skill-appropriate baseline tools.
A preprint audit of 12 physical-AI benchmarks found positive relationships across every pair, with especially strong agreement for two pairs of tests. The analysis covered 51 models and suggests that benchmark averages can repeat the same strengths, while a compact suite may preserve useful model separation.
A two-specimen laboratory study found that a distributed measure of vegetation stiffness stayed stable across robot push heights, while a lumped stiffness varied.
A preprint describes SpecMine, a public GitHub corpus of software specification artifacts that combines broad and Kiro censuses with commit histories, pull requests and typed links to code.
A preprint reports that BLIP can reduce the material used to check an LLM answer to around 9.8% of the full text while retaining an accuracy score of 1 in its evaluation.
The arXiv work claims a deterministic polynomial-time certificate for every nonnegative rational square matrix, with an exponential approximation base below the canonical Bethe guarantee.
FlashVLA, a streaming method for vision-language-action policies, cut reported inference time while maintaining or improving task success across selected simulated benchmarks and three physical-arm tests.
A snapshot of Nordic catalogue metadata found widespread gaps in the fields designed to make health datasets easier to find and compare across borders.
An offline benchmark found stronger cold-item ranking from a model that combines user-specific and multi-view item information, with generally more even performance across users.
APEW-Fed stayed 2.7 F1 points below a centralised baseline on the main ERP test while reporting 94% lower communication than FedAvg, but its formal privacy accounting and simulated attack tests answer different questions.
A review maps 230 AI systems and 29 construction benchmarks, finding that evaluations cover final artifacts more reliably than trajectories, use outcomes, or performance beyond tested cases.