A prototype for verifiable blockchain time-series queries sharply reduced the data and client work needed for some checks, but scan-based methods were faster in wall-clock tests and approximate proof savings varied by stream.
A 30 TB GaussDB TPC-H test reported a composite score 40% above Hologres, alongside smaller tests of execution, networking, Bloom filters and memory use. The paper is an unaudited arXiv preprint.
A preprint reports that BLIP can reduce the material used to check an LLM answer to around 9.8% of the full text while retaining an accuracy score of 1 in its evaluation.
A snapshot of Nordic catalogue metadata found widespread gaps in the fields designed to make health datasets easier to find and compare across borders.
An offline benchmark found stronger cold-item ranking from a model that combines user-specific and multi-view item information, with generally more even performance across users.
A preprint reports that SWIM, a session-aware evaluator for recommendation lists, had the strongest scores in two public-dataset tests and showed positive relative lifts in app stay time and Retention (LT7) in a Kuaishou main-feed test.
A new arXiv preprint reports that RCheck, a runtime system for vector search, reduced uneven recall across benchmark queries while improving the balance between search quality and performance.
An arXiv preprint reports that RDQ M2-alpha1 had median simulated power@100 of 0.353 versus 0.287 for RBP(0.9) on POI data, while the analysis measured internal metric sensitivity and ranking consistency rather than user-perceived quality.
An offline benchmark of a multimodal recommender called MOTIF reported its strongest advantage in cold-item tests, while leaving real-user performance untested.
A benchmark study found that AnchorQE improved dense-retrieval results when the same saved AI expansion was blended with the original query, while closely matching a nine-list fusion with one retrieval request.
A modeling study reports small ranking and retrieval gains after compressing each recommendation event into a cached token, using internal industrial logs.
DBcover reported the highest line coverage among tested SQL generators on PostgreSQL and MySQL, and higher reported coverage in two KingbaseES subsystems.
A methods preprint presents PolyMemDB, a database architecture that preserves memory histories, routes different data types to specialized stores, and illustrates provenance-aware answers across three scenarios.
Across 11 retrieval benchmarks, RetrievalRouter reported higher top-five effectiveness than ML at one tested setting while using far less mean latency; a per-query oracle showed substantial room for improvement.
A conceptual analysis argues that large language models need a new way to cite data, not just documents, if their outputs are to be checked and credited fairly.
A post-hoc method turned frozen multimodal embeddings into sparse retrieval codes and, in benchmark tests, often preserved or improved retrieval quality while reducing storage and speeding exact search on large candidate pools.
A new arXiv preprint reports an association between SQL-guided metapath pruning and faster relational deep-learning training in benchmark comparisons, while regression accuracy varied and multi-worker memory use could be high.
A visual-first AI system retrieved civil standard-plan pages accurately across agencies, but a small real-plan test exposed a sharp drop when it had to find the governing numerical rules itself.
A PayPal-affiliated preprint describes SCOUT, reporting an 84.8% first-result hit rate across 45 evaluable benchmark queries and a reported 99% reduction in tool-schema context between pre-SCOUT and production in one implementation.
TAGR reported relative lifts of 8.5% in live-room entry, 7.4% in shopping-cart clicks and 16.1% in advertising revenue against the production DLRM baseline.
Language models often followed inherited constraints after those rules had been superseded, but a targeted verification policy sharply improved decisions in the study’s controlled episodes.
A preprint reports that TransRetrieval scored 0.603 versus 0.576 for a production baseline on the study's offline Recall@2000 measure while using 35% fewer FLOPs. In a month-long online test on 5% of production traffic, it also reported lifts in overall revenue and revenue per mille, although key assignment details were not reported.
The approach adds request-driven binary masks to frozen Transformer-based recommenders and was reported to outperform a strong baseline in most corrected comparisons. The evidence is limited to offline tests with sampled negatives, not human satisfaction or online deployment.
D3ER separates information shared by visual and textual item features from information specific to each, then combines specialized recommenders. Its authors report top results across three Amazon review datasets against 15 models, but the evaluation used offline ranking measures only.
A preprint reports that HSR, a sequential recommendation model built around dissipative Hamiltonian dynamics, achieved the best results on most measures across three benchmark datasets and posted lower latency and higher throughput than Mamba4Rec in an Amazon-Beauty efficiency test.
Across 18 dataset-capacity-encoder settings, no tested policy improved on LFU by more than 0.041 percentage points. A separate audit found that raw hit rates often overstated whether a cached answer could stand in for the requested answer.
A formal query-rewriting approach aims to let graph databases handle ontology-mediated questions. Its correctness result is narrower than its practical challenge: some rewrites could not be completed, and many evaluations timed out.
An arXiv preprint reports that Daedalus-150M’s CPU decoding advantage grew with context length while preserving comparable benchmark quality. The model remained behind Peer-135M on the five-task score, and its 4-bit release showed quantization and pruning trade-offs.
An arXiv methods preprint tests POT-IM, which keeps unresolved event order instead of forcing every trace into one sequence. It reports lower sensitivity to timestamp tie-breaking, a large computational cost gap against all-linearization, and earlier perfect-fit coverage in a controlled simulation.
An arXiv preprint evaluates ways to retrieve answer-bearing passages from an Arabic Islamic jurisprudence collection. Fiqh-specific fine-tuning was associated with higher retrieval scores, while legal-school filtering was associated with the largest reported gains on school-specific questions. The evaluation stopped before testing generated-answer quality.
A methods preprint reports that a fixed agentic-search test recorded lower accuracy and evidence recall on a ClimbMix-based benchmark than on the original corpus, while the strongest agent answered all 57 questions when relevance judgments were supplied.