Text Models Spot Evasion Better Than They Read Vocal Confidence
Models performed strongly on written signs of evasion but struggled more with vocal unconfidence, and speaker-level calibration helped only modestly.
Scientific publication4 min read
Specialty
Models performed strongly on written signs of evasion but struggled more with vocal unconfidence, and speaker-level calibration helped only modestly.
Scientific publication4 min read
A new benchmark finds that language models often handle the meaning of Chinese internet neologisms better than the sounds, characters and source forms used to create them.
Scientific publication5 min read
H-Scale uses calibration activations to choose hardware-valid NVFP4 scales, with higher average scores reported across Qwen3 and LLaMA tests but a narrow evaluation scope.
Scientific publication4 min read
A new preprint benchmarks language models on questions aligned with CFA Levels I to III and FRM Parts I to II. Leading systems exceeded 97% on Easy items but fell sharply on Hard cases, while the gated result was 0.39 percentage points higher on held-out questions.
Scientific publication5 min read
A preprint reports better support-versus-attack retrieval in tests of three embedding models, alongside evidence of reduced sensitivity to the claim topic in some fine-tuned comparisons.
Scientific publication5 min read
A Bangla question-answering system reported strong results on a manually built test set, while a small student review found high satisfaction. The study did not test patient outcomes, clinical safety, or performance on an external benchmark.
Scientific publication4 min read
A benchmark evaluation of FIRE, a two-stage system that analyzes hate categories before generating replies, reported higher factuality and category accuracy, lower toxicity, greater diversity, and lower peak memory use.
Scientific publication5 min read
The preprint reports that CaRGo-T scored above prompting baselines on benchmark tests of multimodal humor understanding and detection, with gains varying by model and task.
Scientific publication4 min read
A benchmark of 31,316 Sanskrit verse, translation and glossary examples reported the strongest scores for instruction-fine-tuned models, while errors exposed persistent difficulty with phrase boundaries and morphology.
Scientific publication5 min read
An unsupervised system called GUIDE reported lower misspelling rates in a 10-day production comparison and stronger offline scores on two Chinese query datasets, though the study does not establish that the system caused the online gains.
Scientific publication4 min read
A four-language benchmark asked systems to mark hallucinated character spans in image-conditioned text; leading scores were moderate, rankings were uncertain, and empty predictions were common.
Scientific publication5 min read
A prompting method called CritICL uses failure patterns from weaker language models to guide stronger ones at inference time. Tests on Qwen and Llama systems reported modest gains over selected baselines and lower generation and token use, but the results varied across benchmarks.
Scientific publication4 min read
A preprint study found that stylometric features separated human-written from AI-generated text well, but AI-edited text was more often mistaken for human writing and was best separated with editing-aware signals.
Scientific publication5 min read
A preprint reports high detection scores on several adversarial text benchmarks, while results varied across model versions, target generators and attack settings.
Scientific publication3 min read
A benchmark of LLM-generated GPU code found a 2.11× speedup for GPT-5.5 on 22 analytical database queries, while results varied by model and query.
Scientific publication4 min read
A clinician-adjudicated mental-health slice of HealthBench put five leading AI models into a statistical tie. Cross-judge rankings were nearly identical, while only two models produced reproducible empty refusals.
Scientific publication3 min read
A new preprint introduces a dataset and evaluation framework for testing how language models question patients, gather evidence and arrive at a diagnosis across multi-turn conversations.
Chouayfati, Pia, et al. "MTDiag: A Multi-Turn Diagnostic Dataset Towards Clinically Meaningful LLM Evaluation." Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 20264 min read
A preprint study argues that sentence-level negative log-likelihood can make cross-language model comparisons less sensitive to tokenization and other surface differences, although important limits remain.
Scientific publication4 min read
A training-free framework called PACE front-loaded retrieved evidence and adapted its reranking budget, with higher reported recall and lower latency in the tested simulations.
Scientific publication4 min read
A modeling study of networked language-model agents found that rival persuaders were associated with more back-and-forth belief trajectories, while direct and peer-relayed exposure tracked with movement toward a persuader’s position. The result applies to a closed simulation, not to humans or real platforms.
Scientific publication5 min read
An analysis of four transformer language models found that token geometry carried a clear signal about grammatical class, while the layer pattern varied by architecture.
Scientific publication4 min read
A three-model comparison gave BanglaBERT the strongest Macro-F1 and external-test results, while BanglaMamba was faster and used less GPU memory.
Scientific publication6 min read
A preprint reports that GRIN, a mixed-policy reinforcement-learning method, performed better on harder questions that required combining supplied facts or drawing inferences while keeping direct recall comparable.
Scientific publication4 min read
A computational study reports much longer model outputs after targeted expert deactivation and software-applied bit flips, while its proposed hardware step remains untested on a live system.
Scientific publication4 min read
An arXiv preprint reports benchmark comparisons of a lightweight router that selects typed graphs or natural-language handoffs, with lower token use under the reported accounting.
Scientific publication4 min read
An arXiv preprint reports more reproducible generated reports and fewer harmful directional and statistical-result errors when claims are bound to structured evidence, while warning that the underlying analyses remain unvalidated.
Scientific publication4 min read
A mathematical and computational analysis of the J-lens reports stronger average scores for variants focused on a small, high-energy set of positions, but masking did not cleanly separate short-horizon prediction from sparse-concept recall.
Scientific publication3 min read
A modeling study reports that speech-act-informed systems were strongest in low-data tests, while the leading architecture changed across datasets.
Scientific publication4 min read
The preprint finds that adaptive bias correction can preserve more answer accuracy than fixed-timing reflection in some language models, while poorly matched signals may create new errors and extra computing costs.
Scientific publication4 min read
A Chinese-language physics benchmark reports that model accuracy peaks in junior-high material and declines toward university-level questions. It also finds a wide gap between automated and human ratings of generated physics diagrams.
Scientific publication3 min read
A masked-diffusion framework called DCGC reported an average accuracy of 24.8 and the best score on five of six benchmark sets selected from problems an initial solver failed to solve correctly.
Scientific publication4 min read
A new benchmark suggests theorem-proving models can answer many mathematical questions in ordinary language while failing far more often when required to produce proofs that pass Lean 4.
Scientific publication5 min read
A preprint reports a ROC AUC of 0.960 for a Vietnamese AI-text detector, but its benchmark does not replace human judgment on authenticity.
Scientific publication4 min read
A preprint reports that ClueWeaver reached 59.0% overall accuracy and led local methods on all four evaluated long-context narrative benchmarks.
Scientific publication4 min read
A preprint describes TOPAS, a scheduler that coordinates cached agent context and request admission, with reported gains in synthetic and MetaGPT workflow benchmarks.
Scientific publication4 min read
A benchmark of seven vision-language models found sharp accuracy differences when first-person video and dialogue conflicted, along with uneven decisions about assistance.
Scientific publication4 min read
A new arXiv preprint reports an association between lightweight latent-vector steering and higher model-rated emotion scores, alongside lower semantic-preservation scores at stronger settings.
Scientific publication5 min read
A version 1 arXiv preprint reports that selected generative language models made pairwise choices that closely matched human preferences on French ASR hypotheses, while embedding metrics depended on model and setup.
Scientific publication4 min read
A study of Chinese sentence-level metaphor identification found different strengths for fine-tuning and zero-shot prompting with an expert-informed Skill.
Scientific publication4 min read
A model trained on 36 English hate-speech datasets posted the highest average score in the study's main comparison. On cross-lingual tests, the full English-trained model generally scored below zero-shot results, while a minimally trained version was stronger on some tests.
Scientific publication4 min read