A case study of a 5 km Destination Earth simulation compares operational, active-only and life-cycle accounting, showing how each boundary changes the reported energy and carbon totals.
A 20-person comparison reported higher task-success odds, shorter completion times, lower perceived workload and stronger expert ratings with TailorCoPilot than with skill-appropriate baseline tools.
A review maps 230 AI systems and 29 construction benchmarks, finding that evaluations cover final artifacts more reliably than trajectories, use outcomes, or performance beyond tested cases.
In a cross-sectional survey at one teacher-education university, students reported more weekly AI use and higher usefulness scores, while faculty and administrative staff reported stronger integrity concerns. The findings describe group differences and associations, not causes.
Interviews with geometry teachers described a reported split between digital exploration and paper-based formal proof. A separate review of 33 tools found that no reviewed tool covered every activity required for geometric-proof instruction.
A small pilot tested whether an AI tutor could judge student explanations of worked code and guide revisions. It found stronger agreement on correctness than on completeness, but too little posttest data to assess learning.
A Netherlands laboratory study found that participants often reached account settings or companion-app updates while following generic smart-home security advice, then judged the task complete.
A test of 11 language models on 208 rare-disease vignettes found Justice ranked first in every model, with decision-maker framing more closely associated with ethical-value selection than patient type, age or model identity.
A new dataset compares human ratings and AI model performance on asynchronous video interviews. The best multimodal result was only modestly ahead of the best text-based result, while ratings from targeted questions were more reliable than ratings from generic questions.
A two-stage system recorded lower reported harmful-compliance rates than raw generation in a 1,000-dialog synthetic benchmark, while automated-judge variation limits what the result can show.
A preprint proposes a six-attribute, 29-tier framework to make CBRN and offensive-cyber misuse evaluations more explicit before frontier AI models are released.
A 16-person comparison found broader document coverage and fewer concept-map errors with direct browsing, while AI question answering showed more tightly connected exploration.
A small human evaluation found that user-rated odor matches varied by scent and capture material, while stronger perceived replays tended to receive higher similarity scores.
A three-month workplace preprint reports a modest rise in self-reported PERMA scores alongside daily reflection-tool use, with mixed workshop and word-diversity findings.
In a 12-person comparison, Maru's persistent information-architecture rules were associated with steadier approval and more varied layouts, while cross-topic retrieval drew more negative feedback.
A preprint tested four forms of AI support in email and bird classification tasks. Concept displays were more accurate than no support in both tasks, but full-sample results did not show a consistent advantage over label-only support.
A simulated benchmark reports different emotion-recognition scores under matched observation conditions. The active condition scored above the initial view in most configurations but recovered only part of the reference-condition gap.
The 18-item GenAIT test measures conceptual generative AI knowledge and appears promising for group-level research, while its reliability varies across students’ score levels.
A qualitative arXiv preprint finds the assumption that non-great-power conflict matters far less than great-power conflict poorly supported for war and catastrophic terrorism, while treating AI loss of control as a speculative priority.
Six vision-language models showed more agreement when ranking the emotional qualities of 3D objects, but their results varied sharply by category and were not checked against human judgments.
Interview accounts describe Discord as a place where teenage gamers coordinate play, exchange game information and sustain social ties beyond active gaming, while also reporting scams, risky interactions and fragmented moderation.
Elena Koung, Xinning Gui and Yubo Kou. 2026. Gaming Together on Discord: Teen Gamer's Cross-Platform Practices. Proc. ACM Hum.-Comput. Interact. Vol. 10, No. 7 (November 2026), 32 pages4 min read
A simulated Shenzhen neighborhood produced marked differences in modeled travel burden across resident profiles, while average ratings for healthcare access and trip chaining were lower than ratings for travel burden and activity feasibility.
A three-year report on a University of Arkansas course describes a project-based AI curriculum for mechanical engineers. Grade ranges rose descriptively from Fall 2021 through Fall 2023, but changing instruction and incomplete cohort data mean the report cannot show that the curriculum caused the shift.
The proposed task-disentangled method was built to let one framework handle different audio-visual tasks with less cross-task interference, and the paper reports strong benchmark results alongside a clear AVQA caveat.
EVEREST, a two-stage vision-language system, recorded the strongest reported overall cIoU and F1 on SocioSeg, a benchmark built from satellite images, aligned maps and pixel-level masks.
A computational review of Wikipedia articles about modern wars found that language editions can describe their own side and an opposing side in markedly different ways. The strongest differences appeared in power and agency, while accounts written by communities not directly involved in a conflict were broadly similar.
A preprint evaluation found GGSS had the lowest average bias-change score on all four tested models, while MMStar accuracy stayed close to the unsteered baseline.
An exploratory study maps the language and ideas people use to make sense of AI, finding broad debates over development methods, the nature of AI systems and the pace of progress.
A preprint reports that a compact generative model used fewer parameters and substantially less sampling compute than a matched U-Net baseline while recording better scores on several precipitation-retrieval measures. The evidence came from two benchmarks and a qualitative China-scale case, while reported wall-clock latency was slightly higher.
Across a field study and an online experiment, people who felt their privacy expectations were met tended to report greater satisfaction and stronger intentions to keep using their devices, while blocking intentions were inconsistent.
A new preprint proposes a broad framework for teaching and assessing visualization design. The authors built it from instructors’ course objectives and feedback, while stressing that the resulting assessment remains an initial framework rather than an established measure of student knowledge.
In a 33-volunteer study, an interactive 3D graph viewed in virtual reality was compared with a PowerPoint presentation for retrieving linked engineering-model information. The virtual-reality condition took longer, while the other measured outcomes showed no significant differences.
An analysis of two unlinked surveys from active Fledge.Love users found that receptivity to deploying one’s own conversational agent and receptivity to encountering another person’s agent are separate but strongly related. The data do not show how AI agents affect real-world dating behavior.
A methods preprint reviews 66 Chinese mental-health EMA studies and proposes a 43-indicator framework for assessing research platforms. In a Huixin EMAI case, 1,893 of 2,472 prompts were completed, but the study also found missed notifications, uneven platform capabilities and unresolved researcher-side gaps.