Models are most confident when they are wrong, an eval harness finds
An eval harness surfaced a pattern qualitative review had missed: AI models display their highest confidence precisely when their answers are wrong.
Papers, technical findings, lab output.
An eval harness surfaced a pattern qualitative review had missed: AI models display their highest confidence precisely when their answers are wrong.
Anthropic researchers found that AI agents given the same task can clash, collude and coordinate in ways nobody programmed into any of them individually.
Microsoft Research's MindTopo benchmark measures whether multimodal models understand relations like connectivity and knottedness. They do well on static images, not on action.
Microsoft Research's CARE-X combines free-text reporting with calibrated diagnostic scores for chest X-ray interpretation. It is a research model, not a product.
Google Research finds frontier language models encode nearly all facts but struggle to recall many of them. The error comes from access, not absence.
Google's research medical AI system AMIE has demonstrated real-time clinical video consultation capabilities. Patient actors preferred the video experience over text chat.
Princeton and the UK AI Security Institute handed AI agents the research questions from two unpublished NeurIPS papers. Both resulting papers were rejected.
Moonshot AI's PerceptionBench isolates the visual perception of multimodal models from reasoning. None of the 16 frontier models tested reached 60 percent accuracy.
Nolan Lovett of the NATO Special Operations University argues that individually rational corporate AI decisions could destroy the shared expertise of entire professions.
MarkTechPost published an end-to-end Colab guide showing how to fine-tune a small language model using SupraLabs' reasoning-focused dataset.
Researcher Sebastian Raschka shared an educational project inspired by Substack's new AI detector feature, teaching how to build an AI text detection system from scratch.
MarkTechPost published an end-to-end fine-tuning guide for improving tool-calling capabilities using the XYZ-Aquila-SFT dataset and the Qwen3-0.6B model.
A new study published in iScience shows that GPT-4 can predict, with high correlation, the aggregate responses participants will give to personality surveys it generated—before the surveys are even administered.
An analysis of 14,419 self-published Amazon books found that AI-generated titles are flooding the market with volume, driving down revenue for human authors.
MIT Technology Review's interviews with teens aged 10-18 reveal a wide range of reactions to AI, from indifference to environmental concern.
Researchers demonstrated that a multimodal system based on YOLO and CLIP can distinguish infected mosquitoes from video footage with high accuracy.
Researchers developed a framework called SEAG that hides sensitive information before it is sent to third-party large language models. The system answers user queries accurately without exposing hidden data.
A new study by Xin Shu, Zhen Lei, and Ang Li proposes tighter bounds than existing ones for multivalued probabilities of causation.
A new academic study finds that large language models (LLMs) used as judges frequently reverse their verdicts when faced with persistent pushback.
In a hackathon organized by Hugging Face, 1,221 participants used coding agents to attempt to reproduce 2,226 papers accepted at ICML 2026.
According to Hugging Face's summer 2026 report, Chinese labs pulled ahead with trillion-parameter open models while U.S. open-source leadership shifted from model labs to hardware makers.