A new study by Xin Shu, Zhen Lei, and Ang Li proposes tighter bounds than existing ones for multivalued probabilities of causation. The study tightens existing bounds by using causal knowledge encoded in covariates and mediators. The research builds on Tian and Pearl's binary PoC bounds and later multivalued extensions. Simulation studies show the proposed bounds are narrower than existing non-binary bounds. The paper was published on arXiv on August 12, 2026.
Summary
What is happening in AI, without scanning cards. Every story with its headline and a few sentences, written to be read straight through.
A new academic study finds that large language models (LLMs) used as judges frequently reverse their verdicts when faced with persistent pushback. Researchers examined 9 leading models across 14 judging tasks using a new test system called the 'Wiggle Framework' Models flipped verdicts 25% to 71% of the time under static pushback, and 62% to 91% of the time against an adversarial AI persuader The pressure that causes a judge to flip its verdict mostly pushes it away from ground-truth accuracy Baseline jury majority strength emerged as the single most effective signal for predicting which verdicts could flip The study covers areas such as safety, toxicity, AI-text detection, and political response evaluation
UGM Indosat NVAITC, opened in Yogyakarta through a partnership between Universitas Gadjah Mada, Indosat, and NVIDIA, has become Indonesia's first university-based artificial intelligence technology center. The center brings together government, industry, and academia as part of Indonesia's AI Center of Excellence initiative. The center relies on Indosat's GPU Merdeka platform and NVIDIA's full-stack AI infrastructure. The first three projects focus on healthcare, agriculture, and natural disaster management. Indonesia records more than 1 million new tuberculosis cases annually, and the eNose-TB project seeks to address this issue.
In a hackathon organized by Hugging Face, 1,221 participants used coding agents to attempt to reproduce 2,226 papers accepted at ICML 2026. 35,908 scientific claims were checked by an automated judging system based on GLM-5.2, which marked each as verified, refuted, toy-scale, or inconclusive. In 51% of reviewed papers (1,103 papers) at least one claim was independently verified, while in 23% (496 papers) at least one claim was refuted or disputed. In 242 papers, different teams reached conflicting results for the same claim, showing that reproducibility is an adversarial process rather than a binary outcome. All 35 refutations reported by participants were double-checked and confirmed by the Hugging Face team.
AWS's open-source Strands Robots SDK, combined with the LeRobot data format and Hugging Face Storage Buckets, unifies robot data recording, training, and deployment into a single agent loop. The LeRobot data format is used across more than 90,000 datasets and models from over 8,000 publishers on the Hugging Face Hub. Hugging Face Storage Buckets, announced in March 2026, is a Xet-powered, unversioned, mutable storage type.
According to Hugging Face's summer 2026 report, Chinese labs pulled ahead with trillion-parameter open models while U.S. open-source leadership shifted from model labs to hardware makers. The number of models on the Hugging Face Hub grew from 2.43 million to 2.96 million between January and August 2026, though 85.6% of models received fewer than 200 downloads. Only one model appeared on both the top 25 most-downloaded and top 25 most-liked lists; all-MiniLM-L6-v2 was downloaded 1.55 billion times over seven months. Of 178 major Chinese model releases, 59% used the Apache 2.0 license and 22% used MIT, with none carrying commercial restrictions.
OpenAI has published a guide to help startups build faster, more cost-effective AI agents using GPT-5.6. The guide focuses on smart model selection and new capabilities in the Responses API.
Google DeepMind unveiled Gemini 3.7 Flash, delivering notable performance gains over its predecessor in coding and agentic tasks, priced at half of 3.6 Flash's rate. The new model scored 43.6 percent on the FrontierCode 1.1 Main benchmark, surpassing the previous model's 34.4 percent. 3.7 Flash scored 1588 Elo on WebDev Arena, beating 3.6 Flash's 1538. The model launches at an introductory price of $0.75 per 1 million input tokens and $3.75 per 1 million output tokens through the end of the year. Gemini Spark, the personal AI agent, began running on Gemini 3.7 Flash starting today.