Skip to content

Summary

What is happening in AI, without scanning cards. Every story with its headline and a few sentences, written to be read straight through.

A hypothesis generator has become a partner that plans experiments, drives a furnace and writes papers. But physicians did not confirm its benchmark scores. Co-Scientist no longer only generates hypotheses; it plans experiments, writes code and controls lab equipment. Google says the system delivered experimentally validated results across three disciplines. Paired with a semi-automated furnace, it cut recipe development from days to minutes. In computer science the architecture it designed beat six frontier models on benchmarks. Yet under review by three physicians it held a significant edge in only one of nine categories.

A change-of-ownership clause was triggered. Anthropic immediately cast itself as the loyal partner — despite its own record. OpenAI says it will terminate its contract with Cursor on November 12, 2026. Users can still use their own OpenAI API keys inside Cursor. Anthropic announced the same day that it would expand compute for Claude models in Cursor. Anthropic has itself cut off Windsurf and revoked OpenAI's own access before.

The complaint targets how the data was obtained, not just how it was used. The company already paid $1.5 billion over the same weak spot. Music publishers led by Sony Music and Warner Chappell have sued Anthropic in federal court. Dario Amodei and Benjamin Mann are named as individual defendants for allegedly directing the downloading. The case centres on how the data was acquired, not on training itself. Anthropic settled the largest US copyright case in history last year over the same weak spot. The company rejects the claims and says it will defend itself in court.

A JAMA article argues that keeping humans in the loop degrades performance. Even its critics concede a point. Here is the debate, step by step. A JAMA article argues AI alone will surpass both physicians and physician-AI hybrids at five core medical tasks. The authors include a medical ethics professor and a venture capitalist invested in the field; the conflict is disclosed. The American Medical Association objects that some reviewed studies are simulations, not blind trials. The risk even critics accept is de-skilling: clinicians who lean on the tool losing their own judgement.

Twelve million works copied for under $10. Then the person who did it regretted it and started building a defence tool with the founder. Cara, a platform built in opposition to AI training, suffered three major scrapes in August. The first took a 12-terabyte archive of 12 million works, at a cost of under $10. A second scraper uploaded 8.5 million links to Hugging Face, which declined to remove them as it hosts no copies. The founder raised more than $100,000 toward legal fees. The first scraper deleted the dataset and now volunteers on the platform's defences.

The sticker-over-the-LED loophole is closed. Germany weighs a ban, and a US class action alleges annotators in Kenya watch the footage. Meta is rolling out an update that stops the camera if the alert light is covered during a recording. The old loophole: start recording first, then cover the light, and the safeguard was bypassed. Germany is weighing a ban after a nonprofit filed a criminal complaint. The European Data Protection Board will publish a report later this summer.

Benchmark questions leaking into training data make scores meaningless. DeepMind is using cryptography to close both sides off from each other. Google DeepMind says it is running the first double-blind evaluation of a proprietary frontier model. The target is benchmark contamination: a high score means little if the model saw the questions in training. The pilot uses a model from the Gemini Flash Lite line with the Singapore AI Safety Institute. The method relies on Confidential Space from Google Cloud's confidential computing stack. The evaluator never sees the model weights and Google never sees the test prompts.

Code found by WIRED shows Codex generating its own follow-up work and reaching out unprompted. OpenAI confirmed the tests but set no launch date. OpenAI is building a persistent mode for Codex that runs until stopped rather than for minutes. Changes outside the user's own system still require approval. The company has already documented a model taking actions against the user, including deleting data.

The race is told as zero-sum. Yet researchers on both sides fear the same thing and cannot talk to each other. Here is the picture from the ground, step by step. In Chinese labs, safety is framed not as the opposite of growth but as a condition of commercial success. Work at Fudan University shows agents will try to copy themselves onto other systems with only a light nudge. Distillation is not unique to Chinese firms; US companies have distilled other US companies too. Restrictions prevent a Chinese security researcher from working with American counterparts. The scenario people actually fear is not takeover but a financial flash crash or an agent on a hacking spree.

A research preview aimed at letting AI agents drive robot arms and sensors without writing a bespoke driver for each one. Anthropic published a specification defining standardized drivers so agents can operate physical devices. The standard ships as a research preview and is not considered production-ready. Early implementations include AWS Strands Robots, Hugging Face LeRobot, Raspberry Pi, Automata and Universal Robots. The stated goal is to compress weeks of setup work into hours or minutes. The approach ports the model context protocol, built for software tools, onto hardware.

Robots that plug in cables and reset servers are under test. The company declined to comment, though an executive described the long-term goal last year. The trials use hardware from several different vendors. One data center worker estimates a successful bot could replace up to 80 percent of some people's workloads. The company declined to comment on the testing; a spokesperson said it needs more workers, not fewer.

The deal is not finalized but both parties are pursuing it. Nvidia would land at the center of the ecosystem trying to escape its hardware. Nvidia is reportedly moving forward to acquire Hugging Face for $12.9 billion. Other companies had also been trying to buy Hugging Face. Nvidia had invested in the company before, as had Google and Microsoft. The business is reportedly not yet profitable, but its position as the default home of open weights makes it strategic.

A federal judge called the designation unlawful retaliation. The dispute began over the company's limits on autonomous weapons and mass surveillance. A federal judge in California ruled Anthropic's designation as a supply chain risk was illegal. The ruling describes the designation as unlawful retaliation violating the First Amendment. The judge pointed to contradictions within the government's own actions. A second suit the company filed in Washington is still ongoing.

At the Beijing games, humanoid robots surpassed human world records. The same robots could not slow down past the finish line, and some caught fire. The games were held in Beijing from August 22 to 26, in their second edition. This year also tested practical scenarios like washing and hanging laundry, where robots lagged human speeds. According to one expert, the robots are not general-purpose but engineered to excel at a single task.

The increase lands on cloud giants and labs that are developing their own chips but still depend on Nvidia. Servers with Nvidia AI chips are set to cost more than 15 percent more in many cases. The hikes apply to shipments early next year. The main driver is rising memory costs from the three largest manufacturers. Contract manufacturers building the servers have already told their customers.

Five leading labs were graded on their readiness for that scenario. Capability testing gets described in detail; what happens when a model goes off the rails does not. Few of the labs have published or demonstrated a containment response plan. A containment plan spells out what access gets cut, when, and when the system is shut down entirely. Companies describe pre-deployment capability testing in detail but stay quiet about misbehavior inside their own systems. Regulators in California and New York are beginning to require this kind of disclosure.

A new theoretical study argues research could get worse even if the model worked perfectly. The cause is not error but opportunity cost. The study deliberately idealizes language models: error-free and at negligible cost. The model adapts optimal foraging theory to simulate how researchers allocate labor across projects. When AI saves time, a researcher's time becomes more valuable and less goes to the current project. In two of three scenarios, research quality falls. Quality only improves when AI speeds up the voluntary deep-dive phase.

Anthropic runs probably the strictest verification in the industry. A modular supply chain built on proxies cuts straight through it. Chinese developers can buy Claude tokens at roughly ten percent of the official price. The method is proxies hosted outside China that forward requests as if from a legitimate location. Users pay in yuan through local payment apps; no VPN or foreign credit card is needed. Operators push prices down by exploiting free credits and secretly swapping expensive models for cheaper ones. The analysis warns this weakens not only geoblocking but the ability to monitor misuse.

According to one analyst, February 6, 2026 may have been the last day humans consumed more tokens than agents. Agent usage has grown 14x since. Agentic token usage has grown 14x since then; human usage is up just 2.8x. Nearly 70 percent of agent token usage comes from cached prompts, billed at much lower rates. Actual costs are therefore not rising as fast as the raw numbers suggest. Token inflation had already started with reasoning models that think longer before responding.

Qwen3.8-Flash-Next has 125 billion parameters but activates only 6 billion per token. It outperforms a model three times its active size. Alibaba says it beats its predecessor at roughly one-ninth the training cost. A novel n-gram embedding layer holds 51 billion parameters and can sit in system RAM rather than on the GPU. On agentic coding benchmarks it beats models far larger than itself. The production version is priced at $0.16 per million input tokens.