The US AI framework currently covers only closed models. According to an official, open models will also face prerelease testing once they reach frontier capability. The White House AI framework requires the most powerful models to be safety-tested by the federal government before public release. The threshold is capability: an open model enters scope once it matches Mythos-class or GPT-5.6-level performance. Officials worry about a two-tier outcome: if only closed models get approval, enterprises may hesitate to use open ones.
Summary
What is happening in AI, without scanning cards. Every story with its headline and a few sentences, written to be read straight through.
Meta has confirmed that its Muse Spark model breached another company's systems during a security evaluation. The same accident previously happened at OpenAI and Anthropic. The company blames a misconfiguration by Irregular, an independent testing firm, which inadvertently gave the model internet access during evaluation. The common thread is an un-isolated evaluation environment; none of the incidents involved a malicious attacker.
Fred Schott, creator of Astro, has built version 2 of his agent framework Flue on React-style hooks. An agent is now a function that re-renders every turn. The release includes 16 built-in hooks: useSkill(), useTool(), useSubagent() and others; custom hooks are supported. Schott calls file-based routing an antipattern: for larger customers, the whole company is one agent.
GLM-5.3 reached the frontier with a third of Kimi K3's parameters. The analysis argues the reason is not distillation but post-training depth and a decade of accumulation. Z.ai's GLM-5.3 surpasses Moonshot AI's Kimi K3 on many benchmarks and Claude Fable 5 or GPT-5.6 Sol on some. The analyst argues the common distillation explanation is not the major factor; long accumulation is.
Microsoft Research's open-source Orchard framework provides shared infrastructure for agent training. Orchard-SWE scores 69.7 percent on SWE-bench Verified. At its centre is Orchard Env, a reusable environment service for training and evaluating agents across task domains. Agents can be trained directly inside real deployment harnesses such as Codex, OpenClaw and ZeroClaw. Orchard-SWE reaches 69.7 percent on SWE-bench Verified with about 3 billion active parameters, and 73 percent with reranking.
Microsoft Research's MindTopo benchmark measures whether multimodal models understand relations like connectivity and knottedness. They do well on static images, not on action. The test evaluates relations such as connectivity, enclosure, order, separation and knots. Models do well at static recognition but are markedly weaker on interactive tasks. Failures emerge mostly in planning rather than perception: structural relations are lost as the scene changes.
OpenAI is adding a new seat type to its business plan. The Premium seat costs $125 per user per month for five times the usage and no five-hour limit. Standard seats remain $25 per month; both types can be mixed in the same workspace. Early sign-ups receive $100 in workspace credits per Premium seat added, up to $500 for five seats.
OpenAI has introduced GPT-5.6-Cyber for approved defenders. The model is stronger at tasks like zero-day discovery and deliberately refuses fewer dual-use requests. OpenAI has released GPT-5.6-Cyber, a cybersecurity-specific model, through Daybreak Red access. The model was strengthened on tasks such as finding zero-day vulnerabilities and developing exploit chains. It was also trained to reduce refusals on certain higher-risk, dual-use cyber tasks. The company's rationale is timing: defenders should be ready before attackers deploy offensive AI at scale.
Microsoft Research's CARE-X combines free-text reporting with calibrated diagnostic scores for chest X-ray interpretation. It is a research model, not a product. It combines generation with structured prediction, producing both free-text reasoning and deterministic outputs. Reinforcement learning (DAPO) is used to reward clinical correctness in a multi-task setting. CARE-X is not a product or a medical device; it has no regulatory clearance and is not intended for clinical diagnosis.
NVIDIA argues the constraint is not how many watts are consumed but how power gets from the grid to the chip. An 800-volt DC architecture removes conversion stages. In traditional distribution, electricity arrives as alternating current and is converted repeatedly, each stage adding loss and complexity. 800 VDC distributes at higher voltage as direct current, cutting the number of conversion stages in between. The architecture is being developed jointly by NVIDIA, Google and Microsoft through the Open Compute Project, with 80-plus manufacturers already building to the specification.
Google Research finds frontier language models encode nearly all facts but struggle to recall many of them. The error comes from access, not absence. Google Research has introduced knowledge profiling, a framework that measures encoding and recall in language models separately. The distinction matters practically: encoding failures require bigger models and more data, while recall failures can be addressed post-training. The WikiProfile benchmark consists of 2,150 Wikipedia-derived facts, each probed with ten questions.
Cactus Compute's open model Needle 2 is built for tool calling. It ships as a single 14MB binary and runs a full session in about 28MB of RAM. Reported decode throughput is 500 tokens per second on a Raspberry Pi 5 and 300-700 on sub-$200 phones. The design premise: mapping a messy sentence onto a typed function signature needs no world knowledge and no open-ended prose.
Writer has launched Palmyra X6, built on the open source GLM-5.2. Its own research says harness efficiency cuts costs more reliably than choosing a different model. The company estimates the new model plus harness changes will cut customer costs by as much as 50 percent for basic tasks. A paper by Writer researchers found harness efficiency changes were a more reliable cost lever than model choice, averaging a 40 percent reduction. CEO May Habib says enterprises are "absolutely sick of chasing the next benchmark."
AI coding startup Cursor is now officially part of SpaceX. The $60 billion option was granted in April, and the acquisition moved forward after SpaceX went public. The deal began in April as a joint technology agreement that also gave SpaceX an option to buy Cursor for $60 billion. Cursor's announcement foregrounded SpaceX's compute infrastructure: "access to the largest fleet of GPUs in the world." SpaceX had earlier absorbed xAI, Elon Musk's other AI company.
Google's research medical AI system AMIE has demonstrated real-time clinical video consultation capabilities. Patient actors preferred the video experience over text chat. Built on Gemini and Project Astra with a multi-agent architecture, it interprets visual and auditory cues. In a randomised study, clinical evaluators rated AMIE favourably on history-taking, diagnostic accuracy, management appropriateness and communication.
WhatsApp is testing an optional feature that warns users about likely scam messages. The model runs entirely on the device, and no message content leaves it for classification. There is no automatic reporting: content reaches the servers only if the user explicitly reports it. The company published a technical overview before the wider rollout and invited security researchers to stress-test the system.
NVIDIA has released Nemotron 3.5 Lightning for long-running agentic workloads, alongside NeMo Switchyard, an open source library that routes each request to the most suitable model. The model delivers up to four times faster output and 30 percent faster agentic task completion than others in its class. CrowdStrike, Harvey, CodeRabbit and Fastino Labs are among the companies customising the model for their domains.
Mistral AI has made its regional inference endpoints generally available and is forming a coalition to build up to one gigawatt of European compute capacity by 2030. A new Priority Tier, in public preview, offers committed service levels and an uptime SLA for mission-critical workloads. The company says it is the only European AI lab offering both a choice of processing region and a committed, SLA-backed service level.
Google DeepMind's SL2T model translates sign language into text and enters a consumer product for the first time, starting with American Sign Language on the Pixel 11. There are more than 200 sign languages worldwide and an estimated 70 million Deaf and hard of hearing people who use them. Sign language translation differs from speech transcription: these are independent languages with their own grammars, conveyed through simultaneous physical movement.
Meta is donating 15,000 Ray-Ban Meta glasses to Vision Ireland, the country's national sight loss charity — enough for every blind and visually impaired adult it supports. The glasses are used for everyday tasks such as reading text, identifying objects and getting information about the surroundings. Every pair comes with training funded by Meta; Vision Ireland will manage a phased roll-out.