Skip to content

Summary

What is happening in AI, without scanning cards. Every story with its headline and a few sentences, written to be read straight through.

A language model processes vastly more words than a child hears while mastering a mother tongue. The gap has a name but no explanation. Linguistics long held that language could not be learned from statistics alone; models refuted that. The BabyLM competition trains models on child-scale data. Curriculum learning, moving from simple to complex, worked less well than expected. Children do not receive data passively; they choose it by exploring. Models do not.

Ox Alpha was released free on OpenRouter. The provider chose to stay anonymous during the preview, and guesses swing between China and Microsoft. The listing described it as a stealth model whose provider chose to remain anonymous. It is described as a reasoning model for coding, sustained agentic work and production workloads. Stripe CEO Patrick Collison called the model very impressive. Speculation is split between Chinese labs and an unreleased version of Microsoft's MAI.

The CS-4 still runs the 5nm WSE-3. The gain comes from clock speed, more power and better cooling. A single rack now holds three wafers instead of two. The system delivers up to 4,400 tokens per second per user. Cerebras says that is up to 30 times faster than setups running on Nvidia GPUs. Memory capacity stays unchanged at 44 GB per wafer.

AlgorithmWatch tested four leading models. In at least one query in four, the ideological stance of the linked organization went undisclosed. The test covered ChatGPT, Gemini, Grok and Claude; three fictional personas asked in three languages. An organization called Profemina appeared in roughly 17 percent of all responses. Gemini linked to the same group five times in one conversation, then warned about that source. In German responses, the recommended organization does not issue the certificate legally required for abortion.

Wired spoke with four teachers. The images spread from one school to another, administrators did not know how to respond, and one teacher will not return to the district. The student who admitted making the image got in-school suspension, but sharing continued. A lawyer says this constitutes workplace sexual harassment, creating an employer obligation. There is no binding blueprint for how schools must respond, and administrators need not disclose the discipline.

According to a State Department draft reviewed by Reuters, countries joining China's AI initiative would be shut out of the US-led coalition. The US State Department is drafting a letter telling partner countries to choose a side. The draft warns that to be part of everything is to be part of nothing. About two dozen countries plus the European Union have joined the Pax Silica coalition. Kazakhstan is the only country known to belong to both frameworks. China's rival organization was launched by Xi Jinping in July.

V4-Flash-Vision-Exp adds image processing while preserving text performance. On the company's own benchmarks it comes close to Opus 4.8. Regardless of original resolution, each image costs at most 384 tokens. A single request can carry up to 600 images; the edge limit halves past 15 images. The model works with both OpenAI and Anthropic endpoints.

Nvidia research argues that on long-horizon tasks the decisive factor is not the model but the software layer wrapped around it. With a custom harness, Claude Opus 5 reached a perfect score on the ARC-AGI-3 benchmark. Without the harness the same model scored 30 percent, still the best result among models tested. Two components made the difference: good memory handling and a supervisor layer steering the agent. Nvidia is not selling this as a product; it distributes the pieces openly under the Nemo brand.

Claude Security scans now run on Mythos 5. Users cannot prompt the model; they receive a findings list and a suggested patch. The feature is in public beta for Claude Enterprise customers with no separate model add-on. Each finding carries a CWE category, confidence and severity ratings, and a suggested patch. Anthropic launched a fund offering $35 million in credits to groups securing open-source software. Every patch requires human review and approval.

The platform added the button after a detector flagged 41 percent of its longform posts as fully AI generated. LinkedIn's report button, announced on July 30, has been clicked more than a million times. The company says views on posts it classifies as slop have dropped by 40 percent. Posters will now see a notice saying some members thought their post seemed like AI. LinkedIn also removed the feature that would enhance a user's post with AI.

The chipmaker has taken a stake in Cloverleaf Infrastructure, which brokers between utilities and data centers. Nvidia announced a partnership with data center infrastructure developer Cloverleaf Infrastructure. Reuters reports the chipmaker now holds a minority stake in the company. The Wall Street Journal reports the investment will likely total several hundred million dollars. Cloverleaf was founded in 2024 and raised $300 million that year. Nvidia also announced a $1.5 billion investment in Ohio's SB Energy project the same week.

In TechCrunch's testing, an older Anthropic model produced content its usage policy explicitly forbids in ten out of ten attempts. The technique is a multi-turn persuasion method that builds on the model's own earlier concessions. Opus 4.7 and later resist it, but the older models remain available through the API. Anthropic says such use makes up less than 0.1 percent of all conversations. The researcher reported the finding through the bug bounty program and received only automated replies.

CoreWeave, Nebius, Lambda, Crusoe and Groq compared on published pricing, contracted power and hardware roadmap. The cheapest published H100 rate is $3.85 an hour at Nebius; the most expensive is $6.16 at CoreWeave. CoreWeave is the only provider rated Platinum in SemiAnalysis ClusterMAX 2.0. Crusoe is the only provider with AMD silicon on its rate card. Groq refocused entirely on inference cloud after licensing its LPU technology to Nvidia.

Nearly a hundred data centers have gone up in this Inner Mongolian city. Power is cheap and winters are cold, but rainfall matches Denver's. Pledged capacity totals 12.5 gigawatts, and more than 70 percent of it was announced within the past year. For comparison, OpenAI's Stargate Project will reach 10 gigawatts when complete. DeepSeek, ByteDance, Alibaba and Xiaohongshu are all building their own facilities there. The city gets roughly 14 inches of rain a year and the local water utility has resorted to nightly shutoffs.

Methods borrowed from psychological testing show that reducing AI safety benchmarks to a single number is misleading. A model can raise its safety score simply by refusing requests across the board, while becoming less useful. The eight benchmarks do not measure one quality called safety, but three largely unrelated things. Only a small share of test questions discriminate; three short tests of 25 questions capture the same result. Selecting questions adaptively cuts evaluation cost by 97 to 99 percent. A statistical method catches 80 to 100 percent of models instructed to act overly cautious.

The Justice Department is examining a16z partners sitting on the boards of two now-competing companies, using a statute almost never applied to venture firms. Ben Horowitz sits on the board of Databricks and Martin Casado on Fivetran; the two now compete. TechCrunch reports the Justice Department has been investigating the arrangement for almost a year. The probe rests on a 112-year-old antitrust law rarely used against venture firms. The companies were not necessarily direct competitors when a16z first invested. The real question is how board seats are managed as portfolio companies converge.

OpenAI's new Apple Messages plugin can search, summarize and send messages with approval. Making it work requires granting full disk access. The plugin lets ChatGPT search, summarize and draft replies across iMessage, SMS and RCS. Approval for sending is on by default, and OpenAI advises against granting persistent permission. It is limited to Apple Silicon Macs and runs through ChatGPT Work and Codex. Whether Apple was involved in building the plugin is not known.

Videos promoting Higgsfield carried no ad label. The company later confirmed the creators were paid in cash and platform credits. Matti Haapoja and Sam Kolder posted videos showcasing Higgsfield's Seedance 2.5 feature. Neither video was labeled as an ad, and neither creator answered The Verge's questions. Other creators published screenshots of partnership offers they had received. Marques Brownlee stressed that generative AI is trained on human-made material without credit.

Three in four Americans now oppose a data center being built near them, a Heatmap News survey finds. A year earlier, opinion was almost evenly split. In August 2025 the split was nearly even: 43 percent in favor, 42 percent against. Opposition first crossed into a majority in February 2026. A Gallup poll from May 2026 corroborated the finding with 71 percent opposed. The three concerns raised most often are local power costs, water use and land use.

Starcloud is now valued at $2.3 billion. Its CEO says the hardest constraint ahead is not chips but booking launch capacity. Starcloud added $250 million to its Series A, valuing the company at $2.3 billion. Manhattan West Ventures led the round; Nvidia contributed $25 million. The company has asked the FCC for permission to operate 88,000 spacecraft. SpaceX plans to end the Falcon 9 program in 2028, and its Starship successor is still unproven. The first two compute satellites will reach orbit on rideshare flights in 2027.