Skip to content

Summary

What is happening in AI, without scanning cards. Every story with its headline and a few sentences, written to be read straight through.

IBM is bringing OpenAI's models into its global consulting business. The deal covers training tens of thousands of consultants and building sector-specific solutions for finance, government, telecoms and retail. IBM announced a partnership to bring OpenAI's models and tools to enterprise customers; the financial terms were not disclosed. A dedicated OpenAI practice is being set up inside IBM Consulting, with tens of thousands of consultants to be trained and certified over the coming months. Training will focus on Codex, the API, cybersecurity and consultative solution credentials, alongside a group of 'Forward Deployed Experts'. GPT-5.6, Codex and ChatGPT Work will be integrated into IBM Consulting Advantage, the platform IBM's consultants use. The deal follows a similar alliance IBM announced with Anthropic less than a year ago.

AI research company Pathway, which develops models based on what it calls a Post-Transformer BDH architecture, has closed a $30 million seed round at a $500 million valuation. A seed valuation at this level is effectively priced on an architectural claim, not a product. The architecture's claim has not been independently verified.

According to The Information, firms gathering data for AI labs are driving demand to buy or license the internal datasets of startups that are shutting down or being acquired. In one case such an offer arrived eight days after an acquisition agreement. Datasets are becoming an unexpected asset for a company that is winding down.

Artificial Analysis has launched Optima, a platform letting users build custom benchmarks from their own data and workflows, comparing models on cost and time as well as quality. For agent-based applications those metrics can be more informative than raw token pricing. The platform targets the problem of general benchmarks not matching your own work.

Dynatrace is acquiring Arize, which specialises in AI observability and the AI development lifecycle, for $915 million, including around $815 million in cash. Arize builds tools that track how models behave in production. The deal shows the monitoring software market extending into AI systems.

A Bloomberg analysis finds Chinese citizens are more optimistic about AI than Americans, because it is seen there as a practical tool and disrupts a smaller share of the population. The explanation offered is that the technology is seen in China as a practical tool. The argument is that past experience with technology shapes what is expected of it in future.

According to the Wall Street Journal, Nvidia has reworked the deal financing OpenAI's Ohio data centre campus and will initially guarantee only half of the planned $250 billion backstop. The project is 10 gigawatts; the scale is closer to energy infrastructure than to a single data centre. The information rests on Wall Street Journal sources; neither party has confirmed it.

The California Public Utilities Commission has approved Waymo's robotaxi service expanding across 18 counties from Sonoma down to San Diego — its largest expansion yet. This is the largest expansion approval the company has received to date. The approval indicates the confidence driverless vehicles have gained on the regulatory side.

According to the Washington Post, members of Congress and their aides use AI tools with little oversight for everything from writing speeches to sorting constituent mail. The use is largely unsupervised: there is no shared rule or disclosure requirement. Some of the tasks involved feed directly into the legislative process.

OpenAI's Computer History records clicks, keystrokes and app switches on Mac and turns them into a searchable timeline for ChatGPT and Codex. The data is stored locally as unencrypted Markdown files. OpenAI says the data is not used for training; memories that feed into chats may still end up as training data.

Anthropic chief executive Dario Amodei set out his view on AI regulation, arguing open weights are not sufficient on their own and pointing to a crisis of trust in the industry. He defends the necessity of model testing and points to a "crisis of trust" in the industry. The argument turns on whether regulation concentrates power in a few companies or distributes it. The remarks came within a social media exchange rather than a formal policy document.

DeepSeek's leaderboard-topping V4 Flash completed only 53.8% of complex agent tasks in testing. Thirty hard tasks were run across eight different agent harnesses. Composio ran the model through eight agent harnesses, including Claude Code and Codex, on 30 hard tasks using live tools such as Gmail and GitHub. Of 240 total runs 129 passed; only six of the 30 workflows completed across every harness. The gap shows the outcome depends less on the model than on the harness around it.

Fields Medallist Timothy Gowers says almost all the famous mathematics problems solved by language models so far were solved with counterexamples rather than proofs. Finding a counterexample means producing a single instance showing a claim is false; a proof shows the claim holds in every case. The two are cognitively different: one is search, the other is construction. The distinction shows where models are actually contributing in mathematics.

A representative Epoch AI survey finds 20% of employed Americans hand at least one task previously done by a person to AI, usually accepting the output with little or no editing. The tasks in question were previously done by a person, making this a substitution rather than an addition. The finding raises a question less about measuring productivity than about where the review step went.