Epoka

Summary

What is happening in AI, without scanning cards. Every story with its headline and a few sentences, written to be read straight through.

According to the Washington Post, members of Congress and their aides use AI tools with little oversight for everything from writing speeches to sorting constituent mail. The use is largely unsupervised: there is no shared rule or disclosure requirement. Some of the tasks involved feed directly into the legislative process.

OpenAI's Computer History records clicks, keystrokes and app switches on Mac and turns them into a searchable timeline for ChatGPT and Codex. The data is stored locally as unencrypted Markdown files. OpenAI says the data is not used for training; memories that feed into chats may still end up as training data.

Anthropic chief executive Dario Amodei set out his view on AI regulation, arguing open weights are not sufficient on their own and pointing to a crisis of trust in the industry. He defends the necessity of model testing and points to a "crisis of trust" in the industry. The argument turns on whether regulation concentrates power in a few companies or distributes it. The remarks came within a social media exchange rather than a formal policy document.

DeepSeek's leaderboard-topping V4 Flash completed only 53.8% of complex agent tasks in testing. Thirty hard tasks were run across eight different agent harnesses. Composio ran the model through eight agent harnesses, including Claude Code and Codex, on 30 hard tasks using live tools such as Gmail and GitHub. Of 240 total runs 129 passed; only six of the 30 workflows completed across every harness. The gap shows the outcome depends less on the model than on the harness around it.

Fields Medallist Timothy Gowers says almost all the famous mathematics problems solved by language models so far were solved with counterexamples rather than proofs. Finding a counterexample means producing a single instance showing a claim is false; a proof shows the claim holds in every case. The two are cognitively different: one is search, the other is construction. The distinction shows where models are actually contributing in mathematics.

A representative Epoch AI survey finds 20% of employed Americans hand at least one task previously done by a person to AI, usually accepting the output with little or no editing. The tasks in question were previously done by a person, making this a substitution rather than an addition. The finding raises a question less about measuring productivity than about where the review step went.

Anthropic's safety report says the internal system filtering biological and chemical weapons risks was inactive for nearly a year, letting 133 million interactions through unfiltered. During that period roughly 50,000 external feedback contractors ran about 133 million unfiltered interactions with the models. The information comes from the company's own safety report, not from a leak or an outside audit. The case shows the difference between a safeguard existing and its operation being monitored.

Apple is reported to have developed a large language model specific to the Chinese market together with Alibaba, and has registered the service with China's regulator. The information rests on three unnamed sources close to the matter, reported by Reuters. Apple previously used domestic companies' models in China; developing its own marks a change of strategy. The company has registered the service with China's cyberspace regulator; approval would make it the first US company granted that permission.

Publisher and internet pioneer Tim O'Reilly argues that the big AI labs are building an architecture that locks users in, and that real innovation will come from open source. Tim O'Reilly says the big AI labs are trying to lock users in the way Microsoft did in the 1990s. In his view open-source AI means more than publishing weights: it requires a clean separation between model, harness and application. He notes that the largest cybersecurity incidents have come from frontier models.

Energy research firm Noreva forecasts that natural gas prices could triple in parts of the US. Four large companies are planning gigawatt-scale gas plants. Amazon, Google, Meta and Microsoft are planning gigawatt-scale gas plants for their AI data centres. Prices currently between $2 and $4.50 could exceed $10 in some regions, according to the forecast. Because fuel accounts for around half the cost of electricity from a large plant, a rise would feed straight into token prices.

A self-represented plaintiff in Connecticut hid invisible AI instructions in his filings as 3-point white text on white. The judge noticed the unusual whitespace. The hidden text instructed any AI review system to produce output favouring his filing. Judge Walter Spader Jr. spotted the manipulation from unusual whitespace and revoked the plaintiff's electronic filing privileges.

An eval harness surfaced a pattern qualitative review had missed: AI models display their highest confidence precisely when their answers are wrong. The same pattern went unnoticed in qualitative review because fluent, coherent wrong answers read as correct. Most teams skip the verification step: it is tedious, time-consuming and produces nothing visible to end users. The gap between "this sounds right" and "this is verifiably correct" is where enterprise tools fail quietly.

AI coding startup Cognition is reportedly in talks for a new round at a $40 billion valuation, only months after its previous raise. That amounts to a valuation increase of roughly 54% in the space of months. AI coding tools are among the fastest areas of enterprise adoption.

Apple is reportedly talking to publishers about paying for current news to feed Siri. According to the WSJ, the company has considered a nine-figure budget. The move continues the industry's shift from using news content without permission to licensing it. The deal aims to close a long-standing weakness in Siri's handling of questions that need current information.

Google is adding Gemini-based AI summary and reporting tools to its Ads and Analytics platforms. The new tools aim to speed up performance insights and enable benchmarking. Google announced on August 10, 2026 that it added new AI tools to Google Ads and Google Analytics. AI Overviews added to the Analytics homepage show key performance changes since the last login. The new Dashboards feature in Google Ads enables visual report creation via text prompts. A new benchmarking feature within Ask Advisor lets users compare their performance with similar businesses. All the new features are built on the Gemini model.

A leaked Accenture meeting recording shows token consumption is driven by non-engineers, not engineers. One of the biggest cost items is converting PDFs into text. AI costs are shifting from a budget line into an operational problem inside organisations. The source of the problem is not model pricing but internal document formats and working habits.

Nvidia has a plan to keep its GPUs from losing value. The goal is to convince a new class of financiers to keep lending against AI buildouts. The $500 billion scale makes this a financing question beyond any single company's balance sheet. TechCrunch's assessment calls the plan risky but brilliant, particularly for ageing hardware.