Optima lets you benchmark models against your own data
Artificial Analysis has launched Optima, a platform letting users build custom benchmarks from their own data and workflows, comparing models on cost and time as well as quality.
Artificial Analysis has launched Optima, a platform letting users build custom benchmarks from their own data and workflows, comparing models on cost and time as well as quality.
Dynatrace is acquiring Arize, which specialises in AI observability and the AI development lifecycle, for $915 million, including around $815 million in cash.
A Bloomberg analysis finds Chinese citizens are more optimistic about AI than Americans, because it is seen there as a practical tool and disrupts a smaller share of the population.
According to the Wall Street Journal, Nvidia has reworked the deal financing OpenAI's Ohio data centre campus and will initially guarantee only half of the planned $250 billion backstop.
The California Public Utilities Commission has approved Waymo's robotaxi service expanding across 18 counties from Sonoma down to San Diego — its largest expansion yet.
A study involving Google researchers shows that training chatbots not to claim consciousness also changes their stance on animal rights, religion and life satisfaction.
According to the Washington Post, members of Congress and their aides use AI tools with little oversight for everything from writing speeches to sorting constituent mail.
OpenAI's Computer History records clicks, keystrokes and app switches on Mac and turns them into a searchable timeline for ChatGPT and Codex.
Anthropic chief executive Dario Amodei set out his view on AI regulation, arguing open weights are not sufficient on their own and pointing to a crisis of trust in the industry.
DeepSeek's leaderboard-topping V4 Flash completed only 53.8% of complex agent tasks in testing. Thirty hard tasks were run across eight different agent harnesses.
Fields Medallist Timothy Gowers says almost all the famous mathematics problems solved by language models so far were solved with counterexamples rather than proofs.
A representative Epoch AI survey finds 20% of employed Americans hand at least one task previously done by a person to AI, usually accepting the output with little or no editing.