A 3-billion-parameter agent closes in on models ten times its size
Microsoft Research's open-source Orchard framework provides shared infrastructure for agent training. Orchard-SWE scores 69.7 percent on SWE-bench Verified.
Microsoft Research's open-source Orchard framework provides shared infrastructure for agent training. Orchard-SWE scores 69.7 percent on SWE-bench Verified.
Microsoft Research's MindTopo benchmark measures whether multimodal models understand relations like connectivity and knottedness. They do well on static images, not on action.
OpenAI is adding a new seat type to its business plan. The Premium seat costs $125 per user per month for five times the usage and no five-hour limit.
OpenAI has introduced GPT-5.6-Cyber for approved defenders. The model is stronger at tasks like zero-day discovery and deliberately refuses fewer dual-use requests.
Microsoft Research's CARE-X combines free-text reporting with calibrated diagnostic scores for chest X-ray interpretation. It is a research model, not a product.
NVIDIA argues the constraint is not how many watts are consumed but how power gets from the grid to the chip. An 800-volt DC architecture removes conversion stages.
Google Research finds frontier language models encode nearly all facts but struggle to recall many of them. The error comes from access, not absence.
Cactus Compute's open model Needle 2 is built for tool calling. It ships as a single 14MB binary and runs a full session in about 28MB of RAM.
Writer has launched Palmyra X6, built on the open source GLM-5.2. Its own research says harness efficiency cuts costs more reliably than choosing a different model.
AI coding startup Cursor is now officially part of SpaceX. The $60 billion option was granted in April, and the acquisition moved forward after SpaceX went public.
Google's research medical AI system AMIE has demonstrated real-time clinical video consultation capabilities. Patient actors preferred the video experience over text chat.
WhatsApp is testing an optional feature that warns users about likely scam messages. The model runs entirely on the device, and no message content leaves it for classification.