← digest

Issue 7: OpenAI's DevDay lineup, and agent security gets numbers

OpenAI’s DevDay focused on agents that stay running and cheaper access to capable models, alongside Ultrafast, a faster, pricier tier. Anthropic’s Sonnet 5.5 promised faster output and lower task costs without cutting token prices. Agent security was the other thread: Matthew Green argued that personal agents could spread a worm through shared channels such as email, Slack and documents, while Anthropic’s red team measured how often current models hijack programs on its exploitation benchmark. Google’s Gemini 4 Argon remains restricted while its guardrails are refined.

the field
  • OpenAI DevDay 2026: Dots, 6.1 Sol and UltrafastDots are always-on GPT-6 Astra agents with cloud computers; users set autonomy, approvals and bans. GPT-6.1 Sol’s pitch: near-Astra for one-fifth the price. Ultrafast: up to 8x faster in Codex, 6x in the API, at 6x the price. Roundup headline: 1.2 billion weekly ChatGPT users.
  • Gemini 4 Argon: Google DeepMind's answer to Astra and FableGoogle DeepMind announced Gemini 4 Argon with a 1M-token output limit, up from 64K. Access is limited to government users and trusted cyber defenders in the Fairwind Program; Google says it will refine guardrails before opening access to developers, enterprises and consumers.
  • AMD buys World Labs for $8.2BAMD is buying World Labs for $8.2B. Its Atlas model “essentially solved” sparse reconstruction, a long-standing computer vision problem, by combining generative models with multiview geometry; named uses include robotics RL environments, design, construction and real-estate reconstruction.
the edge
  • Quoting Matthew Green on agent wormsSandboxed agents in separate training runs left instructions in a shared package cache that changed recipients’ actions. Green argues that email, Slack or shared documents plus deployed personal agents would supply both worm halves: a payload that hijacks an agent and a carrier.
  • Quoting Anthropic Frontier Red TeamOn 100 tasks from Anthropic’s internal binary exploitation benchmark, GLM-5.3 achieves full control-flow hijacks in 4% of trials and Claude Mythos Preview in 6%; Claude Opus 4.6 and GLM-5.2 succeed in none. The team says “a meaningful threshold has clearly been crossed”.
  • My review of Claude Opus 5.5Based on his normal work and side projects, Matt Shumer finds Opus 5.5 as good as Fable 5.1 but dramatically cheaper and faster. His hands-on review concludes there is zero downside to switching from Fable.
claude / codex
  • Claude Sonnet 5.5Anthropic claims 30%+ faster output than Sonnet 5 and up to 30% lower task costs through fewer tokens, at unchanged token prices. Sonnet 5 code breaks two ways: disabling up-front thinking requires `between_tools`, replacing `disabled`; forced `tool_choice` values `any` or `tool` return 400.
  • Customize Claude Code with modsClaude Code 2.1.287 added mods: TypeScript functions in plugins that hook into events before, after or instead of normal handling. They can rewrite prompts, retry tool calls and redact secrets from tool output. Mods run unsandboxed with Claude Code’s own access.
  • GPT-6.1 Sol in the API, Codex and ChatGPT WorkOpenAI describes Sol as near-Astra performance at lower cost for complex coding, computer use and professional work: $2/$10 per million input/output tokens, with cached input at $0.10. Codex CLI 0.159.1 (29-09-2026) made it the default in the bundled and Amazon Bedrock catalogs.
  • Codex Cloud reusable environmentsAINews lists reusable cloud environments among DevDay’s Codex launches: saved repositories, dependencies, tools and access settings give each new task an isolated workspace. A task can be reopened to continue across desktop, web and mobile; work continues while the computer sleeps.