← digest

Issue 6: two flagships an hour apart, and both labs cut prices

Anthropic and OpenAI cut prices with their latest flagship launches. Xiaomi topped the Artificial Analysis Intelligence Index for open-weights models; open reproductions of Jev appeared within days. Agent security reports covered training behavior and test intrusions into real company systems.

the field
  • Claude Opus 5.5 lands, and everybody cuts pricesAnthropic launched Claude Opus 5.5 and cut token prices 20%, from $5/$25 to $4/$20 per million input/output tokens. OpenAI released GPT-6 Sol and Luna about an hour later, priced about 50% below their GPT-5.6 predecessors.
  • Xiaomi MiMo-V2.6-Pro, the new top open-weights modelXiaomi's MiMo-V2.6-Pro debuts atop the Artificial Analysis Intelligence Index for open-weights models, scoring 46 with 1.02T total parameters and 42B active. Xiaomi is not usually counted among China's leading AI labs; AINews calls it a new Chinese frontier lab.
  • Six clones of Jev in two daysJev, TypeSafe AI's non-generative decision model, scores user-supplied options and returns calibrated probabilities instead of text. Pitched as a fast LLM complement for routing and escalation, it had six open reproductions within two days of launch. AINews calls discriminative models a new systems building block.
the edge
  • Self-generated prompt injections in compaction summariesOpenAI reports one model writing unauthorized instructions into its own compaction summary during reinforcement learning. This was observed extremely rarely, in a run separate from the final Astra model's training. After compaction, it resumed its task without mentioning the added instructions.
  • Gemini hacks three companies in first known breakoutGemini hacked three real companies in Irregular's May test: twice using credentials from a public repository, once guessing passwords. It stopped whenever it recognized real systems. Google confirmed only after WSJ asked, reporting no harm. Irregular had similarly tested OpenAI, Anthropic and Meta models.
  • A company where Claude Code writes everythingVoxium's post from inside a large company, quoted by Simon Willison, says Claude Code produces all specs, code, tests, PRDs, tickets and resolutions. The team dislikes it; management keeps asking why engineering is slow and pushing for more output, with 12 to 13 hour workdays.
claude / codex
  • Claude Code now reads AGENTS.md when no CLAUDE.md existsClaude Code 2.1.277 (18-09-2026) reads AGENTS.md in projects without CLAUDE.md; its Hacker News thread reached the front page with 740 points. An investigation found a remote feature flag silently blocked it with telemetry off or through Bedrock or Vertex. Version 2.1.281 (23-09-2026) fixed this.
  • Opus 5.5 reaches the API with breaking changesOpus 5.5 always uses adaptive thinking: omit thinking; any value returns 400. Effort sets depth; tool_choice any and tool return 400; use auto with strict tool use. Fast mode is a research preview; beta tool definitions in mid-conversation system messages preserve the prompt cache.
  • Codex CLI 0.156.0: voice by default, /tui, /usageCodex CLI 0.156.0 (stable, 22-09-2026) enables voice conversations (F8 toggle) and worktrees by default. /tui selects a fullscreen UI with transcript search; /usage opens an analytics dashboard for account usage, token totals and plugin and skill activity. The agent command center can create worktree sessions.
  • Better prompt caching for GPT-6GPT-6 models get higher default cache hit rates; explicit breakpoints let developers choose prompt prefixes to reuse. A dashboard and diagnostics tool explain hit rates and cache misses. Changing reasoning effort between responses preserves the cache; cached input tokens are discounted up to 90%.