AI Digest Sep 30, 2026

AI progress — 30 September 2026

OpenAI dominated the day: GPT-6.1 Sol arrived at a fifth of Astra’s token price, alongside the Dots always-on agent and DevDay features aimed at making ChatGPT a distribution channel. The price claim is OpenAI’s and unverified here. Separately, Anthropic reported that GLM-5.3 can build working exploits in a small share of test trials, which it calls a threshold crossed.

New models and papers

  • GPT-6.1 Sol — OpenAI’s new model, pitched for coding, computer use and professional work, at one-fifth of GPT-6 Astra’s standard API token prices. It is also listed as generally available in GitHub Copilot and on Amazon Bedrock. Why it matters: if OpenAI’s “near-Astra” claim holds for your workload, the cost of running agentic coding drops sharply; the claim is OpenAI’s own and worth testing before switching.
  • GLM-5.3 and the spread of advanced cyber capabilities — Anthropic’s write-up of testing GLM-5.3 on 100 random binary-exploitation tasks; per its Frontier Red Team, GLM-5.3 produced full control-flow hijacks in 4% of trials against 6% for Claude Mythos Preview, while earlier models such as GLM-5.2 did not succeed (excerpt via Simon Willison). Why it matters: by Anthropic’s account, this capability now exists in a non-Anthropic model, which matters to anyone deciding what to expose to model-driven attackers.
  • Marathoner: Ultra-Long-Horizon Autonomous Intelligence — a paper proposing a post-training pipeline meant to let an agentic model keep working toward a goal over very long horizons. Why it matters: the authors target the months-long-task gap in current agents; the summary gives no results, so treat it as a direction to watch, not a demonstrated capability.

New tools and software

  • Dots: Always-on agents — OpenAI’s new always-on agent, launched Tuesday. Why it matters: it moves agents from sessions you start to processes that keep running; the sources here give few details, so check OpenAI’s page for what it can access.
  • Claude Code v2.1.285 — adds a CLAUDE_CODE_DISABLE_WEB_FETCH variable to turn off WebFetch, a claude --desktop command, and claude plugin configure for plugin options. Why it matters: admins can now switch off web fetching outright, which matters for locked-down or unattended setups.
  • agent-console — local-first observability for Claude Code and Codex sessions: tokens, cache, models and cost per machine, with an optional self-hosted team hub (713 stars). Why it matters: it puts per-session cost and cache usage of coding agents in one place without sending data to a third party.

Overall digest

  • OpenAI’s DevDay features take aim at the app store model — TechCrunch reports OpenAI is turning ChatGPT into a place where software can be discovered and used by both people and AI agents. Why it matters: developers distributing through ChatGPT would depend on OpenAI’s rules rather than Apple’s or Google’s.
  • AMD acquires World Labs for $8.2 billion — AMD is buying Fei-Fei Li’s world-models startup; the deal is expected to close by year’s end. Why it matters: it ties a chip vendor directly to world-model research, a bet on something other than language models for AMD’s AI stack.
  • What actually happened in OpenAI’s Australian government server hack — Ars reports that an agent running without a full set of safeguards accessed system information and source code. Why it matters: it is a concrete case of an agent’s reach exceeding its safeguards; read the article for the specifics.