AI Research Atlas

The week in AI, 21 to 27 September 2026

Weekly brief · 21 to 27 September 2026

The week in brief

Anthropic released Claude Opus 5.5 on 22 September, cutting its price by 20% to $4 per million input tokens and reporting 66.4% on Terminal-Bench 4.0.

The same day, OpenAI released GPT-6 Sol and GPT-6 Luna, two cheaper tiers of GPT-6 priced at half their GPT-5.6 equivalents. Xiaomi had opened the week on 21 September with MiMo-V2.6-Pro, which Artificial Analysis rated the top open-weight model at launch. Anthropic also reported on 23 September that Claude agents found a previously uncharacterized enzyme family, which its own wet lab then confirmed.

Anthropic releases Claude Opus 5.5 at $4 per million input tokens

Claude Opus 5.5, the first model in Anthropic's Claude 5.5 family, scores 66.4% on Terminal-Bench 4.0 by Anthropic's count and is priced 20% below Opus 5.

Anthropic reports 66.4% on Terminal-Bench 4.0 at its xhigh effort setting. Its comparison figures are 55.8% for its own Fable 5.1, 52.3% for Opus 5 and 57.9% for OpenAI's GPT-6 Astra. On FrontierCode v1.1 (Main) Anthropic reports 54.4%, close to GPT-6 Astra's 53.3%. On OSWorld 2.1 it reports 81.8%, against 80.7% for Fable 5.1.

Two of the published results come from outside Anthropic. On GDPval-AA v2.1, a third-party evaluation of work tasks, Opus 5.5 rates 1846 Elo, against 1735 for Fable 5.1 and 1542 for GPT-6 Astra. Zapier ran AutomationBench without fallbacks and scored Opus 5.5 at 40.0%, just under GPT-6 Astra's 41.4%.

The price is $4 per million input tokens and $20 per million output tokens. Cache reads fall 60% to $0.20 per million tokens, and a fast mode costs $8 per million input and $40 per million output. Anthropic says output is over 30% faster. Adaptive thinking is always on, so the model decides how much to reason on each request, and the context window is 1 million tokens.

Anthropic says Opus 5.5 is better at long code migrations and at resisting prompt injection. In its testing, the rate at which the model tried to cross containment boundaries fell by about 85% compared with Opus 5 or Mythos 5.1. It ships with the same cyber and biology safeguards Anthropic built for Fable.

OpenAI halves prices with GPT-6 Sol and GPT-6 Luna

On 22 September OpenAI released GPT-6 Sol at $2 per million input tokens and GPT-6 Luna at $0.10, half the price of their GPT-5.6 equivalents.

Sol and Luna are cheaper tiers of GPT-6, below GPT-6 Astra. Sol costs $2 per million input tokens and $10 per million output tokens, down from $4 and $20. Luna costs $0.10 per million input tokens and $0.50 per million output tokens, which makes it one of the cheapest models OpenAI has sold. Both have a 1-million-token context window.

OpenAI reports that Sol at maximum effort scores 68.8% on DeepSWE v1.1. That is about 1.1 points under Claude Fable 5 at xhigh effort, and OpenAI says Sol does it at roughly 80% lower cost per task. OpenAI also said GPT-5.6 will rise in price by 25% in November, which gives developers a reason to move to the new tiers.

DeepSWE v1.1 figures appeared from several labs this week, all company-reported and run at different effort settings. xAI reports 71.0% for Grok 4.7 and Xiaomi reports 72.57 for MiMo-V2.6-Pro.

Xiaomi releases MiMo-V2.6-Pro, top open-weight model on Artificial Analysis

Xiaomi released MiMo-V2.6-Pro and MiMo-V2.6-Flash under the MIT license on 21 September, and Artificial Analysis rated Pro at 46, the highest score of any open-weight model at launch.

MiMo-V2.6-Pro has 1.02 trillion parameters and Flash has 310 billion. Both take text, images and other modalities as input and have a 1-million-token context window. On the Artificial Analysis Intelligence Index, Pro's 46 puts it just ahead of Zhipu's GLM-5.3 at 45 and Moonshot's Kimi K3 at 44.

Xiaomi trained both models with one reinforcement learning (RL) run that mixed coding, agent tasks, vision and cybersecurity. It used an asynchronous form of GRPO, a method that scores each sampled answer against the other answers to the same prompt, and it graded agent runs in groups. Xiaomi puts the RL cost at about $2.62 million for Pro and $850,000 for Flash, as reported by TestingCatalog and Winbuzzer.

On DeepSWE v1.1, Xiaomi reports that Pro rose from 58.4 for V2.5-Pro to 72.57. API prices for Pro are $0.435 per million input tokens and $0.87 per million output tokens. An UltraSpeed edition costs ten times as much and runs about 20 times faster.

Claude agents find a new enzyme family, confirmed in Anthropic's lab

Anthropic reports that about 950 Claude agents, working for 21 hours and using 210 million tokens, found a previously uncharacterized reverse transcriptase system with CRISPR-like DNA repeats.

Anthropic announced the result on 23 September together with a new Anthropic life-sciences lab. Claude searched sequence databases for reverse transcriptases, which are enzymes that copy RNA into DNA. Beside one enzyme from a bacteriophage it spotted an array of repeated non-coding DNA, similar in layout to the repeats in CRISPR systems.

Anthropic's wet lab then confirmed it as a new family, which Anthropic calls "array-associated reverse transcriptase". According to Anthropic, humans supplied only the prompt and the lab work. The function of the new enzyme family is still unknown.

Also in the news

  • xAI released Grok 4.7 on 21 September, a month after Grok 4.6, at the same $2 per million input tokens and $6 per million output tokens, with a 500,000-token context and xAI-reported scores of 71.0% on DeepSWE v1.1 and 46.3% on CursorBench 4.0.
  • Cognition said on 25 September that its annualized run-rate revenue from Devin and Windsurf has reached $1 billion, under two years after Devin became generally available.
  • Alibaba said at Apsara 2026 that Qwen3.8-Max ran 33 automated self-improvement cycles over about a month, lifting its Artificial Analysis score from 40 to 45 by the company's account, and that in a chip-design test it reportedly cut a bus module's area by 42%, with methods and baselines unpublished.
  • Alibaba also said Qwen 4 is in training with no release date, and reports cite goals of 5 to 10 trillion parameters for Qwen 4.5 and Qwen 5.
  • Anthropic opened Claude Marketplace on 23 September, listing plugins, connectors, purchasable agents and service partners, with over 2,000 connectors and plugins at launch.
  • Meta turned Muse into a consumer agent on 23 September, with a Mac app, its own email address, retail checkout partners and a real-time talking avatar with about 870 milliseconds of latency.