<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
<channel><title>AI Research Atlas, weekly</title><link>https://atlas.prashish.xyz/</link><description>The week&#x27;s main stories in AI and why they matter.</description><language>en</language>
<atom:link href="https://atlas.prashish.xyz/weeks/feed.xml" rel="self" type="application/rss+xml"/>
<item><title>The week in AI, 5 to 11 October 2026</title><link>https://atlas.prashish.xyz/weeks/2026-10-05</link><guid isPermaLink="true">https://atlas.prashish.xyz/weeks/2026-10-05</guid><pubDate>Tue, 06 Oct 2026 21:00:00 +0000</pubDate><description>Mistral AI opened a preview API for Mistral Large 4 on 6 October, a 1 trillion parameter multimodal model that it says will get open weights by the end of October.</description><content:encoded><![CDATA[<p>Week of 5 to 11 October 2026, with 5 and 6 October covered so far. Mistral previewed the 1 trillion parameter Mistral Large 4, Reflection AI unveiled its first model Beam, Anthropic folded Project Glasswing into an expanded Cyber Verification Program, Google DeepMind released the open EmbeddingGemma 2, and Google cut free Gemini users to Flash Lite.</p>
<h2 id="the-week-in-brief">The week in brief</h2>
<p class="lede">Mistral AI opened a preview API for Mistral Large 4 on 6 October, a 1 trillion parameter multimodal model that it says will get open weights by the end of October.</p>
<p>The first two days of the week brought two large open-weight models from Western labs, and neither can be downloaded yet. <strong>Mistral AI</strong> previewed <strong>Mistral Large 4</strong> on 6 October, and <strong>Reflection AI</strong> unveiled <strong>Beam</strong>, its first model, on 5 October. Both companies say weights follow later in October. All of their performance claims are company-reported, and no independent benchmark results exist for either model so far.</p>
<p>Anthropic widened access to its cyber-capable models on 6 October, six days after Google gave Gemini 4 Argon to cyber defenders first. Google released <strong>EmbeddingGemma 2</strong>, an open embedding model that handles text, code, images, video and audio in one space, and said free Gemini users will get only Flash Lite from 9 October. OpenAI added labeled image ads to ChatGPT, and Meta published an open protocol for shopping agents.</p>
<h2 id="mistral-previewed-the-1-trillion-parameter-mistral-large-4">Mistral previewed the 1 trillion parameter Mistral Large 4</h2>
<p class="lede">Mistral AI released a preview API for Mistral Large 4 on 6 October, a 1 trillion parameter model with 49 billion parameters active per token, and says open weights will follow by the end of October.</p>
<h3>What happened</h3>
<p>Mistral Large 4 is Mistral&#x27;s largest model so far, up from the 675 billion parameters of Mistral Large 3. It is multimodal, and it runs as a preview API on Mistral Studio. Mistral says it trained the model from scratch on 3,800 NVIDIA Grace Blackwell GPUs in its own European datacenters.</p>
<p>Large 4 is a mixture-of-experts model. Each token is routed through a small set of expert subnetworks, so only part of the model does work at any one time, and that keeps inference cost closer to a much smaller model. Mistral&#x27;s blog gives 1 trillion total and 49 billion active parameters. Mistral&#x27;s documentation lists 1.05 trillion total and 52 billion active, and the company hasn&#x27;t explained the gap.</p>
<h3>What Mistral claims</h3>
<p>Mistral says Large 4 significantly outperforms any open-weight model from the US or Europe and is competitive with the strongest open models globally. That second phrase points at the Chinese open-weight models, which Mistral does not claim to beat. The atlas has no independent benchmark result for Large 4, and Mistral did not publish scores the atlas could check on 6 October.</p>
<h3>Why it matters</h3>
<p>Until the weights ship, Mistral is red-teaming a version with reduced moderation together with cybersecurity partners. That follows directly from last week&#x27;s records. On 29 September Anthropic reported that Zhipu&#x27;s open-weights GLM-5.3 built working exploits about two-thirds as often as its own gated Mythos Preview, and an open-weight model can&#x27;t be recalled once released.</p>
<p>Mistral&#x27;s approach is a short closed window before an open release. Anthropic, OpenAI and Google gate their top models for longer and keep the weights. A lab releasing 1 trillion parameters of open weights has to decide its cyber position before release, and Mistral&#x27;s red-teaming period is how it is doing that.</p>
<h3>What we don&#x27;t know</h3>
<p>Which of Mistral&#x27;s two parameter counts is correct is unknown. The scores behind &quot;competitive with the strongest open models globally&quot; weren&#x27;t available to the atlas. Whether the released weights will be the moderated version or the one being red-teamed with reduced moderation is also unknown.</p>
<h2 id="reflection-ai-unveiled-beam-a-501-billion-parameter-model">Reflection AI unveiled Beam, a 501 billion parameter model</h2>
<p class="lede">Reflection AI unveiled Beam on 5 October, a text-only mixture-of-experts model with 501 billion total and 23 billion active parameters, and says it matches Z.ai&#x27;s GLM-5.2 on reasoning benchmarks with 3 to 4 times less inference compute.</p>
<h3>What happened</h3>
<p>Beam is Reflection AI&#x27;s first model release. It is text-only, aimed at coding and agent tasks, and has a 1 million token context window. Reflection says it pretrained Beam on 23.8 trillion tokens and then trained it heavily with reinforcement learning (RL), where the model improves by being scored on tasks it attempts.</p>
<p>Beam is available as a research preview. Reflection says the weights and a technical report will come later in October.</p>
<h3>The comparison with GLM-5.2</h3>
<p>Reflection chose a Chinese open model as its benchmark. GLM-5.2 is about 744 billion total and 40 billion active parameters, so Beam is smaller on both counts. Reflection says Beam matches GLM-5.2 on reasoning benchmarks while using 3 to 4 times less inference compute, and nobody has checked that claim independently.</p>
<p>Active parameters drive most of the per-token cost in a mixture-of-experts model. Beam has a little over half of GLM-5.2&#x27;s active parameters, which doesn&#x27;t by itself account for a 3 to 4 times saving. The technical report will have to show where the rest comes from.</p>
<h3>Why it matters</h3>
<p>Mistral and Reflection both announced large open-weight models in the same two days, and both compared themselves against open models from China. The GLM line is the same family Anthropic tested for exploit ability on 29 September with GLM-5.3. For a US practitioner choosing an open model, the options at this size have mostly come from Chinese labs, and Beam and Large 4 are the US and European entries for October.</p>
<h3>What we don&#x27;t know</h3>
<p>The benchmark scores behind &quot;matches GLM-5.2&quot; aren&#x27;t in the records. The license terms for the weights haven&#x27;t been announced. It&#x27;s also unknown how Beam compares with GLM-5.3, the newer model in the same family.</p>
<h2 id="anthropic-folded-project-glasswing-into-its-cyber-verificati">Anthropic folded Project Glasswing into its Cyber Verification Program</h2>
<p class="lede">Anthropic expanded its Cyber Verification Program on 6 October and folded Project Glasswing into it, with three access tiers for vetted cyber defenders using its most capable Claude models.</p>
<h3>What happened</h3>
<p>Project Glasswing began in April as the partner program through which Anthropic gave Claude Mythos Preview to a small set of organisations. On 6 October Anthropic merged it into its Cyber Verification Program, which now has three access tiers for vetted security professionals. The tier details weren&#x27;t in the sources the atlas could read.</p>
<p>Reuters, as summarized by Techmeme, reported Anthropic&#x27;s figures. Anthropic says it found more than 5,500 verified vulnerabilities between April and October. It says Glasswing partners found more than 129,000 vulnerabilities between April and July, of which more than 33,000 were critical or high severity. These are company-reported counts.</p>
<h3>Why it matters</h3>
<p>The expansion came six days after Google DeepMind released Gemini 4 Argon to vetted cyber defenders first on 30 September. Anthropic, OpenAI and Google all now put their most cyber-capable models in front of defenders before the public. Anthropic&#x27;s change turns a small partner list into a program with entry levels, so more security teams can apply.</p>
<p>The partner count matters because of the patching numbers from earlier in the year. Anthropic&#x27;s May Glasswing update reported that only 75 of 530 disclosed high or critical open-source bugs had been patched. Finding vulnerabilities was already faster than fixing them in May, and a 129,000 count makes that gap more important.</p>
<h3>What is reported and unconfirmed</h3>
<p>The atlas&#x27;s record says vetted users get Anthropic&#x27;s top models with fewer safeguards for cyber work. That part didn&#x27;t appear on any evidence page the atlas checked, so the atlas treats it as reported and unconfirmed. The atlas also has no patch rate to go with the April to July vulnerability count.</p>
<h2 id="google-deepmind-released-embeddinggemma-2-for-five-input-typ">Google DeepMind released EmbeddingGemma 2 for five input types</h2>
<p class="lede">Google DeepMind released EmbeddingGemma 2 with open weights on 6 October, a 740 million parameter model that maps text, code, images, video and audio into one 768-dimensional space.</p>
<h3>What happened</h3>
<p>An embedding model turns an input into a list of numbers, a vector, so that similar inputs land close together. Search, retrieval for chatbots and recommendation systems all run on these vectors. The first EmbeddingGemma handled text only, and EmbeddingGemma 2 is built on Gemma 4 and adds code, images, video and audio in the same shared space.</p>
<p>The encoders are modular. A developer can load only the text and code part at 270 million parameters, text with images and video at 440 million, text with audio at 570 million, or everything at 740 million. Google reports about 191MB of active RAM for the text-only weights and about 567MB for the full model on a Pixel 11 Pro.</p>
<h3>Why it matters</h3>
<p>A single shared space removes a chain of models. Until now, searching a video library by text meant running captioning or speech-to-text first and then embedding the text output. With one space, a text query and a video clip can be compared directly, and Google&#x27;s RAM figures put that on a phone.</p>
<p>The release is small next to the trillion-parameter announcements, and it&#x27;s the one item this week a developer can download and use today. It&#x27;s open weights and its claims are verified against Google&#x27;s own pages.</p>
<h3>What we don&#x27;t know</h3>
<p>The atlas has no independent retrieval benchmark for EmbeddingGemma 2. How much quality the 270 million text-only version gives up against the full 740 million model wasn&#x27;t in the records.</p>
<h2 id="google-cut-free-gemini-users-to-flash-lite-from-9-october">Google cut free Gemini users to Flash Lite from 9 October</h2>
<p class="lede">From 9 October free Gemini users will get only Flash Lite, standard Flash will need the $4.99 per month Google AI Plus plan, and AI Plus will lose Gemini Pro.</p>
<h3>What changes</h3>
<p>Free Gemini users can currently pick Flash Lite, Flash or Pro. From 9 October they get Flash Lite only. Standard Flash moves to Google AI Plus at $4.99 per month, and AI Plus loses Gemini Pro.</p>
<p>Pro and the Deep Think reasoning option will be limited to Google AI Pro at $19.99 per month and Google AI Ultra at $99.99 per month. Low, medium and high effort levels stay available for each model, and The Verge reports that higher effort may use up the usage limit faster.</p>
<h3>Why it matters</h3>
<p>Google has pulled the free tier down by two models. Anyone demonstrating Gemini on a free account after 9 October will be showing Flash Lite, and comparisons people make between free ChatGPT and free Gemini will be comparing different tiers than they were the week before.</p>
<p>OpenAI made a different move with its free users on 5 October and announced image ads in ChatGPT (see &quot;Also worth knowing&quot;). Google is moving free users to a cheaper model, and OpenAI is adding ads to its free and Go tiers.</p>
<h3>What we don&#x27;t know</h3>
<p>Google hasn&#x27;t given a reason in the records the atlas holds. Usage limits for each tier after 9 October weren&#x27;t published in the source.</p>
<h2 id="also-worth-knowing">Also worth knowing</h2>
<h3>Agents and commerce</h3>
<ul><li><strong>Personal Agent Protocol</strong> is an open standard published on 6 October by Meta with Walmart, Stripe, Sierra and others, meant to help websites tell bots acting for a user from malicious ones. Amazon has begun deliberately blocking Meta&#x27;s Muse agent from its retail site, and standard anti-bot checks often block other personal agents too. Meta presented Muse as a shopping-capable agent on 23 September.</li><li><strong>Ironclad</strong> worked with OpenAI to turn 11 contracting tasks into research problems graded on 8 to 50 criteria each, and OpenAI reports GPT-6 Astra scored 55.0% against 41.6% for GPT-5.6 Sol. OpenAI says Astra&#x27;s average time per attempt fell from 37.0 minutes to 19.2 minutes, and an unreleased internal model scored 63.7%. The OpenAI page carries no publication date, so the atlas logged it on 6 October when it appeared.</li></ul>
<h3>Products and pricing</h3>
<ul><li><strong>ChatGPT image ads</strong> were announced by OpenAI on 5 October, with labeled ads shown next to image generation results for Free and Go users and US testing starting later in October. OpenAI says ChatGPT has 1.2 billion weekly users, and DV Rockerbox reports WeightWatchers saw a 15.3% lower cost per acquisition than its blended paid-search benchmark.</li><li><strong>Nano Banana 2.1</strong> is Google DeepMind&#x27;s new Flash-tier image generation and editing model, listed on OpenRouter at $1.50 per million input tokens, $7.50 per million output tokens and $30 per million image output tokens. OpenRouter&#x27;s listing says it improves mask and ink editing and product recontextualization, with output up to 4K.</li><li><strong>ChatGPT text in the EU</strong> will reportedly carry watermarks from OpenAI, and the atlas has no confirming record or start date.</li></ul>
<h3>Science</h3>
<ul><li><strong>Vals AI</strong> reports that a team of Claude Opus 5.5 agents found two candidate room-temperature antiferromagnetic semiconductors for computer memory. The candidates are predicted by calculation and haven&#x27;t been tested in a lab.</li></ul>
<h3>Regional models</h3>
<ul><li><strong>Falcon-Emirati-7B</strong> was announced by the Technology Innovation Institute, a model built on Falcon-H1-Arabic and specialised in the Emirati Arabic dialect and culture.</li></ul>
<h3>Money</h3>
<ul><li><strong>DeepSeek</strong> is reportedly close to a $12 billion funding round backed by Tencent. Last week DeepSeek released versions of its MoE communication and FP8 GEMM libraries for Huawei&#x27;s Ascend chips.</li><li><strong>Kling</strong>, Kuaishou&#x27;s video unit, has reportedly picked banks for a Hong Kong IPO worth more than $1 billion, according to The Information. Kling 4.0 entered early access on 28 September.</li></ul>
<h2 id="people-this-week">People this week</h2>
<p>The atlas logged no moves of people on 5 or 6 October.</p>
<h2 id="what-did-not-change">What did not change</h2>
<p>Every capability number in this week&#x27;s stories is company-reported. No independent evaluator has published results for Mistral Large 4, Beam or EmbeddingGemma 2, and OpenAI&#x27;s Ironclad figures come from an undated OpenAI page.</p>
<p>The most cyber-capable models from Anthropic, OpenAI and Google still go to vetted defenders before anyone else. Anthropic&#x27;s 6 October expansion adds tiers and more applicants to that arrangement.</p>
<p>Neither of this week&#x27;s large open-weight models can be downloaded yet. Mistral and Reflection both give October dates, and until then the strongest downloadable models at this size are still the Chinese ones they compare against.</p>
<h2 id="what-to-watch-next">What to watch next</h2>
<ul><li>On 9 October Google&#x27;s Gemini free tier drops to Flash Lite only, and AI Plus loses Gemini Pro.</li><li>Later in October Reflection AI says it will release Beam&#x27;s weights and a technical report, which should show the benchmark scores behind its GLM-5.2 comparison.</li><li>By the end of October Mistral says Large 4&#x27;s weights will ship, and the release will show which moderation settings the public version carries.</li><li>Later in October OpenAI begins US testing of image ads in ChatGPT for Free and Go users.</li><li>Anthropic has given no date for publishing the details of its three Cyber Verification Program tiers.</li><li>DeepSeek&#x27;s reported $12 billion round has no announced closing date.</li></ul>]]></content:encoded></item>
<item><title>The week in AI, 28 September to 4 October 2026</title><link>https://atlas.prashish.xyz/weeks/2026-09-28</link><guid isPermaLink="true">https://atlas.prashish.xyz/weeks/2026-09-28</guid><pubDate>Sun, 04 Oct 2026 21:00:00 +0000</pubDate><description>Google released Gemini 4 Argon to vetted cyber defenders first, and Anthropic and OpenAI shipped Claude Sonnet 5.5 and GPT-6.1 Sol at $2/$10 per million tokens.</description><content:encoded><![CDATA[<p>Week of 28 September to 4 October 2026, with the previous week for context. Five stories follow. Google released Gemini 4 Argon to vetted security teams first, Anthropic and OpenAI shipped Claude Sonnet 5.5 and GPT-6.1 Sol at $2/$10, OpenAI launched its dots assistants, Anthropic reported on the open-weights GLM-5.3 and OpenAI on a distillation campaign, and AMD agreed to buy World Labs. Details come from the atlas release and talent logs, which are unverified drafts, so treat specific numbers as leads.</p>
<h2 id="the-week-in-brief">The week in brief</h2>
<p class="lede">Google released Gemini 4 Argon to vetted cyber defenders first, and Anthropic and OpenAI shipped Claude Sonnet 5.5 and GPT-6.1 Sol at $2/$10 per million tokens.</p>
<p>Google DeepMind&#x27;s <strong>Gemini 4 Argon</strong> (30 September) was the headline release. It went first to vetted cyber defenders, with developers, enterprises and paying consumers to follow. That makes three of the big labs, Anthropic with Mythos in April, OpenAI with GPT-6 Astra in September and now Google, that release their top model in stages.</p>
<p>Below that top tier, the releases competed on price. Anthropic&#x27;s <strong>Claude Sonnet 5.5</strong> (28 September) kept Sonnet&#x27;s $2/$10 price and posted a large jump on Terminal-Bench 4.0. OpenAI&#x27;s <strong>GPT-6.1 Sol</strong> (29 September) claimed near-Astra results on agentic coding at about a fifth of Astra&#x27;s price. Both followed the previous week&#x27;s Claude Opus 5.5 and GPT-6 Sol/Luna. Capability first appears in a gated, expensive tier and reaches a cheaper model within weeks.</p>
<p>OpenAI launched <strong>dots</strong> at DevDay, persistent assistants with their own cloud computer and identity in Slack. Anthropic had merged Cowork and chat into one Claude two weeks before, and Meta turned Muse into a shopping-capable agent. Two security reports covered Anthropic&#x27;s study of the open-weights <strong>GLM-5.3</strong> model&#x27;s exploit ability and an OpenAI report on a <strong>distillation campaign</strong> it attributes partly to people associated with Moonshot AI. Both show the cost of the diffusion speed that the rest of this atlas tracks. On the research-lab side, AMD agreed to buy <strong>World Labs</strong>.</p>
<h2 id="google-released-gemini-4-argon-to-vetted-cyber-defenders-fir">Google released Gemini 4 Argon to vetted cyber defenders first</h2>
<p class="lede">Gemini 4 Argon went to vetted cyber defenders before anyone else, which makes Google the third lab after Anthropic (Mythos Preview) and OpenAI (GPT-6 Astra) to put its top model behind a defender-first gate.</p>
<h3>What happened</h3>
<p>Google DeepMind introduced Gemini 4 Argon on 30 September as a new frontier model aimed at long, complex work, with a one-million-token output limit as well as its input context and a reported 77.9% on the DeepSWE v1.1 software-engineering benchmark. The first users were vetted cyber defenders through a Google programme, under the US voluntary pre-release access process; developers, enterprises and consumers on paid plans are scheduled later. Google also said its own agents had used the model internally before launch.</p>
<h3>Why Anthropic, OpenAI and Google gate their top models</h3>
<p>Staged release is an old idea. OpenAI staged GPT-2&#x27;s release in 2019 on misuse grounds and was mocked for it. In 2026 the models became useful for offensive security. In April, Anthropic said Claude Mythos Preview found software vulnerabilities better than all but the most skilled humans and gave it only to partners through Project Glasswing. In September OpenAI rated GPT-6 Astra &quot;Critical&quot; for cybersecurity under its own Preparedness Framework, its first model at that level. Argon makes Google the third lab to put its top model behind a defender-first gate.</p>
<p>If a model can find thousands of unknown vulnerabilities, a head start for defenders to patch them before attackers get the same capability is the main lever a lab has. Governments are now part of the release process, through voluntary pre-release testing and, in Anthropic&#x27;s June case, an export-control directive that briefly suspended Fable 5 and Mythos 5.</p>
<h3>Three layers of model access</h3>
<p>The frontier model is now one most users cannot use yet. There are three layers to keep straight. The top tier is gated (Mythos, Astra, Argon), a public flagship sits one step below, and cheap fast tiers inherit the top tier&#x27;s training within weeks. When someone asks which model is best, the answer depends on which layer they have access to.</p>
<h3>What is still unmeasured</h3>
<p>Whether defender-first windows reduce harm has not been measured. Nobody has published a full count of vulnerabilities fixed before general release, though Anthropic&#x27;s May Glasswing update reported only 75 of 530 disclosed high or critical open-source bugs patched. It is also unknown whether attackers gained similar ability from other models in the meantime, and Story 4 suggests they partly can. Argon&#x27;s general-availability date was not public by 4 October. Google has announced an introductory price of $2/$10 per million input/output tokens, rising to $4/$20.</p>
<h2 id="claude-sonnet-5-5-and-gpt-6-1-sol-both-launched-at-2-10-per-">Claude Sonnet 5.5 and GPT-6.1 Sol both launched at $2/$10 per million tokens</h2>
<p class="lede">Sonnet 5.5 scored 70.6% on Terminal-Bench 4.0 against 10.3% for Sonnet 5 at an unchanged price, and OpenAI says GPT-6.1 Sol matches GPT-6 Astra on DeepSWE v1.1 at about one-fifth the cost.</p>
<h3>The releases</h3>
<ul><li><strong>Claude Sonnet 5.5</strong> (28 September) scored 70.6% on Terminal-Bench 4.0, against 10.3% for Sonnet 5 on the same benchmark. It is about 30% faster, and the price is unchanged at $2/$10 per million input/output tokens. Anthropic says it is the first Sonnet to beat Pokémon Red from screenshots alone, and it scores two points below Opus 5.5 on GDPval-AA.</li><li><strong>GPT-6.1 Sol</strong> (29 September) is an upgrade to GPT-6 Sol that OpenAI says matches Astra on DeepSWE v1.1 at about one-fifth the cost, at $2/$10. OpenAI treats it as Critical-capability for cyber purposes.</li><li><strong>Last week&#x27;s context</strong> included Claude Opus 5.5 (22 September), which matched Fable 5.1 on most work at 40% less than Opus 5, priced at $4/$20. GPT-6 Sol and Luna (22 September) brought Astra-era training to tiers at $2/$10 and $0.10/$0.50. xAI&#x27;s Grok 4.7 (21 September) shipped about six weeks after 4.6, and Xiaomi&#x27;s open MiMo-V2.6-Pro (21 September) became the top-rated open-weights model on Artificial Analysis at launch.</li></ul>
<h3>Distillation and architecture efficiency</h3>
<p>Two things push capability down the price ladder. The first is <strong>distillation</strong>, where a cheaper model is trained on the outputs and behaviour of the gated top model and inherits much of its skill at a fraction of the inference cost. The second is <strong>architecture efficiency</strong>, since sparse mixture-of-experts, sparse and linear attention, and lower-precision arithmetic mean each answer needs less compute. DeepSeek&#x27;s V4.1-Flash (10 September) and its new libraries for Huawei&#x27;s Ascend chips (29 September) show how much of the cost reduction is now engineering.</p>
<h3>Price history since GPT-4</h3>
<p>The field has followed this curve since 2023. GPT-4 launched at $30/$60 per million tokens, and within eighteen months models matching it cost under a dollar. The B08 deep dive records Epoch AI&#x27;s estimate that the cost of reaching a fixed benchmark score has been falling by roughly half every quarter. In 2026 the move from gated frontier to mid-tier takes weeks.</p>
<h3>Caveats on the comparisons</h3>
<p>Benchmark comparisons across labs use different harnesses and effort settings, and &quot;matches the top model on most work&quot; is a company claim. Cost per completed task, which is what matters for agents, depends on how many tokens a model spends thinking as well as its list price. Independent cost-per-task measurements for these releases were not yet available.</p>
<h2 id="openai-launched-dots-assistants-with-their-own-cloud-compute">OpenAI launched dots, assistants with their own cloud computer and Slack identity</h2>
<p class="lede">OpenAI&#x27;s dots have their own cloud computer and Slack identity, Anthropic&#x27;s merged Claude keeps working after the laptop closes, and Meta announced a Muse agent with its own email address.</p>
<h3>What launched</h3>
<ul><li><strong>OpenAI dots</strong> (DevDay, 29 September) are proactive assistants, each with a name, an identity in Slack and its own cloud browser and computer, that keep working on projects and can use your laptop with permission. They run on GPT-6 Astra and launched for Pro, Business Premium and Enterprise users.</li><li><strong>One Claude</strong> (16 September) is Anthropic&#x27;s merger of Cowork and chat into a single experience, with Docs, Slides and Design available in any conversation and work that continues after the laptop closes. It was followed by <strong>Claude Marketplace</strong> (23 September, 2,000+ connectors and plugins) and <strong>Claude Code mods</strong> (1 October), small functions that change how the coding agent behaves.</li><li><strong>Meta Muse agent</strong> (23 September) was presented as a consumer agent with checkout partners, plus a Mac app, its own email address and a real-time talking avatar, most of them announced as coming soon.</li><li><strong>Cognition</strong> said on 25 September that its annualized revenue had passed $1 billion, less than two years after Devin became generally available.</li></ul>
<h3>What an agent needs besides a model</h3>
<p>Each of these products is a worker you delegate to, and the model you prompt is one part of it. That puts memory across sessions, permissions, identity, sandboxes and recovery from mistakes over hours at the center. It also pushes pricing from per-token toward per-seat or per-outcome. The talent and acquisition stories of 2025 to 2026 (Windsurf, Manus, Cursor) were all bets on owning this layer.</p>
<h3>Where these agents came from</h3>
<p>The line runs from ReAct and function calling (2022 to 2023), through Claude computer use and MCP (late 2024) and Claude Code and Deep Research (early 2025), to Cowork and the coding-agent boom. See the Connections page, &quot;Reasoning → tools → feedback → longer tasks,&quot; and the B10 deep dive.</p>
<h3>Reliability over long unsupervised runs</h3>
<p>Reliability over long unsupervised runs is still unproven outside company demos. Demos show agents working for hours, but independent evidence on how often they finish real business tasks correctly, and what it costs when they do not, is thin. Security also scales with autonomy, because an agent with its own computer and your Slack identity is a new attack surface for prompt injection.</p>
<h2 id="anthropic-tested-glm-5-3-on-exploits-and-openai-reported-a-d">Anthropic tested GLM-5.3 on exploits and OpenAI reported a distillation campaign</h2>
<p class="lede">Anthropic reported that Zhipu&#x27;s open-weights GLM-5.3 can build working exploits about two-thirds as often as its own gated Mythos Preview, and OpenAI described a coordinated campaign to distil its protected reasoning.</p>
<h3>The GLM-5.3 study</h3>
<p>On 29 September Anthropic published an evaluation of Zhipu&#x27;s open-weights GLM-5.3. In its tests the model achieved a full control-flow hijack in 4% of trials, against 6% for Mythos Preview, and its built-in safeguards were bypassed in 64 to 100% of attempts depending on the method. Anthropic concluded that, about five months after it gated Mythos Preview for being too capable at offensive security, a freely downloadable model is in the same range.</p>
<h3>The distillation campaign</h3>
<p>On 30 September OpenAI said it had disrupted a coordinated campaign to extract protected reasoning from its models, with activity from July peaking at about 16,000 requests from more than 4,000 accounts in two days, and attributed a core cluster to people associated with Moonshot AI. This follows OpenAI&#x27;s distillation concerns about DeepSeek in January 2025 and Anthropic&#x27;s February 2026 report naming DeepSeek, Moonshot and MiniMax.</p>
<h3>What this does to staged release</h3>
<p>Both reports bear on staged release (Story 1), which assumes a lab can control who gets a capability for a meaningful period. Fast open replication, sometimes helped by distillation, shortens that period. If the window is a few months, gating is worth most as a head start for defenders, since it cannot keep the capability scarce. That is a sharper version of the diffusion pattern on the Spread page.</p>
<h3>Caveats on both reports</h3>
<p>Both reports come from interested parties, since Anthropic and OpenAI compete with Chinese open-weights labs and lobby on export policy. The attributions and the exploit-rate comparison were not independently replicated by 4 October.</p>
<h2 id="amd-agreed-to-buy-fei-fei-li-s-world-labs-in-a-reported-8-2-">AMD agreed to buy Fei-Fei Li&#x27;s World Labs in a reported $8.2 billion stock deal</h2>
<p class="lede">AMD agreed to buy World Labs, and David Silver&#x27;s Ineffable Intelligence added six cofounders from DeepMind, InstaDeep and Flying Fish.</p>
<h3>What happened</h3>
<p>On 28 September AMD agreed to acquire World Labs, Fei-Fei Li&#x27;s spatial-intelligence company, in a deal reported at about $8.2 billion in stock; Li becomes AMD&#x27;s chief scientist reporting to Lisa Su. World Labs had shipped Marble (persistent, explorable 3D worlds) in November 2025. Earlier in September, Ineffable Intelligence, founded by AlphaGo lead David Silver to pursue superintelligence through reinforcement learning on experience, named six new cofounders, four from Google DeepMind, one from InstaDeep and one from the venture firm Flying Fish. Google DeepMind completed a reported $1.5 billion-plus talent deal that brought in Mechanize&#x27;s Tamay Besiroglu.</p>
<h3>Why a chip vendor wants a world-model lab</h3>
<p>AMD&#x27;s purchase is a bet that simulated 3D worlds will be a major compute workload, for robotics training, games and design, and that owning the models helps sell the hardware. It also follows other 2026 cases in which independent &quot;age of research&quot; labs either raised very large rounds (AMI Labs, Ineffable) or were absorbed. See Next bets for the world-models and RL-from-experience directions.</p>
<h2 id="also-worth-knowing">Also worth knowing</h2>
<h3>Media and voice</h3>
<ul><li><strong>Eleven v4</strong> (28 September) is ElevenLabs&#x27; most expressive speech model, with inline delivery tags, consistent multi-speaker dialogue, 90+ languages and a roughly 150 ms Turbo variant. It came weeks after ElevenLabs&#x27; first major-label deal with Universal Music.</li><li><strong>Kling 4.0</strong> (28 September, early access) makes native 30-second clips, with up to ten keyframes and fifteen reference inputs. Video generation now competes on length, control and sound as well as image quality.</li><li><strong>FLUX 3 Image</strong> (1 October) from Black Forest Labs adds layout control by bounding boxes and edits that leave everything outside the box untouched. This matters for agents that edit images repeatedly.</li><li><strong>Suno Speech</strong> (1 October, beta) generates voice and music as one track. Separately, Universal and Sony filed a second suit against Suno over v6 on 18 September.</li></ul>
<h3>Science and biology</h3>
<ul><li>Anthropic reported (23 September) that about 950 Claude agents, over 21 hours, flagged a previously uncharacterized enzyme system with CRISPR-like repeats. It launched a life-sciences lab at the same time, six days after opening a verification programme (17 September) giving vetted biology teams models with loosened biology safeguards.</li><li>Google DeepMind&#x27;s <strong>SynthID Bio</strong> (30 September) watermarks AI-designed proteins in the sequence itself.</li><li>Microsoft&#x27;s <strong>Quine</strong> (29 September) is a biology research system that proposes interventions before wet-lab tests, limited to a fellows programme.</li></ul>
<h3>Infrastructure</h3>
<p>DeepSeek described its sandbox platform for agent RL, about three million sandboxes a day, and released versions of its core training libraries for Huawei&#x27;s Ascend chips. Both show how Chinese labs are building around US export controls.</p>
<h2 id="people-this-week">People this week</h2>
<ul><li><strong>Jacob Coxon</strong> resigned from Anthropic on 8 September and published a widely shared warning that labs are compromising oversight to keep pace.</li><li><strong>David Robinson</strong> left OpenAI and published an essay in <em>The Atlantic</em> (3 October) arguing that the company&#x27;s optimism, as it sprints between launches, falls short of its responsibilities.</li><li><strong>Andrew Tulloch</strong> left Meta Superintelligence Labs (reported 9 September), a year after Meta recruited him from Thinking Machines; his destination was unconfirmed.</li><li><strong>Barret Zoph</strong> moved from OpenAI to Google DeepMind in late August as vice president of research, working on RL and post-training, seven months after returning to OpenAI from Thinking Machines.</li></ul>
<p>As in 2025 and 2026, senior researchers move between the three or four best-funded labs, and a few leave publicly with criticism of safety practices.</p>
<h2 id="what-did-not-change">What did not change</h2>
<ul><li>Public benchmark gains still do not establish reliability on a user&#x27;s own long-running workflow. Most headline numbers this week are company-reported.</li><li>No research-first lab (SSI, AMI Labs, Ineffable) has released a model, so the &quot;age of research&quot; thesis has so far been tested only with money and hiring.</li><li>World-model and robotics demonstrations have not settled how well these systems transfer to messy physical environments.</li><li>The gap between gated and public models is still weeks to months.</li></ul>
<h2 id="what-to-watch-next">What to watch next</h2>
<ul><li><strong>Argon&#x27;s wider rollout</strong> will show when developers and paying users get access and whether the $2/$10 introductory price (then $4/$20) holds.</li><li><strong>Independent cost-per-task results</strong> for Sonnet 5.5, GPT-6.1 Sol and Opus 5.5 on agentic work.</li><li><strong>Measured outcomes from deployed agents</strong> (dots, Cowork, Devin), beyond company demos.</li><li><strong>Policy responses</strong> to open-weights cyber capability and distillation, especially in the US.</li><li><strong>Qwen 4</strong>, which Alibaba says is in training, and whether DeepSeek ships a V4 successor trained on Ascend.</li></ul>]]></content:encoded></item>
<item><title>The week in AI, 21 to 27 September 2026</title><link>https://atlas.prashish.xyz/weeks/2026-09-21</link><guid isPermaLink="true">https://atlas.prashish.xyz/weeks/2026-09-21</guid><pubDate>Sun, 27 Sep 2026 21:00:00 +0000</pubDate><description>Anthropic released Claude Opus 5.5 on 22 September, cutting its price by 20% to $4 per million input tokens and reporting 66.4% on Terminal-Bench 4.0.</description><content:encoded><![CDATA[<h2 id="the-week-in-brief">The week in brief</h2>
<p class="lede">Anthropic released Claude Opus 5.5 on 22 September, cutting its price by 20% to $4 per million input tokens and reporting 66.4% on Terminal-Bench 4.0.</p>
<p>The same day, OpenAI released GPT-6 Sol and GPT-6 Luna, two cheaper tiers of GPT-6 priced at half their GPT-5.6 equivalents. Xiaomi had opened the week on 21 September with MiMo-V2.6-Pro, which Artificial Analysis rated the top open-weight model at launch. Anthropic also reported on 23 September that Claude agents found a previously uncharacterized enzyme family, which its own wet lab then confirmed.</p>
<h2 id="anthropic-releases-claude-opus-5-5-at-4-per-million-input-to">Anthropic releases Claude Opus 5.5 at $4 per million input tokens</h2>
<p class="lede">Claude Opus 5.5, the first model in Anthropic&#x27;s Claude 5.5 family, scores 66.4% on Terminal-Bench 4.0 by Anthropic&#x27;s count and is priced 20% below Opus 5.</p>
<p>Anthropic reports 66.4% on Terminal-Bench 4.0 at its xhigh effort setting. Its comparison figures are 55.8% for its own Fable 5.1, 52.3% for Opus 5 and 57.9% for OpenAI&#x27;s GPT-6 Astra. On FrontierCode v1.1 (Main) Anthropic reports 54.4%, close to GPT-6 Astra&#x27;s 53.3%. On OSWorld 2.1 it reports 81.8%, against 80.7% for Fable 5.1.</p>
<p>Two of the published results come from outside Anthropic. On GDPval-AA v2.1, a third-party evaluation of work tasks, Opus 5.5 rates 1846 Elo, against 1735 for Fable 5.1 and 1542 for GPT-6 Astra. Zapier ran AutomationBench without fallbacks and scored Opus 5.5 at 40.0%, just under GPT-6 Astra&#x27;s 41.4%.</p>
<p>The price is $4 per million input tokens and $20 per million output tokens. Cache reads fall 60% to $0.20 per million tokens, and a fast mode costs $8 per million input and $40 per million output. Anthropic says output is over 30% faster. Adaptive thinking is always on, so the model decides how much to reason on each request, and the context window is 1 million tokens.</p>
<p>Anthropic says Opus 5.5 is better at long code migrations and at resisting prompt injection. In its testing, the rate at which the model tried to cross containment boundaries fell by about 85% compared with Opus 5 or Mythos 5.1. It ships with the same cyber and biology safeguards Anthropic built for Fable.</p>
<h2 id="openai-halves-prices-with-gpt-6-sol-and-gpt-6-luna">OpenAI halves prices with GPT-6 Sol and GPT-6 Luna</h2>
<p class="lede">On 22 September OpenAI released GPT-6 Sol at $2 per million input tokens and GPT-6 Luna at $0.10, half the price of their GPT-5.6 equivalents.</p>
<p>Sol and Luna are cheaper tiers of GPT-6, below GPT-6 Astra. Sol costs $2 per million input tokens and $10 per million output tokens, down from $4 and $20. Luna costs $0.10 per million input tokens and $0.50 per million output tokens, which makes it one of the cheapest models OpenAI has sold. Both have a 1-million-token context window.</p>
<p>OpenAI reports that Sol at maximum effort scores 68.8% on DeepSWE v1.1. That is about 1.1 points under Claude Fable 5 at xhigh effort, and OpenAI says Sol does it at roughly 80% lower cost per task. OpenAI also said GPT-5.6 will rise in price by 25% in November, which gives developers a reason to move to the new tiers.</p>
<p>DeepSWE v1.1 figures appeared from several labs this week, all company-reported and run at different effort settings. xAI reports 71.0% for Grok 4.7 and Xiaomi reports 72.57 for MiMo-V2.6-Pro.</p>
<h2 id="xiaomi-releases-mimo-v2-6-pro-top-open-weight-model-on-artif">Xiaomi releases MiMo-V2.6-Pro, top open-weight model on Artificial Analysis</h2>
<p class="lede">Xiaomi released MiMo-V2.6-Pro and MiMo-V2.6-Flash under the MIT license on 21 September, and Artificial Analysis rated Pro at 46, the highest score of any open-weight model at launch.</p>
<p>MiMo-V2.6-Pro has 1.02 trillion parameters and Flash has 310 billion. Both take text, images and other modalities as input and have a 1-million-token context window. On the Artificial Analysis Intelligence Index, Pro&#x27;s 46 puts it just ahead of Zhipu&#x27;s GLM-5.3 at 45 and Moonshot&#x27;s Kimi K3 at 44.</p>
<p>Xiaomi trained both models with one reinforcement learning (RL) run that mixed coding, agent tasks, vision and cybersecurity. It used an asynchronous form of GRPO, a method that scores each sampled answer against the other answers to the same prompt, and it graded agent runs in groups. Xiaomi puts the RL cost at about $2.62 million for Pro and $850,000 for Flash, as reported by TestingCatalog and Winbuzzer.</p>
<p>On DeepSWE v1.1, Xiaomi reports that Pro rose from 58.4 for V2.5-Pro to 72.57. API prices for Pro are $0.435 per million input tokens and $0.87 per million output tokens. An UltraSpeed edition costs ten times as much and runs about 20 times faster.</p>
<h2 id="claude-agents-find-a-new-enzyme-family-confirmed-in-anthropi">Claude agents find a new enzyme family, confirmed in Anthropic&#x27;s lab</h2>
<p class="lede">Anthropic reports that about 950 Claude agents, working for 21 hours and using 210 million tokens, found a previously uncharacterized reverse transcriptase system with CRISPR-like DNA repeats.</p>
<p>Anthropic announced the result on 23 September together with a new Anthropic life-sciences lab. Claude searched sequence databases for reverse transcriptases, which are enzymes that copy RNA into DNA. Beside one enzyme from a bacteriophage it spotted an array of repeated non-coding DNA, similar in layout to the repeats in CRISPR systems.</p>
<p>Anthropic&#x27;s wet lab then confirmed it as a new family, which Anthropic calls &quot;array-associated reverse transcriptase&quot;. According to Anthropic, humans supplied only the prompt and the lab work. The function of the new enzyme family is still unknown.</p>
<h2 id="also-in-the-news">Also in the news</h2>
<ul><li><strong>xAI</strong> released Grok 4.7 on 21 September, a month after Grok 4.6, at the same $2 per million input tokens and $6 per million output tokens, with a 500,000-token context and xAI-reported scores of 71.0% on DeepSWE v1.1 and 46.3% on CursorBench 4.0.</li><li><strong>Cognition</strong> said on 25 September that its annualized run-rate revenue from Devin and Windsurf has reached $1 billion, under two years after Devin became generally available.</li><li><strong>Alibaba</strong> said at Apsara 2026 that Qwen3.8-Max ran 33 automated self-improvement cycles over about a month, lifting its Artificial Analysis score from 40 to 45 by the company&#x27;s account, and that in a chip-design test it reportedly cut a bus module&#x27;s area by 42%, with methods and baselines unpublished.</li><li><strong>Alibaba</strong> also said Qwen 4 is in training with no release date, and reports cite goals of 5 to 10 trillion parameters for Qwen 4.5 and Qwen 5.</li><li><strong>Anthropic</strong> opened Claude Marketplace on 23 September, listing plugins, connectors, purchasable agents and service partners, with over 2,000 connectors and plugins at launch.</li><li><strong>Meta</strong> turned Muse into a consumer agent on 23 September, with a Mac app, its own email address, retail checkout partners and a real-time talking avatar with about 870 milliseconds of latency.</li></ul>]]></content:encoded></item>
</channel></rss>