<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
<channel><title>AI Research Atlas, daily</title><link>https://atlas.prashish.xyz/</link><description>The day&#x27;s AI news, read and checked against the sources.</description><language>en</language>
<atom:link href="https://atlas.prashish.xyz/feed.xml" rel="self" type="application/rss+xml"/>
<item><title>Mistral previewed Large 4, a 1 trillion parameter model, with open weights due October</title><link>https://atlas.prashish.xyz/daily/2026-10-06</link><guid isPermaLink="true">https://atlas.prashish.xyz/daily/2026-10-06</guid><pubDate>Tue, 06 Oct 2026 21:00:00 +0000</pubDate><description>Mistral AI opened a preview API for Mistral Large 4, a 1 trillion parameter multimodal model, and says open weights will follow by the end of October.</description><content:encoded><![CDATA[<p class="lede">Mistral AI opened a preview API for Mistral Large 4, a 1 trillion parameter multimodal model, and says open weights will follow by the end of October.</p>
<p>Mistral Large 4 is Mistral&#x27;s largest model so far, up from the 675 billion parameters of Mistral Large 3. It is a mixture-of-experts model, so each token goes through a small set of expert subnetworks and only 49 billion parameters are active at a time. Those are the figures in Mistral&#x27;s blog. Mistral&#x27;s documentation lists 1.05 trillion total and 52 billion active, and the company hasn&#x27;t explained the gap.</p>
<p>Mistral says it trained the model from scratch on 3,800 NVIDIA Grace Blackwell GPUs in its own European datacenters. It claims Large 4 significantly outperforms any open-weight model from the US or Europe and is competitive with the strongest open models globally. No independent benchmark results are available yet. For now the preview runs only through the API on Mistral Studio, and Mistral is red-teaming a version with reduced moderation with cybersecurity partners until the weights ship.</p>
<p>On 5 October, Reflection AI unveiled <strong>Beam</strong>, its first model. Beam is a text-only mixture-of-experts model with 501 billion total and 23 billion active parameters, a 1 million token context window and 23.8 trillion pretraining tokens, and it is aimed at coding and agent tasks. Reflection says Beam matches Z.ai&#x27;s GLM-5.2 on reasoning benchmarks while using 3 to 4 times less inference compute, and nobody has checked that claim independently. GLM-5.2 is about 744 billion total and 40 billion active parameters. Reflection says the weights and a technical report will come later in October.</p>
<p>Google DeepMind released <strong>EmbeddingGemma 2</strong> with open weights on 6 October. An embedding model turns inputs into vectors so that similar things sit close together. This version puts text, code, images, video and audio into one shared 768-dimensional space, where the first EmbeddingGemma handled text only. The encoders are modular, so a developer can load the 270 million parameter text part alone or go up to 740 million with vision and audio. Google reports about 191MB of active RAM for the text-only weights and about 567MB for the full model on a Pixel 11 Pro.</p>
<h2 id="also-in-the-news">Also in the news</h2>
<ul><li><strong>Google</strong> will limit free Gemini users to Flash Lite from 9 October, and the $4.99 per month AI Plus plan will lose Gemini Pro, which will stay on the $19.99 AI Pro and $99.99 Ultra plans.</li><li><strong>Google DeepMind</strong> released Nano Banana 2.1, an image generation and editing model priced on OpenRouter at $1.50 per million input tokens, $7.50 per million output tokens and $30 per million image output tokens.</li><li><strong>Vals AI</strong> reports that a team of Claude Opus 5.5 agents found two candidate room-temperature antiferromagnetic semiconductors for computer memory, predicted by calculation and not yet tested in a lab.</li><li><strong>Meta</strong> published the Personal Agent Protocol with Walmart, Stripe, Sierra and others, an open standard meant to help websites tell legitimate user agents from malicious bots.</li><li><strong>OpenAI</strong> announced labeled image ads that appear next to image generation results for ChatGPT Free and Go users, with US testing starting later in October.</li><li><strong>OpenAI</strong> reports that GPT-6 Astra scored 55.0% on 11 contracting tasks built with Ironclad, against 41.6% for GPT-5.6 Sol, and the page carries no publication date.</li><li><strong>Anthropic</strong> expanded its Cyber Verification Program to absorb Project Glasswing, with three access tiers for vetted cyberdefenders, six days after Google gave Gemini 4 Argon to cyber defenders first.</li><li><strong>Technology Innovation Institute</strong> announced Falcon-Emirati-7B, a model built on Falcon-H1-Arabic and specialised in Emirati Arabic dialect and culture.</li><li><strong>DeepSeek</strong> is reportedly close to a $12 billion funding round backed by Tencent.</li><li><strong>Kuaishou</strong> has reportedly picked banks for a Hong Kong IPO of its Kling video unit worth more than $1 billion, according to The Information.</li><li><strong>OpenAI</strong> will reportedly start watermarking ChatGPT text in the EU.</li></ul>]]></content:encoded></item>
<item><title>David Robinson left OpenAI, then criticized its pace in The Atlantic</title><link>https://atlas.prashish.xyz/daily/2026-10-02</link><guid isPermaLink="true">https://atlas.prashish.xyz/daily/2026-10-02</guid><pubDate>Fri, 02 Oct 2026 21:00:00 +0000</pubDate><description>David Robinson left OpenAI and, the next day, published an essay in The Atlantic that criticizes the company&#x27;s pace.</description><content:encoded><![CDATA[<p>Friday. Written from the atlas logs (checked against sources 6 Oct 2026).</p>
<p class="lede">David Robinson left OpenAI and, the next day, published an essay in The Atlantic that criticizes the company&#x27;s pace.</p>
<p>Robinson argued that OpenAI&#x27;s optimism, as it sprints between launches, falls short of its responsibilities. Jacob Coxon resigned from Anthropic in September with a similar warning that labs are compromising oversight to keep pace. Public exits criticizing safety practice have recurred since Jan Leike left OpenAI in 2024, and none so far has changed a lab&#x27;s release schedule.</p>]]></content:encoded></item>
<item><title>Claude Code added mods, FLUX 3 Image added bounding-box control, and Suno launched Speech</title><link>https://atlas.prashish.xyz/daily/2026-10-01</link><guid isPermaLink="true">https://atlas.prashish.xyz/daily/2026-10-01</guid><pubDate>Thu, 01 Oct 2026 21:00:00 +0000</pubDate><description>Claude Code gained &quot;mods&quot; written as small TypeScript functions, Black Forest Labs&#x27; FLUX 3 Image gained bounding-box layout control, and Suno released Speech in beta.</description><content:encoded><![CDATA[<p>Thursday. Written from the atlas logs (checked against sources 6 Oct 2026).</p>
<p class="lede">Claude Code gained &quot;mods&quot; written as small TypeScript functions, Black Forest Labs&#x27; FLUX 3 Image gained bounding-box layout control, and Suno released Speech in beta.</p>
<p>Claude Code &quot;mods&quot; let users, or Claude itself, change how the coding agent behaves with small TypeScript functions shipped in plugins, without forking it. Black Forest Labs&#x27; FLUX 3 Image added layout control by bounding boxes and edits that leave everything outside the box numerically unchanged, which matters for agents that edit images many times in a row. Suno Speech, in beta, generates voice and music together as one track.</p>
<h2 id="also-in-the-news">Also in the news</h2>
<ul><li><strong>Microsoft</strong> released streaming transcription and faster voice models for real-time voice agents.</li></ul>]]></content:encoded></item>
<item><title>Google released Gemini 4 Argon to cyber defenders first; OpenAI disrupted a distillation campaign</title><link>https://atlas.prashish.xyz/daily/2026-09-30</link><guid isPermaLink="true">https://atlas.prashish.xyz/daily/2026-09-30</guid><pubDate>Wed, 30 Sep 2026 21:00:00 +0000</pubDate><description>Google released Gemini 4 Argon first to vetted cyber defenders, the third lab to stage a top model&#x27;s release around cyber risk, and OpenAI said it disrupted a campaign to copy its models&#x27; reasoning.</description><content:encoded><![CDATA[<p>Wednesday. Written from the atlas logs (checked against sources 6 Oct 2026); numbers are company-reported unless noted.</p>
<p class="lede">Google released Gemini 4 Argon first to vetted cyber defenders, the third lab to stage a top model&#x27;s release around cyber risk, and OpenAI said it disrupted a campaign to copy its models&#x27; reasoning.</p>
<p>Gemini 4 Argon has a one-million-token output limit and a reported 77.9% on DeepSWE v1.1. Its first users are vetted cyber defenders, under the US voluntary pre-release access process. Developers, enterprises and paying consumers come later. Anthropic staged Mythos this way in April, and OpenAI did the same with GPT-6 Astra in September.</p>
<p>OpenAI said it disrupted a campaign to extract protected reasoning from its models. The campaign involved about 16,000 requests from more than 4,000 users in a two-day spike in July, and OpenAI attributed a core cluster to people associated with Moonshot AI. Read with yesterday&#x27;s GLM-5.3 study, the day shows labs gating their strongest models while others close the gap fast.</p>
<h2 id="also-in-the-news">Also in the news</h2>
<ul><li><strong>SynthID Bio</strong> from Google DeepMind is a proof of concept for watermarking AI-designed proteins in the sequence itself.</li><li><strong>Cohere Embed 5</strong> (Pro and Fast) produces multimodal embeddings with a 128K context.</li></ul>]]></content:encoded></item>
<item><title>OpenAI launched dots assistants, and Anthropic found open GLM-5.3 near Mythos on exploits</title><link>https://atlas.prashish.xyz/daily/2026-09-29</link><guid isPermaLink="true">https://atlas.prashish.xyz/daily/2026-09-29</guid><pubDate>Tue, 29 Sep 2026 21:00:00 +0000</pubDate><description>OpenAI&#x27;s DevDay introduced &quot;dots&quot;, persistent assistants with their own cloud computer, and GPT-6.1 Sol claimed near-Astra results at a fifth of the price. Anthropic reported that Zhipu&#x27;s open-weights GLM-5.3 can now build working exploits.</description><content:encoded><![CDATA[<p>Tuesday. Written from the atlas logs (checked against sources 6 Oct 2026); numbers are company-reported unless noted.</p>
<p class="lede">OpenAI&#x27;s DevDay introduced &quot;dots&quot;, persistent assistants with their own cloud computer, and GPT-6.1 Sol claimed near-Astra results at a fifth of the price. Anthropic reported that Zhipu&#x27;s open-weights GLM-5.3 can now build working exploits.</p>
<p>At DevDay OpenAI launched dots, assistants that each have a name, an identity in Slack and their own cloud browser and computer. They keep working on projects and can use your laptop with permission, and they run on GPT-6 Astra. The same day GPT-6.1 Sol claimed Astra-level results on the DeepSWE coding benchmark at about one-fifth the cost ($2/$10).</p>
<p>Anthropic published an evaluation of Zhipu&#x27;s open-weights GLM-5.3. The model achieved full control-flow hijacks in 4% of trials, against 6% for Anthropic&#x27;s own gated Mythos Preview, and its safeguards were bypassed 64 to 100% of the time. Five months after Anthropic gated Mythos for cyber reasons, a downloadable model is in the same range.</p>
<h2 id="also-in-the-news">Also in the news</h2>
<ul><li><strong>DeepSeek</strong> released versions of its core MoE communication and FP8 matrix libraries for Huawei&#x27;s Ascend chips.</li><li><strong>Microsoft Quine</strong> is a biology research system that proposes and ranks interventions before wet-lab tests, and it is limited to a fellows programme.</li><li><strong>Bolt</strong> made its first acquisition, the agent startup Dokai.</li></ul>]]></content:encoded></item>
<item><title>Anthropic released Claude Sonnet 5.5, and AMD agreed to buy World Labs</title><link>https://atlas.prashish.xyz/daily/2026-09-28</link><guid isPermaLink="true">https://atlas.prashish.xyz/daily/2026-09-28</guid><pubDate>Mon, 28 Sep 2026 21:00:00 +0000</pubDate><description>Anthropic released Claude Sonnet 5.5 at $2/$10 per million tokens, scoring 70.6% on Terminal-Bench 4.0, and AMD agreed to buy Fei-Fei Li&#x27;s World Labs in a deal reported at about $8.2 billion in stock.</description><content:encoded><![CDATA[<p>Monday. Written from the atlas logs (checked against sources 6 Oct 2026); numbers are company-reported unless noted.</p>
<p class="lede">Anthropic released Claude Sonnet 5.5 at $2/$10 per million tokens, scoring 70.6% on Terminal-Bench 4.0, and AMD agreed to buy Fei-Fei Li&#x27;s World Labs in a deal reported at about $8.2 billion in stock.</p>
<p>Claude Sonnet 5.5 kept Sonnet&#x27;s $2/$10 per-million-token price and reported 70.6% on Terminal-Bench 4.0, against 10.3% for Sonnet 5 on the same test, with output about 30% faster. It arrived six days after Claude Opus 5.5, which fits the usual 2026 pattern. The gated and flagship tiers move first, and the mid tier inherits their training within weeks.</p>
<p>AMD&#x27;s agreement to acquire World Labs, reported at about $8.2 billion in stock, makes Fei-Fei Li AMD&#x27;s chief scientist. AMD is betting that persistent 3D worlds, used to train robots and to build games and simulations, will be a large compute workload that a chip maker wants to own.</p>
<h2 id="also-in-the-news">Also in the news</h2>
<ul><li><strong>Eleven v4</strong> from ElevenLabs is its most expressive speech model, with inline delivery tags, multi-speaker dialogue, 90+ languages and a roughly 150 ms Turbo variant.</li><li><strong>Kling 4.0</strong> (early access) from Kuaishou supports native 30-second clips, up to ten keyframes and fifteen reference inputs.</li></ul>]]></content:encoded></item>
</channel></rss>