AI Research Atlas

The week in AI, 5 to 11 October 2026

Weekly brief · 5 to 11 October 2026

This is the current week. The page is revised every day until Sunday.

Week of 5 to 11 October 2026, with 5 and 6 October covered so far. Mistral previewed the 1 trillion parameter Mistral Large 4, Reflection AI unveiled its first model Beam, Anthropic folded Project Glasswing into an expanded Cyber Verification Program, Google DeepMind released the open EmbeddingGemma 2, and Google cut free Gemini users to Flash Lite.

The week in brief

Mistral AI opened a preview API for Mistral Large 4 on 6 October, a 1 trillion parameter multimodal model that it says will get open weights by the end of October.

The first two days of the week brought two large open-weight models from Western labs, and neither can be downloaded yet. Mistral AI previewed Mistral Large 4 on 6 October, and Reflection AI unveiled Beam, its first model, on 5 October. Both companies say weights follow later in October. All of their performance claims are company-reported, and no independent benchmark results exist for either model so far.

Anthropic widened access to its cyber-capable models on 6 October, six days after Google gave Gemini 4 Argon to cyber defenders first. Google released EmbeddingGemma 2, an open embedding model that handles text, code, images, video and audio in one space, and said free Gemini users will get only Flash Lite from 9 October. OpenAI added labeled image ads to ChatGPT, and Meta published an open protocol for shopping agents.

Mistral previewed the 1 trillion parameter Mistral Large 4

Mistral AI released a preview API for Mistral Large 4 on 6 October, a 1 trillion parameter model with 49 billion parameters active per token, and says open weights will follow by the end of October.

What happened

Mistral Large 4 is Mistral's largest model so far, up from the 675 billion parameters of Mistral Large 3. It is multimodal, and it runs as a preview API on Mistral Studio. Mistral says it trained the model from scratch on 3,800 NVIDIA Grace Blackwell GPUs in its own European datacenters.

Large 4 is a mixture-of-experts model. Each token is routed through a small set of expert subnetworks, so only part of the model does work at any one time, and that keeps inference cost closer to a much smaller model. Mistral's blog gives 1 trillion total and 49 billion active parameters. Mistral's documentation lists 1.05 trillion total and 52 billion active, and the company hasn't explained the gap.

What Mistral claims

Mistral says Large 4 significantly outperforms any open-weight model from the US or Europe and is competitive with the strongest open models globally. That second phrase points at the Chinese open-weight models, which Mistral does not claim to beat. The atlas has no independent benchmark result for Large 4, and Mistral did not publish scores the atlas could check on 6 October.

Why it matters

Until the weights ship, Mistral is red-teaming a version with reduced moderation together with cybersecurity partners. That follows directly from last week's records. On 29 September Anthropic reported that Zhipu's open-weights GLM-5.3 built working exploits about two-thirds as often as its own gated Mythos Preview, and an open-weight model can't be recalled once released.

Mistral's approach is a short closed window before an open release. Anthropic, OpenAI and Google gate their top models for longer and keep the weights. A lab releasing 1 trillion parameters of open weights has to decide its cyber position before release, and Mistral's red-teaming period is how it is doing that.

What we don't know

Which of Mistral's two parameter counts is correct is unknown. The scores behind "competitive with the strongest open models globally" weren't available to the atlas. Whether the released weights will be the moderated version or the one being red-teamed with reduced moderation is also unknown.

Reflection AI unveiled Beam, a 501 billion parameter model

Reflection AI unveiled Beam on 5 October, a text-only mixture-of-experts model with 501 billion total and 23 billion active parameters, and says it matches Z.ai's GLM-5.2 on reasoning benchmarks with 3 to 4 times less inference compute.

What happened

Beam is Reflection AI's first model release. It is text-only, aimed at coding and agent tasks, and has a 1 million token context window. Reflection says it pretrained Beam on 23.8 trillion tokens and then trained it heavily with reinforcement learning (RL), where the model improves by being scored on tasks it attempts.

Beam is available as a research preview. Reflection says the weights and a technical report will come later in October.

The comparison with GLM-5.2

Reflection chose a Chinese open model as its benchmark. GLM-5.2 is about 744 billion total and 40 billion active parameters, so Beam is smaller on both counts. Reflection says Beam matches GLM-5.2 on reasoning benchmarks while using 3 to 4 times less inference compute, and nobody has checked that claim independently.

Active parameters drive most of the per-token cost in a mixture-of-experts model. Beam has a little over half of GLM-5.2's active parameters, which doesn't by itself account for a 3 to 4 times saving. The technical report will have to show where the rest comes from.

Why it matters

Mistral and Reflection both announced large open-weight models in the same two days, and both compared themselves against open models from China. The GLM line is the same family Anthropic tested for exploit ability on 29 September with GLM-5.3. For a US practitioner choosing an open model, the options at this size have mostly come from Chinese labs, and Beam and Large 4 are the US and European entries for October.

What we don't know

The benchmark scores behind "matches GLM-5.2" aren't in the records. The license terms for the weights haven't been announced. It's also unknown how Beam compares with GLM-5.3, the newer model in the same family.

Anthropic folded Project Glasswing into its Cyber Verification Program

Anthropic expanded its Cyber Verification Program on 6 October and folded Project Glasswing into it, with three access tiers for vetted cyber defenders using its most capable Claude models.

What happened

Project Glasswing began in April as the partner program through which Anthropic gave Claude Mythos Preview to a small set of organisations. On 6 October Anthropic merged it into its Cyber Verification Program, which now has three access tiers for vetted security professionals. The tier details weren't in the sources the atlas could read.

Reuters, as summarized by Techmeme, reported Anthropic's figures. Anthropic says it found more than 5,500 verified vulnerabilities between April and October. It says Glasswing partners found more than 129,000 vulnerabilities between April and July, of which more than 33,000 were critical or high severity. These are company-reported counts.

Why it matters

The expansion came six days after Google DeepMind released Gemini 4 Argon to vetted cyber defenders first on 30 September. Anthropic, OpenAI and Google all now put their most cyber-capable models in front of defenders before the public. Anthropic's change turns a small partner list into a program with entry levels, so more security teams can apply.

The partner count matters because of the patching numbers from earlier in the year. Anthropic's May Glasswing update reported that only 75 of 530 disclosed high or critical open-source bugs had been patched. Finding vulnerabilities was already faster than fixing them in May, and a 129,000 count makes that gap more important.

What is reported and unconfirmed

The atlas's record says vetted users get Anthropic's top models with fewer safeguards for cyber work. That part didn't appear on any evidence page the atlas checked, so the atlas treats it as reported and unconfirmed. The atlas also has no patch rate to go with the April to July vulnerability count.

Google DeepMind released EmbeddingGemma 2 for five input types

Google DeepMind released EmbeddingGemma 2 with open weights on 6 October, a 740 million parameter model that maps text, code, images, video and audio into one 768-dimensional space.

What happened

An embedding model turns an input into a list of numbers, a vector, so that similar inputs land close together. Search, retrieval for chatbots and recommendation systems all run on these vectors. The first EmbeddingGemma handled text only, and EmbeddingGemma 2 is built on Gemma 4 and adds code, images, video and audio in the same shared space.

The encoders are modular. A developer can load only the text and code part at 270 million parameters, text with images and video at 440 million, text with audio at 570 million, or everything at 740 million. Google reports about 191MB of active RAM for the text-only weights and about 567MB for the full model on a Pixel 11 Pro.

Why it matters

A single shared space removes a chain of models. Until now, searching a video library by text meant running captioning or speech-to-text first and then embedding the text output. With one space, a text query and a video clip can be compared directly, and Google's RAM figures put that on a phone.

The release is small next to the trillion-parameter announcements, and it's the one item this week a developer can download and use today. It's open weights and its claims are verified against Google's own pages.

What we don't know

The atlas has no independent retrieval benchmark for EmbeddingGemma 2. How much quality the 270 million text-only version gives up against the full 740 million model wasn't in the records.

Google cut free Gemini users to Flash Lite from 9 October

From 9 October free Gemini users will get only Flash Lite, standard Flash will need the $4.99 per month Google AI Plus plan, and AI Plus will lose Gemini Pro.

What changes

Free Gemini users can currently pick Flash Lite, Flash or Pro. From 9 October they get Flash Lite only. Standard Flash moves to Google AI Plus at $4.99 per month, and AI Plus loses Gemini Pro.

Pro and the Deep Think reasoning option will be limited to Google AI Pro at $19.99 per month and Google AI Ultra at $99.99 per month. Low, medium and high effort levels stay available for each model, and The Verge reports that higher effort may use up the usage limit faster.

Why it matters

Google has pulled the free tier down by two models. Anyone demonstrating Gemini on a free account after 9 October will be showing Flash Lite, and comparisons people make between free ChatGPT and free Gemini will be comparing different tiers than they were the week before.

OpenAI made a different move with its free users on 5 October and announced image ads in ChatGPT (see "Also worth knowing"). Google is moving free users to a cheaper model, and OpenAI is adding ads to its free and Go tiers.

What we don't know

Google hasn't given a reason in the records the atlas holds. Usage limits for each tier after 9 October weren't published in the source.

Also worth knowing

Agents and commerce

  • Personal Agent Protocol is an open standard published on 6 October by Meta with Walmart, Stripe, Sierra and others, meant to help websites tell bots acting for a user from malicious ones. Amazon has begun deliberately blocking Meta's Muse agent from its retail site, and standard anti-bot checks often block other personal agents too. Meta presented Muse as a shopping-capable agent on 23 September.
  • Ironclad worked with OpenAI to turn 11 contracting tasks into research problems graded on 8 to 50 criteria each, and OpenAI reports GPT-6 Astra scored 55.0% against 41.6% for GPT-5.6 Sol. OpenAI says Astra's average time per attempt fell from 37.0 minutes to 19.2 minutes, and an unreleased internal model scored 63.7%. The OpenAI page carries no publication date, so the atlas logged it on 6 October when it appeared.

Products and pricing

  • ChatGPT image ads were announced by OpenAI on 5 October, with labeled ads shown next to image generation results for Free and Go users and US testing starting later in October. OpenAI says ChatGPT has 1.2 billion weekly users, and DV Rockerbox reports WeightWatchers saw a 15.3% lower cost per acquisition than its blended paid-search benchmark.
  • Nano Banana 2.1 is Google DeepMind's new Flash-tier image generation and editing model, listed on OpenRouter at $1.50 per million input tokens, $7.50 per million output tokens and $30 per million image output tokens. OpenRouter's listing says it improves mask and ink editing and product recontextualization, with output up to 4K.
  • ChatGPT text in the EU will reportedly carry watermarks from OpenAI, and the atlas has no confirming record or start date.

Science

  • Vals AI reports that a team of Claude Opus 5.5 agents found two candidate room-temperature antiferromagnetic semiconductors for computer memory. The candidates are predicted by calculation and haven't been tested in a lab.

Regional models

  • Falcon-Emirati-7B was announced by the Technology Innovation Institute, a model built on Falcon-H1-Arabic and specialised in the Emirati Arabic dialect and culture.

Money

  • DeepSeek is reportedly close to a $12 billion funding round backed by Tencent. Last week DeepSeek released versions of its MoE communication and FP8 GEMM libraries for Huawei's Ascend chips.
  • Kling, Kuaishou's video unit, has reportedly picked banks for a Hong Kong IPO worth more than $1 billion, according to The Information. Kling 4.0 entered early access on 28 September.

People this week

The atlas logged no moves of people on 5 or 6 October.

What did not change

Every capability number in this week's stories is company-reported. No independent evaluator has published results for Mistral Large 4, Beam or EmbeddingGemma 2, and OpenAI's Ironclad figures come from an undated OpenAI page.

The most cyber-capable models from Anthropic, OpenAI and Google still go to vetted defenders before anyone else. Anthropic's 6 October expansion adds tiers and more applicants to that arrangement.

Neither of this week's large open-weight models can be downloaded yet. Mistral and Reflection both give October dates, and until then the strongest downloadable models at this size are still the Chinese ones they compare against.

What to watch next

  • On 9 October Google's Gemini free tier drops to Flash Lite only, and AI Plus loses Gemini Pro.
  • Later in October Reflection AI says it will release Beam's weights and a technical report, which should show the benchmark scores behind its GLM-5.2 comparison.
  • By the end of October Mistral says Large 4's weights will ship, and the release will show which moderation settings the public version carries.
  • Later in October OpenAI begins US testing of image ads in ChatGPT for Free and Go users.
  • Anthropic has given no date for publishing the details of its three Cyber Verification Program tiers.
  • DeepSeek's reported $12 billion round has no announced closing date.