Command A
Command A is a 111B enterprise model with 256K context that runs on two GPUs and, Cohere claims, matches GPT-4o and DeepSeek-V3 on agentic tasks.
Optimized for tool use, RAG with citations and agents across 23 languages. Uses sliding-window plus global attention layers. Technical report describes decentralized training with self-refinement and model merging. Weights are CC-BY-NC, so non-commercial.
- Date
- Thursday, 13 March 2025
- Lab
- Cohere
- Kind
- open-weights
- Access
- open weights (restricted license)
Figures
| Measure | Value | Measured by |
|---|---|---|
| Parameters / context | 111B / 256K tokens deployable on two GPUs per model card | company |
'On par or better than GPT-4o and DeepSeek-V3' is Cohere's own claim. Technical report arXiv 2504.00698, v1 2025-04-01 (55 pages).
Sources
This record was checked against its sources on 6 October 2026. How we check