← News

Cohere releases North Small Translate open-weight MT model

10 Sep 2026: Cohere published North Small Translate — a mixture-of-experts machine translation model (218B total / 25B active parameters) for 50+ languages, open weights under CC BY-NC 4.0. Company WMT26 and throughput figures stay Cohere’s claims.

10 Sep 2026: Cohere published “Introducing North Small Translate: A leading sovereign open-weight machine translation model.” That dated company blog is the filing event.

Model on the page: North-Small-Translate-1.0 — mixture-of-experts architecture, 218B total / 25B active parameters, 16k input and 16k output context, text-in / text-out. Cohere says it supports 50+ languages; the full list is on Cohere’s page, not reprinted here.

License on the post: open weights for research and non-commercial use under CC BY-NC 4.0. Hugging Face weights — including near-lossless quantizations — and a Hugging Face Space demo are linked from the post. This is not unrestricted commercial open source. Cohere points enterprises that need a commercial path to RWS Language Weaver.

Cohere’s WMT26 “All Languages” score claim: 83.6 for North Small Translate versus listed comparators on the page (DeepL NextGen 81.37, Google Translate 68.20, Gemma 4 31B (on) 79.46, GLM 5.2 FP8 76.50, Qwen 3.5 397B A17B 81.56). An “Agentic” variant that finds and fixes translation errors is listed at 84.36. Those scores are Cohere’s published evaluations; the page footnotes GPT-5.6-Sol as judge. This desk did not rerun the benches.

Throughput claim, company internal tests: up to 1.4× higher output throughput than Gemma 4 31B TP1 under identical concurrency and hardware in Cohere’s chart — 112 vs 81 TOPS at low concurrency; 39 vs 30 at high. Attribute as Cohere’s test, not an independent throughput audit.

Long-context claim on the same company table: 48.9 on Cohere’s book-chapter eval versus Google Translate 21.3 and Gemma 4 31B 19.4. That is Cohere’s long-context framing — not an external audit.

Hardware minimums stated on the page — for example 1× B200 @ W4A4 or 2× H100s @ W4A4 — are Cohere’s guidance, not a desk hardware check.

Partnership: Cohere says the model was developed with RWS / Language Weaver research, science, and language-expert teams. Commercial enterprises are pointed to Language Weaver for secured deployment beyond the non-commercial Hugging Face weights. This filing does not name customer logos or a deal size.

A dated Cohere primary puts a dedicated open-weight (non-commercial) MoE translation model into researchers’ hands with published WMT26 claims — readers get a sovereign-MT product beat without treating company judge scores as settled science.

ONLINE

article thread

guidelines

warming…

warming…

Sources