Alibaba (Qwen) Qwen-Image-2.1-Turbo
8-step accelerated checkpoint of Qwen-Image-2.1 (same 7B architecture); Pro and Turbo APIs live on Alibaba Cloud Model Studio. Qwen Research License.
Frontier LLMs, reasoning models, open weights, and the image, video and audio models that matter — from 45 labs across the US, China, Europe, Asia-Pacific and the Middle East. Mirrors startrise.io/models. Updated twice a day.
The 10 most recent releases across every provider we track.
8-step accelerated checkpoint of Qwen-Image-2.1 (same 7B architecture); Pro and Turbo APIs live on Alibaba Cloud Model Studio. Qwen Research License.
Streaming Arabic speech-to-text for MSA and dialects; fine-tuned from Voxtral-Mini-4B-Realtime-2602. Apache 2.0.
12B MoE (2.5B active) Apache-2.0 coding-agent model with RL in real repos; HF JetBrains/Mellum2.1-12B-A2.5B-Thinking.
Next-gen Vidu video model with native audio, reference voices, up to 15 reference images and 4K 10-bit output.
MIT-licensed late-interaction multimodal retrieval models for text, images and pages, sharing one embedding space across sizes.
Open edge decision models built on LFM2.5 backbones; d1-3B takes text and images, experimental d1-omni-600M adds audio.
Anthropic’s cheapest, fastest small model; first Haiku with an adjustable effort setting, about 75% cheaper to run than Haiku 4.5.
7B Emirati-dialect LLM on Falcon-H1-Arabic (Falcon Chat), plus a 1.6B multilingual ASR model and a 270M Arabic OCR model.
1T-parameter natively multimodal model (49B active), Mistral's largest yet; public preview API on Mistral Studio, weights promised by end of October.
Open-sourced 3B Lite and 29B Pro text/image-to-audio-video diffusion models: 5s clips with synced 44 kHz audio and lip-sync, upscaling to 1080p. MIT.
Providers ordered by their most recent release. Filter by region, or jump straight to a lab.
15 releases · latest Oct 9, 2026
8-step accelerated checkpoint of Qwen-Image-2.1 (same 7B architecture); Pro and Turbo APIs live on Alibaba Cloud Model Studio. Qwen Research License.
Unified text-to-image and editing model with a 7B generation component and native transparent (RGBA) output. Qwen Research License.
Open-weight 125B multimodal MoE (6B active, plus 51B n-gram embeddings) previewing the Qwen4 architecture; served as Qwen3.8-Flash on Qwen Cloud.
Open weights followed on August 12.
1 release · latest Oct 8, 2026
12B MoE (2.5B active) Apache-2.0 coding-agent model with RL in real repos; HF JetBrains/Mellum2.1-12B-A2.5B-Thinking.
20 releases · latest Oct 8, 2026
Streaming Arabic speech-to-text for MSA and dialects; fine-tuned from Voxtral-Mini-4B-Realtime-2602. Apache 2.0.
1T-parameter natively multimodal model (49B active), Mistral's largest yet; public preview API on Mistral Studio, weights promised by end of October.
Open weights under a Modified MIT license.
Unifies instruct, reasoning and coding.
Apache 2.0.
First Mistral audio models.
First Mistral reasoning models (Small is open-weight).
Arabic and South Asian languages.
17 releases · latest Oct 7, 2026
Anthropic’s cheapest, fastest small model; first Haiku with an adjustable effort setting, about 75% cheaper to run than Haiku 4.5.
$5 / $25 per million input / output tokens.
First hybrid reasoning Claude.
1 release · latest Oct 7, 2026
Open edge decision models built on LFM2.5 backbones; d1-3B takes text and images, experimental d1-omni-600M adds audio.
1 release · latest Oct 7, 2026
MIT-licensed late-interaction multimodal retrieval models for text, images and pages, sharing one embedding space across sizes.
1 release · latest Oct 7, 2026
Next-gen Vidu video model with native audio, reference voices, up to 15 reference images and 4K 10-bit output.
6 releases · latest Oct 6, 2026
3.35B open research model that reasons in the prompt language across 60 languages; CC-BY-NC with Cohere Labs AUP.
Multimodal, 100+ language embedding family sharing one embedding space; Pro for quality, Fast for latency.
218B total / 25B active MoE, Apache 2.0.
27 releases · latest Oct 6, 2026
GA update to Nano Banana 2 (gemini-nano-banana-2.1): better visual quality, prompt adherence, text rendering and new panoramic aspect ratios up to 4K.
740M Apache-2.0 on-device embedder built on Gemma 4; maps text, code, images, audio and video into one space, 8K-token context.
Initially limited to trusted cyber defenders.
Full-length song generation.
Conversational video generation and editing.
Moved to Apache 2.0.
Music generation.
Preview.
Announced at Google I/O 2025.
On-device preview.
1 release · latest Oct 6, 2026
Open-sourced 3B Lite and 29B Pro text/image-to-audio-video diffusion models: 5s clips with synced 44 kHz audio and lip-sync, upscaling to 1080p. MIT.
4 releases · latest Oct 6, 2026
7B Emirati-dialect LLM on Falcon-H1-Arabic (Falcon Chat), plus a 1.6B multilingual ASR model and a 270M Arabic OCR model.
Hybrid Transformer-Mamba family.
5 releases · latest Oct 5, 2026
Speech-to-speech model for real-time voice agents with improved reasoning, instruction following and tool calling; GA on Amazon Bedrock.
1 release · latest Oct 5, 2026
501B-total / 23B-active MoE for coding, reasoning and agents; early-access preview, Apache 2.0 weights promised later in October.
3 releases · latest Oct 5, 2026
1 release · latest Oct 3, 2026
English-German MoE, Apache 2.0.
5 releases · latest Oct 2, 2026
Image side of FLUX 3: text-to-image, pixel-exact local edits, up to 10 references, bounding-box layout, up to 4K.
Open-weight 7B world-action (robotics) model.
Unified video, image and audio model.
1 release · latest Oct 1, 2026
Cloudflare's first trained models: 27B and 9B decision models (Qwen backbones) returning typed answer probabilities; Jev-API compatible, Apache 2.0, on Workers AI.
8 releases · latest Oct 1, 2026
Public preview in Microsoft Foundry.
Small-tier coding model in GitHub Copilot; adds native vision over MAI-Code-1-Flash at a 73% lower list price.
Seven-model MAI family announced at Build 2026.
First in-house MAI models.
1 release · latest Sep 30, 2026
Ant InclusionAI ~560B MoE (~25B active) hybrid-reasoning model; API exposes 256K now, up to 1M planned; weights promised soon.
21 releases · latest Sep 29, 2026
Initially limited to select organizations.
Limited preview from June 26, GA July 9.
API availability followed on April 24.
Speech-to-speech model.
First open-weight OpenAI LLMs since GPT-2.
Research preview.
1 release · latest Sep 22, 2026
35B MoE (3B active) cost-efficient agent model with Korean/English/Japanese, reasoning mode, and 512K context.
3 releases · latest Sep 22, 2026
Native omni-modal MoE; Pro is 1.02T total / 42B active. MIT-licensed weights, plus a Pro-UltraSpeed API variant.
Previewed on OpenRouter as "Hunter Alpha".
14 releases · latest Sep 21, 2026
Speech-to-text model.
Speech-to-speech voice model.
SpaceXAI's own announcement post is dated July 16.
4 releases · latest Sep 20, 2026
600B MoE (27B active) flagship for agentic coding and knowledge work with vision and 1M context; open weights planned Oct 15.
2 releases · latest Sep 11, 2026
1 release · latest Sep 11, 2026
744B MoE agentic model built on GLM-5.2, MIT license. Technical report followed on Sep 14.
12 releases · latest Sep 10, 2026
Natively multimodal; first model in a new architecture family.
Experimental vision model; weights followed on Hugging Face.
Public beta of the re-post-trained V4-Flash.
Introduced DeepSeek Sparse Attention.
Hybrid thinking / non-thinking model.
6 releases · latest Sep 2, 2026
Improved agentic and coding vs 1.2; max reasoning on Muse Code and Meta Model API.
30B open-weight model for local use.
Released alongside the Muse Code agent.
First model from Meta Superintelligence Labs.
5 releases · latest Aug 28, 2026
3 releases · latest Aug 25, 2026
7 releases · latest Aug 14, 2026
Open weights followed on August 28.
4 releases · latest Jul 31, 2026
11 releases · latest Jul 31, 2026
Open multimodal video model.
1080p, up to 10-second clips.
6 releases · latest Jul 16, 2026
Full weights released July 27 under the Kimi K3 License.
6 releases · latest Jul 8, 2026
Unified audio-video generation.
2 releases · latest Jun 30, 2026
2 releases · latest Jun 8, 2026
3 releases · latest Jun 4, 2026
3 releases · latest May 20, 2026
4 releases · latest May 9, 2026
2 releases · latest Feb 18, 2026
1 release · latest Feb 11, 2026
Trained entirely on domestic Chinese compute.
3 releases · latest Jan 8, 2026
0 releases · latest —
No notable model releases tracked since January 2025.