Home / Models / Step 3.7 Flash

Step 3.7 Flash

Maker: StepFun · reasoning

Sparse multimodal mixture of experts: a 196B language backbone plus a 1.8B vision encoder, activating roughly 11B parameters per token. Native image and video understanding, a 256K context and about 400 tokens per second. Apache 2.0.

выход 29.05.2026; производителя StepFun в справочнике нет

Understands

text, images, video

Produces

text, code

Contents verified with the vendor: 2026-08-18

Included in subscriptions

No catalogue entry marks this model as part of a plan yet. It will appear here once the plan contents are confirmed with the vendor.