Home / Models / Doubao-Seed-2.0-Lite

Doubao-Seed-2.0-Lite

Maker: ByteDance · language

Doubao's first omni-modal understanding model with unified handling of text, images, audio and video. It transcribes speech in 19 languages, translates, captures emotion, ambient sound and music, and reasons across audio and vision together; it also recognises GUI elements and controls a browser.

предлагается по API на платформе Volcano Ark (火山方舟) от Volcengine

Understands

text, images, audio, video

Produces

text, code

Contents verified with the vendor: 2026-08-18

Included in subscriptions

No catalogue entry marks this model as part of a plan yet. It will appear here once the plan contents are confirmed with the vendor.