Home / Models / ERNIE 5.0

ERNIE 5.0

Maker: Baidu · language

Baidu's natively omni-modal 2.4-trillion-parameter model: a single framework both understands and generates text, images, video and audio, alongside language tasks such as reasoning, coding and agentic tool use.

Understands

text, images, audio, video

Produces

text, images, video, speech

Contents verified with the vendor: 2026-08-18

Included in subscriptions

No catalogue entry marks this model as part of a plan yet. It will appear here once the plan contents are confirmed with the vendor.