ERNIE 5.0
Maker: Baidu · language
Baidu's natively omni-modal 2.4-trillion-parameter model: a single framework both understands and generates text, images, video and audio, alongside language tasks such as reasoning, coding and agentic tool use.
Understands
text, images, audio, video
Produces
text, images, video, speech
Contents verified with the vendor: 2026-08-18
Included in subscriptions
No catalogue entry marks this model as part of a plan yet. It will appear here once the plan contents are confirmed with the vendor.