ERNIE-4.5-VL-28B-A3B-Thinking
ERNIE-4.5-VL-28B-A3B-Thinking augments Baidu's 28B vision-language model with explicit reasoning for visual-language tasks. It targets multi-step reasoning over images, chart analysis, causal reasoning from visual content, and video understanding with temporal awareness and event localization. Despite activating only 3 billion parameters, it aims to compete with larger models on document and visual reasoning workloads, with a 131K token context for long document analysis. Key features include visual grounding with precise instruction execution, tool-calling for image search, and thinking modes for harder visual reasoning. It is trained with multimodal reinforcement learning using GSPO and IcePop strategies and dynamic difficulty sampling, released as open weights under the Apache 2.0 license with support for multiple inference frameworks (FastDeploy, vLLM, Transformers).
Key info
Available routes
No routes currently available — ERNIE-4.5-VL-28B-A3B-Thinking isn't routed through the Opper gateway right now. It may return.
Contact us about this model →