Alibaba International Digital Commerce Group (AIDC-AI) has open-sourced its latest multimodal large model, Ovis2.6-80B-A3B. The model upgrades the language model backbone to a mixture-of-experts (MoE) architecture, bringing total parameters to 80 billion — but only activating about 3 billion parameters per inference.
The biggest breakthrough is the introduction of a “Think with Image” mechanism. Previous multimodal models passively take in a full image and generate an answer in one go. But Ovis2.6, while generating a chain of thought, can actively call built-in visual tools like cropping and rotation — zooming in on image regions and repeatedly cross-checking, much like a human. This self-reflective multi-round reasoning dramatically improves accuracy on complex visual tasks.
On the performance side, Ovis2.6 expands its context window to 64K tokens and natively supports input images up to 2880×2880 resolution. Combined with enhanced optical character recognition (OCR) and chart analysis, the system can gather clues across multi-page documents and handle detail-heavy long-form question answering.
With 80 billion parameters ensuring high cognitive capacity and just 3 billion activated per inference controlling costs, Ovis2.6 offers a cost-effective formula: let large models handle massive financial statements, lengthy research reports, and other information-intensive tasks — seeing clearly while catching details, without the exploding compute bill.