Menu

Categories

Tags

Jina AI's V5-Omni: Four-Modal Search by Training Just 0.35% of Parameters

May 14, 2026 | Source: jina | AI | 110 views 0 comments

Jina AI has open-sourced jina-embeddings-v5-omni, a four-modal embedding model that handles text, images, audio, and video. The breakthrough isn't just the multimodality — it's how cheaply it's achieved. By freezing the existing text backbone and only training the connection layers between the vision/audio encoders and the model, just 0.35% of total parameters are updated.

That means any company already using Jina's text-only v5 embeddings doesn't have to rebuild their text index. The same text input produces the same vector in both v5-text and v5-omni. You just add separate indexes for images, audio, and video to unlock four-modal search on top of your existing system.

The training savings are dramatic: up to 64% less GPU memory and 3.9x faster training. Benchmarks show the 1.57B-parameter Small version matches LCO-Embedding-Omni-7B, a model nearly six times its size. Video retrieval still lags, but the underlying logic is compelling: if your text foundation is strong enough, adding extra senses can be nearly free.

Leave a Reply

Your email address will not be published. Required fields are marked *