Jina AI has open-sourced jina-embeddings-v5-omni, a four-modal embedding model that handles text, images, audio, and video. The breakthrough isn't just the multimodality — it's how cheaply it's achieved. By freezing the existing text backbone and only training the connection layers between the vision/audio encoders and the model, just 0.35% of total parameters are updated.
That means any company already using Jina's text-only v5 embeddings doesn't have to rebuild their text index. The same text input produces the same vector in both v5-text and v5-omni. You just add separate indexes for images, audio, and video to unlock four-modal search on top of your existing system.
The training savings are dramatic: up to 64% less GPU memory and 3.9x faster training. Benchmarks show the 1.57B-parameter Small version matches LCO-Embedding-Omni-7B, a model nearly six times its size. Video retrieval still lags, but the underlying logic is compelling: if your text foundation is strong enough, adding extra senses can be nearly free.