Volcengine today launched its "Agent Plan" (Agent套餐包), a subscription that deeply integrates large language models with the peripheral tools agents need — what it calls the Harness. Competitors like MiniMax have offered bundles of multimodal models before, but those are essentially just inference APIs. Volcengine's Agent Plan goes further by wrapping the agent's external retrieval system — web search and memory — into the billing loop, and it's already natively compatible with Claude Code, OpenClaw, and Hermes Agent.
At the model layer, the plan includes ByteDance's own Seed family of multimodal models (handling code, images, and video) and also pulls in third-party models like GLM-5.1 and Kimi-K2.6, with an Auto mode for intelligent routing. The real differentiation lives in the Harness tool layer: the plan directly bundles the same real-time web search that powers Doubao (ByteDance's consumer AI assistant) and Doubao-embedding-vision memory retrieval.
To smooth over the pricing differences between calling various models and external search tools, Volcengine introduced a new unified billing unit called AFP (Agent Fuel Points). Whether you're generating video or making a search query, everything burns AFP credits. By packaging models and tools together, developers building agents no longer have to separately buy search engine APIs or set up a standalone vector database. It's a clear signal that the delivery model for foundation model companies is shifting from selling a single API to supplying a complete agent runtime base.