
AI agent infrastructure company Composio ran the same Kimi K3 model through three different coding frameworks — Kimi Code, Hermes, and Claude Code — on 28 identical tasks. The success rates were nearly the same, with a gap of just two completed tasks. But the token usage and cost? Wildly different.
Kimi Code completed 22 tasks, Hermes 21, and Claude Code 20. Median token usage per task: 61,000 for Kimi Code, 67,000 for Hermes, and 340,000 for Claude Code. That translates to an average cost per task of about $0.22, $0.28, and $2 respectively. Some individual tasks saw token usage differ by as much as 30x.
Hermes was the fastest with a median task time of 179 seconds, compared to 297 seconds for Kimi Code and 348 seconds for Claude Code.
A separate study from Writer found similar results. Researchers ran 22 enterprise tasks across 6 models, swapping only the framework. Token usage dropped by 38%, per-task cost by 41%, and time by 44%, while quality remained essentially the same.
https://twitter.com/composio/status/2082452269522378858?s=46