Hermes Agent plugs into Hugging Face: cloud AI or fully offline
Hugging Face and open-source agent Hermes Agent are now fully integrated. Hermes, developed by Nous Research, is a local AI assistant that can work autonomously in the background and connect to messaging platforms like Telegram and Discord. With this integration, developers have two ways to run AI tasks with Hermes, and Hugging Face Hub now natively supports visualizing Hermes execution traces.
Cloud path: Hugging Face as an 'OpenRouter' Hermes lists Hugging Face's Inference Providers as a first-tier inference supplier. Developers just need to enter their HF Token, and Hermes will automatically call over 20 cloud open-source models like Qwen and DeepSeek through Hugging Face's unified routing interface (router.huggingface.co/v1). Behind the scenes, it picks the fastest among compute providers like Groq, Together, and SambaNova. Free credit is $0.10 per month, no markup.
Local path: Download models and run offline Hugging Face has added a Hermes section in the 'Local Applications' part of its documentation, and introduced a model filter tag apps=hermes-agent. Developers can directly filter compatible quantized models (GGUF format), use llama.cpp to run a local inference server, and point Hermes to localhost—no internet connection required.
Hugging Face's team says the move is to accelerate the spread of local agents: "The vast majority of agents will soon run locally, and we want to accelerate that process as much as possible."