Kimi long-context setup
Kimi-1M handles huge codebases / long contracts / research docs alongside GLM.
Configure KimiHermes Agent China-region LLM provider guide
Zhipu GLM-4 is the fastest LLM option for Hermes Agent users in mainland China: 3-step setup, 18 RMB free trial credit, prices an order of magnitude below GPT-4, side-by-side comparison with Claude/GPT, no proxy needed. Synced with v2026.9.14 and the GLM-4-Plus / Air / Flash tier lineup.
Zhipu AI (Tsinghua University spin-off) runs one of the most stable LLM API platforms in mainland China. GLM-4 is its flagship closed-source model. For Hermes Agent users in China, it solves three pain points at once:
Short version: if your Hermes Agent runs primarily inside China and you care about cost + network stability, GLM-4 is the most cost-effective default today.
You need a working API key before configuring GLM. The flow takes about 5 minutes.
1. Sign up at Zhipu Open Platform
Go to https://open.bigmodel.cn โ sign up with a phone number โ complete real-name verification (individual or business). Verification is mandatory โ without it you cannot claim the trial credit.
2. Create and copy an API key
After login: Console โ API Management โ Create API Key. Copy the key (format: abcdef123456.xyz). The key is shown once only โ save it to 1Password / your system keychain immediately.
3. Verify Hermes โฅ v0.9.5
Native GLM provider support was added in v0.9.5. Run hermes --version. If lower, run hermes update first (see /install).
Ready to go. The next section wires GLM into Hermes in 3 commands.
The config file lives at ~/.hermes/config.toml, but you do not need to edit it by hand โ everything goes through the CLI.
hermes config set provider zhipuSwitches the default LLM provider to Zhipu. Hermes ships with a built-in zhipu adapter that uses the official OpenAI-compatible endpoint.
1. Switch provider
The command above flips ~/.hermes/config.toml [llm].provider to zhipu and seeds an empty [llm.zhipu] section.
2. Write the API key
Run hermes config set provider.zhipu.api_key "YOUR_API_KEY" (mind the double quotes โ the key may contain dots). It lands encrypted in ~/.hermes/secrets.toml (mode 0600), so it will not be committed to git by accident.
3. Pick a model + verify
hermes config set provider.zhipu.model glm-4-plus (recommended: top-tier capability). Then run hermes doctor --provider โ the script sends a ping request and you should see โ within 5 seconds.
Done. Start the agent with hermes and any task will now hit GLM. Next section: how to pick between the three tiers and what it costs.
Zhipu charges per token, with input and output priced separately. Current three-tier lineup and target scenarios (2026-05 data):
These numbers shift as Zhipu adjusts pricing; refer to the official pricing page at https://open.bigmodel.cn/pricing for the latest. Hermes Agent does not handle billing โ all charges go directly to your Zhipu account.
GLM-4 vs Claude / GPT-4 / Kimi in real Hermes Agent scenarios. Numbers are from the Hermes teamโs 2026-04 measurements across 4 machines in China (50 runs per task, averaged). For sizing only.
Rule of thumb: daily + in-China โ GLM-4-Plus or Air. Need ultra-long context (>200K) โ Kimi. Heavy English code + budget allows โ Claude. Hermes supports multiple providers in parallel โ see /docs to switch per Skill.
Once GLM is wired into Hermes, four practices keep your bill low and your output quality high.
Combined with the Hermes Memory module (see /docs#memory), high-frequency context can be compressed under 5K โ long-session bills drop another ~60%.
90% of day-to-day GLM errors fall into these five buckets. Walk through and match.
401 Unauthorized / Invalid API key
ๅๅ / Cause: Key copied with whitespace, quotes, or a trailing newline; or generated under a different account; or disabled due to long inactivity.
ไฟฎๅค / Fix: Generate a fresh key in the open-platform console, run hermes config set provider.zhipu.api_key "NEW_KEY". Use ASCII double quotes, no surrounding whitespace.
429 Too Many Requests / Rate limit exceeded
ๅๅ / Cause: RPM (requests per minute) exceeded. Free tier 5 RPM hits the ceiling fast, especially in Gateway mode with concurrent tasks.
ไฟฎๅค / Fix: Short-term: add a global cap with hermes config set llm.zhipu.rpm 4 (leave 1 RPM buffer). Long-term: open a ticket in the console to raise RPM. Paid users typically get 600 RPM within 24h.
Insufficient balance
ๅๅ / Cause: Trial credit exhausted or prepaid balance hit zero.
ไฟฎๅค / Fix: Top up via the console โ Account Center โ Top up. Keep at least ~7 USD buffer to avoid service interruption mid-task. Set a low-balance email alert (< 10 RMB).
Model not found: glm-4-plus / glm-4-air
ๅๅ / Cause: Misspelled model ID, or your account lacks access (some new models are opt-in by default).
ไฟฎๅค / Fix: Check the current allowlist at https://open.bigmodel.cn/dev/api, copy the exact ID to hermes config set provider.zhipu.model <id>. GLM-4-Flash sometimes needs manual enable in the "Service activation" page.
Timeout after 60s / context too long
ๅๅ / Cause: 128K window stuffed too full causing inference timeout, or in-flight network blip.
ไฟฎๅค / Fix: Run hermes context summarize to compress the session. For non-RAG large-doc cases, enable streaming with hermes config set llm.zhipu.stream true. If timeouts persist, switch to Kimi-128k (see /providers/kimi).
Kimi-1M handles huge codebases / long contracts / research docs alongside GLM.
Configure KimiHermes Skills system: 4 registries, self-learning loop, pairing with GLM function calls.
Skills guidehermes config / skills / gateway โ 50+ commands grouped and cross-linked.
CLI referenceComplete documentation synced with v2026.9.14: install, Skills, Memory, Gateway.
Docs hub