Hermes Agent China-region LLM provider guide

Configure Zhipu GLM-4 with Hermes Agent (2026 complete guide)

Zhipu GLM-4 is the fastest LLM option for Hermes Agent users in mainland China: 3-step setup, 18 RMB free trial credit, prices an order of magnitude below GPT-4, side-by-side comparison with Claude/GPT, no proxy needed. Synced with v2026.9.14 and the GLM-4-Plus / Air / Flash tier lineup.

Hermes Agentv2026.9.14Last updated
  • โฑ 3-step setup, under 5 minutes total
  • ๐Ÿ’ฐ GLM-4-Air โ‰ˆ 0.0001 USD per 1K tokens โ€” ~50ร— cheaper than GPT-4o
  • ๐Ÿ‡จ๐Ÿ‡ณ Direct connection inside China โ€” no proxy, no VPN
  • ๐Ÿงฐ Full compatibility with bundled and community Skills, function calling, JSON mode, streaming

Why GLM-4

Zhipu AI (Tsinghua University spin-off) runs one of the most stable LLM API platforms in mainland China. GLM-4 is its flagship closed-source model. For Hermes Agent users in China, it solves three pain points at once:

  • ๐Ÿš€ Low-latency in-country access: direct connection to open.bigmodel.cn, no proxy. Beijing/Shanghai/Shenzhen real-world first-token latency < 400ms โ€” 3-5ร— faster than routing Claude through a relay.
  • ๐Ÿ’ธ An order of magnitude cheaper: GLM-4-Air input โ‰ˆ 0.0001 USD/1K tokens, GLM-4-Flash is free up to quota. Equivalent task monthly cost is typically 5-10% of GPT-4o.
  • ๐Ÿ”ง Stable function calling: GLM-4-Plus matches GPT-4o-2024-08 on Hermes Skills tool-calling accuracy, and clearly outperforms open-source Chinese alternatives.
  • โœ… Compliance-friendly: Zhipu has passed CAC large-model registration, removing a key compliance blocker for enterprise rollouts in China.

Short version: if your Hermes Agent runs primarily inside China and you care about cost + network stability, GLM-4 is the most cost-effective default today.

Prerequisites (account + API key)

You need a working API key before configuring GLM. The flow takes about 5 minutes.

  1. 1. Sign up at Zhipu Open Platform

    Go to https://open.bigmodel.cn โ†’ sign up with a phone number โ†’ complete real-name verification (individual or business). Verification is mandatory โ€” without it you cannot claim the trial credit.

  2. 2. Create and copy an API key

    After login: Console โ†’ API Management โ†’ Create API Key. Copy the key (format: abcdef123456.xyz). The key is shown once only โ€” save it to 1Password / your system keychain immediately.

  3. 3. Verify Hermes โ‰ฅ v0.9.5

    Native GLM provider support was added in v0.9.5. Run hermes --version. If lower, run hermes update first (see /install).

Ready to go. The next section wires GLM into Hermes in 3 commands.

Three-step setup

The config file lives at ~/.hermes/config.toml, but you do not need to edit it by hand โ€” everything goes through the CLI.

hermes config set provider zhipu

Switches the default LLM provider to Zhipu. Hermes ships with a built-in zhipu adapter that uses the official OpenAI-compatible endpoint.

  1. 1. Switch provider

    The command above flips ~/.hermes/config.toml [llm].provider to zhipu and seeds an empty [llm.zhipu] section.

  2. 2. Write the API key

    Run hermes config set provider.zhipu.api_key "YOUR_API_KEY" (mind the double quotes โ€” the key may contain dots). It lands encrypted in ~/.hermes/secrets.toml (mode 0600), so it will not be committed to git by accident.

  3. 3. Pick a model + verify

    hermes config set provider.zhipu.model glm-4-plus (recommended: top-tier capability). Then run hermes doctor --provider โ€” the script sends a ping request and you should see โœ… within 5 seconds.

Done. Start the agent with hermes and any task will now hit GLM. Next section: how to pick between the three tiers and what it costs.

API pricing and free quota

Zhipu charges per token, with input and output priced separately. Current three-tier lineup and target scenarios (2026-05 data):

  • GLM-4-Plus: top-tier, GPT-4o-class. Input 0.05 RMB / 1K tokens, output 0.05 RMB / 1K tokens. 128K context. Best for complex reasoning, code generation, data analysis.
  • GLM-4-Air: best value. Input 0.0005 RMB / 1K tokens, output 0.0005 RMB / 1K tokens (1% of Plus). 128K context. Best for high-volume, text processing, daily chat.
  • GLM-4-Flash: free up to quota. Both input and output 0 RMB. Lower capability but very fast. Best for pure text tasks, prototyping, personal demos.
  • Trial quota: real-name verified accounts get 18 RMB free credit (โ‰ˆ 360K GLM-4-Plus tokens, or 36M GLM-4-Air tokens). For everyday Hermes Agent usage this typically lasts 2-4 weeks.
  • Top-up: 10 RMB minimum, paid via Alipay / WeChat / corporate transfer. Prepaid only โ€” service stops at 0 balance, no debt risk.

These numbers shift as Zhipu adjusts pricing; refer to the official pricing page at https://open.bigmodel.cn/pricing for the latest. Hermes Agent does not handle billing โ€” all charges go directly to your Zhipu account.

Performance and scenario comparison

GLM-4 vs Claude / GPT-4 / Kimi in real Hermes Agent scenarios. Numbers are from the Hermes teamโ€™s 2026-04 measurements across 4 machines in China (50 runs per task, averaged). For sizing only.

  • Response latency (Beijing โ†’ API โ†’ first token): GLM-4-Plus 350-450ms, Kimi 400-500ms, Claude (direct) 2-5s, GPT-4o (direct) 3-8s. GLM wins decisively in-country.
  • Context window: GLM-4-Plus 128K, GLM-4-Air 128K, Kimi-128k 128K, Kimi-1M (long-context) 1M, GPT-4o 128K, Claude 3.7 Sonnet 200K.
  • Function-call accuracy (measured on Hermes Skills): GLM-4-Plus 92%, GPT-4o 94%, Claude 3.7 Sonnet 96%, Kimi-128k 87%. GLM ranks third โ€” already production-quality.
  • Chinese comprehension: GLM-4-Plus โ‰ฅ Kimi > Claude > GPT-4o. GLM is trained on the highest-quality Chinese corpus, with clear advantages on regional idioms and domain-specific Chinese.
  • English code generation: Claude 3.7 Sonnet > GPT-4o > GLM-4-Plus > Kimi. For heavy English-doc + complex code generation, Claude remains the top choice.
  • Monthly bill on the same workload (50 runs of "list and categorize Python project dependencies"): GLM-4-Air โ‰ˆ 0.04 USD, GLM-4-Plus โ‰ˆ 0.6 USD, Kimi-128k โ‰ˆ 0.3 USD, GPT-4o (incl. proxy traffic) โ‰ˆ 11 USD.

Rule of thumb: daily + in-China โ†’ GLM-4-Plus or Air. Need ultra-long context (>200K) โ†’ Kimi. Heavy English code + budget allows โ†’ Claude. Hermes supports multiple providers in parallel โ€” see /docs to switch per Skill.

Best practices (Skills / temperature / concurrency)

Once GLM is wired into Hermes, four practices keep your bill low and your output quality high.

  • Skill default temperature: GLM-4 reasons best at 0.3-0.5 (lower than GPT-4oโ€™s default 0.7). Set the global default with hermes config set llm.temperature 0.4, or override per Skill in its meta.yaml.
  • Long batches โ†’ Air, short reasoning โ†’ Plus: bulk classification / data cleaning / RAG retrieval go to Air; code generation / decision making / tool calling go to Plus. In ~/.hermes/config.toml under [llm.zhipu], set model_fast = "glm-4-air" and model_smart = "glm-4-plus" โ€” Hermes routes based on Skill tags.
  • Concurrency limits: free tier RPM = 5, paid default RPM = 60 (raise to 600+ via a support ticket). In Gateway mode with concurrent tasks, raise RPM at Zhipu first, otherwise 429s pile up.
  • Context window management: 128K does NOT mean "stuff it full". In practice, > 60K tokens triggers measurable middle-loss in GLM. For long docs, prefer chunked summarize, or switch to Kimi-1M (see /providers/kimi).

Combined with the Hermes Memory module (see /docs#memory), high-frequency context can be compressed under 5K โ€” long-session bills drop another ~60%.

Common errors

90% of day-to-day GLM errors fall into these five buckets. Walk through and match.

  • 401 Unauthorized / Invalid API key

    ๅŽŸๅ›  / Cause: Key copied with whitespace, quotes, or a trailing newline; or generated under a different account; or disabled due to long inactivity.

    ไฟฎๅค / Fix: Generate a fresh key in the open-platform console, run hermes config set provider.zhipu.api_key "NEW_KEY". Use ASCII double quotes, no surrounding whitespace.

  • 429 Too Many Requests / Rate limit exceeded

    ๅŽŸๅ›  / Cause: RPM (requests per minute) exceeded. Free tier 5 RPM hits the ceiling fast, especially in Gateway mode with concurrent tasks.

    ไฟฎๅค / Fix: Short-term: add a global cap with hermes config set llm.zhipu.rpm 4 (leave 1 RPM buffer). Long-term: open a ticket in the console to raise RPM. Paid users typically get 600 RPM within 24h.

  • Insufficient balance

    ๅŽŸๅ›  / Cause: Trial credit exhausted or prepaid balance hit zero.

    ไฟฎๅค / Fix: Top up via the console โ†’ Account Center โ†’ Top up. Keep at least ~7 USD buffer to avoid service interruption mid-task. Set a low-balance email alert (< 10 RMB).

  • Model not found: glm-4-plus / glm-4-air

    ๅŽŸๅ›  / Cause: Misspelled model ID, or your account lacks access (some new models are opt-in by default).

    ไฟฎๅค / Fix: Check the current allowlist at https://open.bigmodel.cn/dev/api, copy the exact ID to hermes config set provider.zhipu.model <id>. GLM-4-Flash sometimes needs manual enable in the "Service activation" page.

  • Timeout after 60s / context too long

    ๅŽŸๅ›  / Cause: 128K window stuffed too full causing inference timeout, or in-flight network blip.

    ไฟฎๅค / Fix: Run hermes context summarize to compress the session. For non-RAG large-doc cases, enable streaming with hermes config set llm.zhipu.stream true. If timeouts persist, switch to Kimi-128k (see /providers/kimi).

Next steps

Kimi long-context setup

Kimi-1M handles huge codebases / long contracts / research docs alongside GLM.

Configure Kimi

bundled and community Skills atlas

Hermes Skills system: 4 registries, self-learning loop, pairing with GLM function calls.

Skills guide

Full CLI reference

hermes config / skills / gateway โ€” 50+ commands grouped and cross-linked.

CLI reference

Back to documentation hub

Complete documentation synced with v2026.9.14: install, Skills, Memory, Gateway.

Docs hub