Back to the GLM pillar
GLM-4 is the everyday / short-context default โ pairs with Kimi via multi-provider config.
GLM setupHermes Agent long-context LLM guide
Kimi (Moonshot) is the long-context LLM of choice for Hermes Agent users in mainland China: 8K / 32K / 128K / 1M context tiers cover huge codebases, long contracts, and research-doc dialog. This guide covers API key sign-up, Hermes wiring, model selection, and 5 real-world long-context scenarios.
Kimi is developed by Moonshot AI, the first Chinese LLM vendor to market long context as a primary capability. For Hermes Agent users, the value is clear: when GLM-4 / Claude 128K-200K windows fill up, Kimi-1M takes the baton.
Kimi is not a GLM replacement โ itโs a complement. Daily: GLM-4-Plus. Long-context (> 60K) tasks: switch to Kimi. Hermes lets you configure multiple providers in parallel.
Moonshot Open Platform flows like Zhipu โ about 5 minutes end to end.
1. Sign up at Moonshot Open Platform
Visit https://platform.moonshot.cn โ sign up with phone โ real-name verify. Business accounts get higher RPM ceilings.
2. Generate and copy an API key
After login โ Account Settings โ API Key Management โ New API Key. Copy the key (format: sk-xxxxxx) to a password manager. The key is shown once only.
3. Top up (optional)
New accounts get a 15 RMB trial credit. Kimi-128k input is โ 0.060 RMB / 1K tokens, so the trial covers ~250K tokens (about one medium PDF). To continue, top up via Finance Center (10 RMB minimum).
Recommend burning trial credit on 1-2 long-context tasks before paying โ see if the capability fits your workflow.
Hermes ships a built-in moonshot adapter using the official OpenAI-compatible endpoint. 3 commands and you are done.
hermes config set provider.moonshot.api_key "YOUR_KIMI_KEY"Writes the Moonshot key encrypted to ~/.hermes/secrets.toml (mode 0600). Use ASCII double quotes.
1. Write the API key
Once the command runs the key lands encrypted on disk โ it will not appear in git diffs. With multiple providers, Kimi and GLM keys coexist without interference.
2. Pick the default model
hermes config set provider.moonshot.model moonshot-v1-128k (recommended: best 128K cost/perf). Switch to moonshot-v1-1m only when ultra-long context is needed.
3. Make Kimi the global default OR per-Skill
Global switch: hermes config set provider moonshot. Per-Skill switch (keep GLM default, use Kimi only for long-context tasks): add provider_override: moonshot inside the Skill meta.yaml (see /skills).
Run hermes doctor --provider to verify โ a โ within 5 seconds means the connection is good.
Moonshot exposes 4 tiers, differentiated by context window and pricing. They all look like "Kimi" but performance and cost vary by an order of magnitude โ picking wrong can multiply your bill 10ร.
Default to moonshot-v1-128k. Only escalate to 1M when a single call demonstrably needs > 80K. Refer to the official Moonshot pricing page for the latest โ https://platform.moonshot.cn/docs/pricing .
Once Kimi is wired in, these 5 task types are where it actually outperforms shorter-context options. Each maps to a built-in / recommended Hermes Skill.
If your daily tasks stay under 60K tokens, you do not need Kimi โ GLM-4-Plus is cheaper and faster. Pull Kimi out only when long context is the actual requirement.
Four error buckets cover most Kimi day-to-day issues.
context_length_exceeded
ๅๅ / Cause: The current dialog or single request exceeds the chosen modelโs window. e.g. moonshot-v1-8k but the input is 12K tokens.
ไฟฎๅค / Fix: Two options: (a) escalate: hermes config set provider.moonshot.model moonshot-v1-128k; (b) compress: hermes context summarize to compress prior turns. If you really need > 128K, switch to moonshot-v1-1m.
Model not found
ๅๅ / Cause: Your account lacks access. moonshot-v1-1m is request-only by default.
ไฟฎๅค / Fix: Console โ Service Activation โ request 1M model. Approval is usually 1-2 business days. For 128k and below, confirm real-name verification first.
401 Invalid API key
ๅๅ / Cause: Key copy error, expired, or generated under another account.
ไฟฎๅค / Fix: Regenerate in the console, run hermes config set provider.moonshot.api_key "NEW_KEY". Moonshot keys start with sk- โ do not strip the prefix.
429 Rate limit / TPM exceeded
ๅๅ / Cause: Moonshot throttles by TPM (tokens per minute), not just RPM. One long-context request can blow the TPM ceiling.
ไฟฎๅค / Fix: Console โ Account Center โ Rate Limits โ request a higher TPM. Short-term workaround: hermes config set llm.moonshot.concurrent 1 to serialize requests until you get the quota bump.
GLM-4 is the everyday / short-context default โ pairs with Kimi via multi-provider config.
GLM setupbundled and community Skills + 4 registries โ best practices when pairing with Kimi long context.
Skills guidehermes config / context / gateway โ 50+ commands grouped.
CLI referenceComplete documentation synced with v2026.9.14: Memory / Gateway / multi-provider routing.
Docs hub