Hermes Agent long-context LLM guide

Configure Kimi (Moonshot) with Hermes Agent: 128K-1M long context

Kimi (Moonshot) is the long-context LLM of choice for Hermes Agent users in mainland China: 8K / 32K / 128K / 1M context tiers cover huge codebases, long contracts, and research-doc dialog. This guide covers API key sign-up, Hermes wiring, model selection, and 5 real-world long-context scenarios.

Hermes Agentv2026.9.14Last updated
  • ๐ŸชŸ 128K starting tier, 1M context dedicated model for ultra-long jobs
  • ๐Ÿ‡จ๐Ÿ‡ณ Direct connection inside China to platform.moonshot.cn, no proxy
  • ๐Ÿ’ฐ Kimi-128k โ‰ˆ 0.012 USD / 1K tokens, Kimi-1M โ‰ˆ 0.04 USD / 1K tokens
  • ๐Ÿค Co-exists with GLM-4 โ€” Hermes routes per Skill automatically

Why Kimi

Kimi is developed by Moonshot AI, the first Chinese LLM vendor to market long context as a primary capability. For Hermes Agent users, the value is clear: when GLM-4 / Claude 128K-200K windows fill up, Kimi-1M takes the baton.

  • ๐ŸชŸ 1M tokens of context: ~750K Chinese characters, enough to hold a complete medium codebase (100-150K LOC) or a 200-page PDF.
  • ๐Ÿ” High long-context recall: measured Kimi-1M needle-in-haystack accuracy > 95% across the 500K-1M token range, clearly ahead of Chinese peers that nominally support 128K but degrade past 60K.
  • ๐Ÿ‡จ๐Ÿ‡ณ Stable in-China access: moonshot.cn direct, Beijing / Shanghai / Shenzhen first-token latency โ‰ˆ 400-500ms. No proxy required.
  • โšก Streaming-friendly: long-output streaming is stable (low jitter, few drops), good fit for Hermes Gateway real-time echo scenarios.

Kimi is not a GLM replacement โ€” itโ€™s a complement. Daily: GLM-4-Plus. Long-context (> 60K) tasks: switch to Kimi. Hermes lets you configure multiple providers in parallel.

Get a Moonshot API key

Moonshot Open Platform flows like Zhipu โ€” about 5 minutes end to end.

  1. 1. Sign up at Moonshot Open Platform

    Visit https://platform.moonshot.cn โ†’ sign up with phone โ†’ real-name verify. Business accounts get higher RPM ceilings.

  2. 2. Generate and copy an API key

    After login โ†’ Account Settings โ†’ API Key Management โ†’ New API Key. Copy the key (format: sk-xxxxxx) to a password manager. The key is shown once only.

  3. 3. Top up (optional)

    New accounts get a 15 RMB trial credit. Kimi-128k input is โ‰ˆ 0.060 RMB / 1K tokens, so the trial covers ~250K tokens (about one medium PDF). To continue, top up via Finance Center (10 RMB minimum).

Recommend burning trial credit on 1-2 long-context tasks before paying โ€” see if the capability fits your workflow.

Configure Hermes to use Kimi

Hermes ships a built-in moonshot adapter using the official OpenAI-compatible endpoint. 3 commands and you are done.

hermes config set provider.moonshot.api_key "YOUR_KIMI_KEY"

Writes the Moonshot key encrypted to ~/.hermes/secrets.toml (mode 0600). Use ASCII double quotes.

  1. 1. Write the API key

    Once the command runs the key lands encrypted on disk โ€” it will not appear in git diffs. With multiple providers, Kimi and GLM keys coexist without interference.

  2. 2. Pick the default model

    hermes config set provider.moonshot.model moonshot-v1-128k (recommended: best 128K cost/perf). Switch to moonshot-v1-1m only when ultra-long context is needed.

  3. 3. Make Kimi the global default OR per-Skill

    Global switch: hermes config set provider moonshot. Per-Skill switch (keep GLM default, use Kimi only for long-context tasks): add provider_override: moonshot inside the Skill meta.yaml (see /skills).

Run hermes doctor --provider to verify โ€” a โœ… within 5 seconds means the connection is good.

Model selection (8k / 32k / 128k / 1M)

Moonshot exposes 4 tiers, differentiated by context window and pricing. They all look like "Kimi" but performance and cost vary by an order of magnitude โ€” picking wrong can multiply your bill 10ร—.

  • moonshot-v1-8k: 8K context, input โ‰ˆ 0.0024 USD/1K tokens, output 0.0024. Best for short chats, simple classification, prototypes. Lowest cost daily-chat tier.
  • moonshot-v1-32k: 32K context, input โ‰ˆ 0.0048 USD/1K tokens. Best for medium-length doc summary, few-file RAG.
  • moonshot-v1-128k: 128K context, input โ‰ˆ 0.012 USD/1K tokens. Hermes Agent default long-context choice, best cost/perf. Holds a 50-80K-character project doc.
  • moonshot-v1-1m (kimi-1m): 1M context, input โ‰ˆ 0.04 USD/1K tokens. Dedicated model โ€” slower and pricier but most capable. For ingesting an entire codebase or thick contract. Requires manual enable via console "Service activation".
  • kimi-k1.5 (reasoning-enhanced): time-limited availability for complex reasoning tasks, pay-as-you-go.

Default to moonshot-v1-128k. Only escalate to 1M when a single call demonstrably needs > 80K. Refer to the official Moonshot pricing page for the latest โ†’ https://platform.moonshot.cn/docs/pricing .

5 long-context scenarios where Kimi shines

Once Kimi is wired in, these 5 task types are where it actually outperforms shorter-context options. Each maps to a built-in / recommended Hermes Skill.

  • Scenario 1 โ€” Whole-codebase audit: Skill code-audit + Kimi-1M. Drop ~100K LOC into a single context window so the agent can find architectural smells, circular deps, security gaps. Models that degrade after 60K (most 128K-nominal Chinese LLMs) wobble here.
  • Scenario 2 โ€” Long contract / legal doc summary: Skill doc-summarize + Kimi-128k. Feed in a 200-page PDF (~300K chars), get per-chapter summaries + key-clause risk flags. More accurate than naive chunk-and-merge RAG.
  • Scenario 3 โ€” Multi-turn research doc dialog: Skill research-chat + Kimi-128k. Import 10-20 papers, hold multi-turn dialog with cross-paper reasoning and consistent context. Hermes Memory module can compress those long sessions down to 5K on reload.
  • Scenario 4 โ€” Multi-file Skill chain: Hermes Gateway running "read 50 YAML configs โ†’ produce a comparison table โ†’ write a migration script" runs end-to-end without intermediate chunking on Kimi-128k.
  • Scenario 5 โ€” 200+ turn conversation without context loss: Hermes Chat mode with Kimi-128k + Hermes auto-context management beats GLM-4 / GPT-4o, which start forgetting after ~30 turns.

If your daily tasks stay under 60K tokens, you do not need Kimi โ€” GLM-4-Plus is cheaper and faster. Pull Kimi out only when long context is the actual requirement.

Common errors

Four error buckets cover most Kimi day-to-day issues.

  • context_length_exceeded

    ๅŽŸๅ›  / Cause: The current dialog or single request exceeds the chosen modelโ€™s window. e.g. moonshot-v1-8k but the input is 12K tokens.

    ไฟฎๅค / Fix: Two options: (a) escalate: hermes config set provider.moonshot.model moonshot-v1-128k; (b) compress: hermes context summarize to compress prior turns. If you really need > 128K, switch to moonshot-v1-1m.

  • Model not found

    ๅŽŸๅ›  / Cause: Your account lacks access. moonshot-v1-1m is request-only by default.

    ไฟฎๅค / Fix: Console โ†’ Service Activation โ†’ request 1M model. Approval is usually 1-2 business days. For 128k and below, confirm real-name verification first.

  • 401 Invalid API key

    ๅŽŸๅ›  / Cause: Key copy error, expired, or generated under another account.

    ไฟฎๅค / Fix: Regenerate in the console, run hermes config set provider.moonshot.api_key "NEW_KEY". Moonshot keys start with sk- โ€” do not strip the prefix.

  • 429 Rate limit / TPM exceeded

    ๅŽŸๅ›  / Cause: Moonshot throttles by TPM (tokens per minute), not just RPM. One long-context request can blow the TPM ceiling.

    ไฟฎๅค / Fix: Console โ†’ Account Center โ†’ Rate Limits โ†’ request a higher TPM. Short-term workaround: hermes config set llm.moonshot.concurrent 1 to serialize requests until you get the quota bump.

Next steps

Back to the GLM pillar

GLM-4 is the everyday / short-context default โ€” pairs with Kimi via multi-provider config.

GLM setup

Skills system guide

bundled and community Skills + 4 registries โ€” best practices when pairing with Kimi long context.

Skills guide

Full CLI reference

hermes config / context / gateway โ€” 50+ commands grouped.

CLI reference

Back to documentation hub

Complete documentation synced with v2026.9.14: Memory / Gateway / multi-provider routing.

Docs hub