Skip to content

Command-line agents

Hermes

Set up Nous Research's Hermes Agent to use NoviaHub as a custom OpenAI-compatible endpoint.

Hermes Agent is an AI agent from Nous Research. You can chat with it in the terminal or use its desktop app, and connect it to messaging platforms such as Telegram, Discord and Slack (hermes gateway setup) and to editors (hermes acp). In Hermes’ own words, it works with any OpenAI-compatible API endpoint: if a server implements /v1/chat/completions, Hermes can use it. NoviaHub connects as a “Custom endpoint”.

终端窗口
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash

These commands install the command line only. There is also a desktop app: a DMG for macOS (Apple Silicon only; Intel Macs are not supported) and an .appinstaller for Windows. See the installation docs.

After installing, reload your shell config (source ~/.bashrc or source ~/.zshrc); then you can run hermes.

The examples use deepseek-v4-flash. Use either method.

Section titled “Method 1: the hermes model wizard (recommended by Hermes)”
  1. In a terminal (not inside a Hermes chat), run:

    终端窗口
    hermes model
  2. Choose “Custom endpoint (self-hosted / VLLM / etc.)”.

  3. Answer the prompts:

    Prompt Value
    API base URL https://noviahub.com/v1 (include /v1)
    API key Your NoviaHub key (starts with sk-)
    Model name deepseek-v4-flash
  4. The wizard also asks for the API mode and the context length:

    • API mode (the prompt reads “Select API compatibility mode”): choose Chat Completions (stored in the config as chat_completions).
    • Context length: enter the model’s context length, shown as Context on the model’s page under Models & Pricing (deepseek-v4-flash listed 1,000,000 on 2026-09-28). See Notes for why.
  1. Save the key. Run the command below; Hermes stores it in ~/.hermes/.env (every UPPER_SNAKE name goes to .env, never to config.yaml):

    终端窗口
    hermes config set NOVIAHUB_API_KEY "sk-..."
  2. Edit ~/.hermes/config.yaml and write (or change) the model section:

    ~/.hermes/config.yaml
    model:
    default: deepseek-v4-flash
    provider: custom
    base_url: https://noviahub.com/v1
    key_env: NOVIAHUB_API_KEY
    api_mode: chat_completions
    context_length: 1000000
    Setting Purpose
    default The default model ID. It must match the model ID on NoviaHub exactly.
    provider custom, meaning a custom endpoint.
    base_url The API address, including /v1: https://noviahub.com/v1.
    key_env The environment variable that holds the key (instead of writing api_key directly).
    api_mode The protocol; chat_completions means Chat Completions. Optional: Hermes already uses Chat Completions for an address like NoviaHub’s.
    context_length The model’s context length, stated explicitly; see Notes for why.
  3. Run hermes.

  1. Run hermes. It starts with a welcome banner showing the current model, available tools and skills.

  2. Send something easy to check, such as “Introduce yourself in one sentence.” A reply means the setup works.

  3. If something is wrong, run hermes doctor to check the configuration.

  4. Check the call and its cost under Usage Logs in the NoviaHub console.

  • Within a chat: type /model custom:<model ID>, for example /model custom:kimi-k3.
  • Two different commands: hermes model (run in the terminal) is the full setup wizard for adding providers and entering keys; /model (typed in a chat) only switches between providers and models you have already set up and cannot add a new provider.
  • Which models: Hermes calls the Chat Completions API, so pick models that carry the Chat label under Models & Pricing. Every model on the site carried it on 2026-09-28.
  • At least 64,000 tokens of context. Hermes says agent use with tools needs a model context of at least 64,000 tokens; smaller windows are rejected at startup. Check Context on the model’s page before choosing it.
  • Set context_length explicitly. Hermes tries to read the context length from the endpoint’s /v1/models, but the model entries NoviaHub’s /v1/models returns don’t include a context length. Hermes’ own fix for this is to pin it with context_length in config.yaml.
  • Use Chat Completions with a hand-written config. In Hermes’ source, a provider: custom config ignores api_mode: codex_responses unless the address is one of a few official hosts such as OpenAI’s, and sends Chat Completions requests anyway. That is why this page covers Chat Completions only; pick models with the Chat label.
  • OPENAI_BASE_URL has no effect on custom endpoints. That variable only applies to Hermes’ built-in openai-api provider; set a custom endpoint with hermes model or model.base_url.

Checked on 2026-09-28 and 2026-09-29: