Skip to content

Guides

Reasoning and thinking

Control how long a model "thinks": each protocol's own reasoning parameters, and NoviaHub's model-name modifiers @thinking and @effort.

Many newer models “think” for a while before they answer. More thinking gives better answers to hard questions, but takes longer and uses more tokens, so it costs more.

On NoviaHub you can control thinking in two ways:

  1. The protocol’s own parameters: put that protocol’s reasoning parameters in the request body, and NoviaHub passes them to the model.
  2. NoviaHub’s model-name modifiers: add a suffix such as @effort:high to the model name. The advantage is that the request body stays unchanged; you only change the model name. This suits tools that let you enter a model name but not custom parameters.
Protocol Parameter Example
OpenAI Chat Completions reasoning_effort "reasoning_effort": "low"
OpenAI Responses reasoning "reasoning": {"effort": "low"}
Anthropic Messages thinking (and output_config) As described in Anthropic’s documentation
Gemini generationConfig.thinkingConfig As described in Google’s documentation

In our tests, reasoning_effort was passed to the upstream unchanged. Which values each model accepts is up to its vendor; see the vendor’s documentation.

Append one or more @key:value pairs to the model name, for example:

deepseek-v4-flash@effort:high
claude-sonnet-5@thinking:off
deepseek-v4-flash@temperature:0.2@topp:0.9

When NoviaHub receives the request, it first strips these suffixes from the model name, then turns them into parameters the target model understands. The upstream receives the model name without the suffix.

Modifier Values Effect
@effort: none, minimal, low, medium, high, xhigh, max Thinking effort, from no thinking to the maximum
@thinking: on, adaptive, off, or an integer Turn thinking on / adaptive / off; an integer is a thinking budget in tokens, and 0 means off
@temperature: A number, such as 0.2 Same as the temperature request parameter
@topp: A number, such as 0.9 Same as the top_p request parameter

Rules:

  • Modifiers must come at the end of the model name. You can chain several, in any order.
  • If the same key appears twice, the last one wins.
  • Keys are case-insensitive.
  • Modifiers take precedence over the request body. For example, with both @temperature:0.2 and "temperature": 1 in the request, 0.2 is used. With @effort or @thinking, the reasoning parameters already in the request body are replaced by what the modifier translates to.
  • A mistake makes the request fail with HTTP 400 and the error code convert_request_failed. For example:
    • @foo:bar gives unsupported model modifier "foo";
    • @effort:huge gives invalid effort modifier value "huge".

These are the translations we observed in a test environment, that is, the parameters NoviaHub actually sent to the upstream. How a modifier is translated depends on which kind of thinking control the target model supports.

You write Called through The upstream receives
deepseek-v4-flash@effort:high Chat Completions "reasoning_effort": "high"
gpt-6-sol@effort:low Responses "reasoning": {"effort": "low"}
claude-sonnet-5@effort:high Anthropic Messages "thinking": {"type": "adaptive"} and "output_config": {"effort": "high"}
gemini-3-flash@thinking:off Chat Completions "thinkingConfig": {"thinkingLevel": "minimal"}
deepseek-v4-flash@temperature:0.2@topp:0.9 Chat Completions "temperature": 0.2 and "top_p": 0.9

Some models can’t turn thinking off completely, for example the Gemini 3 family. With @thinking:off or @effort:none, NoviaHub uses the model’s lowest thinking level instead (minimal in the table above) rather than returning an error.

  • Modifiers don’t work in the URL of the native Gemini API. The native Gemini API puts the model name in the URL (/v1beta/models/{model}:generateContent), and NoviaHub cuts the model name at the first colon. gemini-3-flash@thinking:off:generateContent is read as the model gemini-3-flash@thinking and gets a 503 “no available channel”. With the native Gemini API, use thinkingConfig in the request body instead, or switch to Chat Completions and put the modifier in the model field.
  • Billing uses the base model. NoviaHub currently has no separate prices for modifiers, so you pay the price of the model without the suffix. Usage logs also record the model name without the suffix, and Reasoning Effort in the details shows the effort used. With thinking on, the model generates extra thinking content, which usually counts as output tokens, so the cost goes up accordingly.
  • An API key’s Model Limits check the base model. For example, a key that only allows deepseek-v4-flash can also call deepseek-v4-flash@effort:high.
  • Some model names already end in -high or -low. For example gemini-3.6-flash-high, gemini-3.8-flash-high and gemini-3.1-pro-low on NoviaHub are complete model names; the trailing -high and -low are not modifiers. Enter them as they are.

For backward compatibility, NoviaHub also recognises some suffixes without @:

  • An effort word after a GPT model name: -none, -minimal, -low, -medium, -high, -xhigh or -max. For example, gpt-6-sol-high is the same as gpt-6-sol@effort:high; in our tests the upstream received "reasoning_effort": "high".
  • -thinking after a Claude model name turns thinking on. Whether this form works depends on platform settings.

Legacy suffixes are easier to confuse with real model names, so use the @ form for new integrations.