Guides
Reasoning and thinking
Control how long a model "thinks": each protocol's own reasoning parameters, and NoviaHub's model-name modifiers @thinking and @effort.
Many newer models “think” for a while before they answer. More thinking gives better answers to hard questions, but takes longer and uses more tokens, so it costs more.
On NoviaHub you can control thinking in two ways:
- The protocol’s own parameters: put that protocol’s reasoning parameters in the request body, and NoviaHub passes them to the model.
- NoviaHub’s model-name modifiers: add a suffix such as
@effort:highto the model name. The advantage is that the request body stays unchanged; you only change the model name. This suits tools that let you enter a model name but not custom parameters.
Option 1: the protocol’s own parameters
Section titled “Option 1: the protocol’s own parameters”| Protocol | Parameter | Example |
|---|---|---|
| OpenAI Chat Completions | reasoning_effort |
"reasoning_effort": "low" |
| OpenAI Responses | reasoning |
"reasoning": {"effort": "low"} |
| Anthropic Messages | thinking (and output_config) |
As described in Anthropic’s documentation |
| Gemini | generationConfig.thinkingConfig |
As described in Google’s documentation |
In our tests, reasoning_effort was passed to the upstream unchanged. Which values each model accepts is up to its vendor; see the vendor’s documentation.
Option 2: model-name modifiers
Section titled “Option 2: model-name modifiers”Append one or more @key:value pairs to the model name, for example:
deepseek-v4-flash@effort:highclaude-sonnet-5@thinking:offdeepseek-v4-flash@temperature:0.2@topp:0.9When NoviaHub receives the request, it first strips these suffixes from the model name, then turns them into parameters the target model understands. The upstream receives the model name without the suffix.
Available modifiers
Section titled “Available modifiers”| Modifier | Values | Effect |
|---|---|---|
@effort: |
none, minimal, low, medium, high, xhigh, max |
Thinking effort, from no thinking to the maximum |
@thinking: |
on, adaptive, off, or an integer |
Turn thinking on / adaptive / off; an integer is a thinking budget in tokens, and 0 means off |
@temperature: |
A number, such as 0.2 |
Same as the temperature request parameter |
@topp: |
A number, such as 0.9 |
Same as the top_p request parameter |
Rules:
- Modifiers must come at the end of the model name. You can chain several, in any order.
- If the same key appears twice, the last one wins.
- Keys are case-insensitive.
- Modifiers take precedence over the request body. For example, with both
@temperature:0.2and"temperature": 1in the request,0.2is used. With@effortor@thinking, the reasoning parameters already in the request body are replaced by what the modifier translates to. - A mistake makes the request fail with HTTP 400 and the error code
convert_request_failed. For example:@foo:bargivesunsupported model modifier "foo";@effort:hugegivesinvalid effort modifier value "huge".
What they turn into
Section titled “What they turn into”These are the translations we observed in a test environment, that is, the parameters NoviaHub actually sent to the upstream. How a modifier is translated depends on which kind of thinking control the target model supports.
| You write | Called through | The upstream receives |
|---|---|---|
deepseek-v4-flash@effort:high |
Chat Completions | "reasoning_effort": "high" |
gpt-6-sol@effort:low |
Responses | "reasoning": {"effort": "low"} |
claude-sonnet-5@effort:high |
Anthropic Messages | "thinking": {"type": "adaptive"} and "output_config": {"effort": "high"} |
gemini-3-flash@thinking:off |
Chat Completions | "thinkingConfig": {"thinkingLevel": "minimal"} |
deepseek-v4-flash@temperature:0.2@topp:0.9 |
Chat Completions | "temperature": 0.2 and "top_p": 0.9 |
Some models can’t turn thinking off completely, for example the Gemini 3 family. With @thinking:off or @effort:none, NoviaHub uses the model’s lowest thinking level instead (minimal in the table above) rather than returning an error.
Things to know
Section titled “Things to know”- Modifiers don’t work in the URL of the native Gemini API. The native Gemini API puts the model name in the URL (
/v1beta/models/{model}:generateContent), and NoviaHub cuts the model name at the first colon.gemini-3-flash@thinking:off:generateContentis read as the modelgemini-3-flash@thinkingand gets a 503 “no available channel”. With the native Gemini API, usethinkingConfigin the request body instead, or switch to Chat Completions and put the modifier in themodelfield. - Billing uses the base model. NoviaHub currently has no separate prices for modifiers, so you pay the price of the model without the suffix. Usage logs also record the model name without the suffix, and Reasoning Effort in the details shows the effort used. With thinking on, the model generates extra thinking content, which usually counts as output tokens, so the cost goes up accordingly.
- An API key’s Model Limits check the base model. For example, a key that only allows
deepseek-v4-flashcan also calldeepseek-v4-flash@effort:high. - Some model names already end in
-highor-low. For examplegemini-3.6-flash-high,gemini-3.8-flash-highandgemini-3.1-pro-lowon NoviaHub are complete model names; the trailing-highand-loware not modifiers. Enter them as they are.
Legacy suffixes
Section titled “Legacy suffixes”For backward compatibility, NoviaHub also recognises some suffixes without @:
- An effort word after a GPT model name:
-none,-minimal,-low,-medium,-high,-xhighor-max. For example,gpt-6-sol-highis the same asgpt-6-sol@effort:high; in our tests the upstream received"reasoning_effort": "high". -thinkingafter a Claude model name turns thinking on. Whether this form works depends on platform settings.
Legacy suffixes are easier to confuse with real model names, so use the @ form for new integrations.