Skip to content

OpenAI-compatible

Responses

Call models on NoviaHub in the OpenAI Responses format, plus the related Responses Compact and /v1/alpha/search endpoints.

Compatible with OpenAI’s Create a model response. Tools such as Codex use this endpoint.

POST https://noviahub.com/v1/responses
Header Required Description
Authorization Yes Bearer <your API key>
Content-Type Yes application/json

These are the parameters NoviaHub recognises and forwards. Whether a parameter takes effect depends on the model.

Parameter Type Required Description
model string Yes Model ID, such as gpt-6-sol.
input string / array No Input: a string, or an array of messages or content items as defined by OpenAI.
instructions string No System-level instructions.
max_output_tokens integer No Maximum tokens to generate.
temperature number No Sampling temperature.
top_p number No Nucleus sampling.
stream boolean No true returns a stream (SSE).
stream_options object No Streaming options.
tools array No Available tools.
tool_choice string / object No How tools are chosen.
parallel_tool_calls boolean No Allow several tool calls at once.
max_tool_calls integer No Maximum number of tool calls.
reasoning object No Reasoning settings, such as {"effort": "medium", "summary": "auto"}.
text object No Text output settings (format and so on).
previous_response_id string No Continue from an earlier reply.
include array No Extra content to return.
store boolean No Whether the upstream stores this reply.
truncation string No What to do when the context is too long.
prompt_cache_key string No Prompt cache key.
top_logprobs integer No Number of candidate token probabilities to return.
metadata, user — No Extra data and user identifier, forwarded as is.

NoviaHub also recognises conversation, context_management, prompt, prompt_cache_options, prompt_cache_retention, frequency_penalty, presence_penalty, moderation, client_metadata, and extension fields such as enable_thinking, thinking_budget, chat_template_kwargs, top_k, min_p, repetition_penalty and stop.

终端窗口
curl https://noviahub.com/v1/responses \
-H "Authorization: Bearer $NOVIAHUB_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-6-sol",
"input": "Explain what an API is in one sentence."
}'
Field Description
id ID of this reply.
object Always response.
created_at Creation time (Unix seconds).
status Status; completed when finished normally.
model Model ID.
output Output items. The answer is in items of type message, in content[] entries of type output_text, under text.
usage.input_tokens / output_tokens / total_tokens Input, output and total tokens.
usage.input_tokens_details.cached_tokens Input tokens served from cache.
usage.output_tokens_details.reasoning_tokens Tokens spent on reasoning.
Example response (test instance; content from a simulated upstream)
{
"id": "resp_mock123",
"object": "response",
"created_at": 1790609867,
"status": "completed",
"model": "gpt-6-sol",
"output": [
{
"type": "message",
"id": "msg_mock123",
"status": "completed",
"role": "assistant",
"content": [{ "type": "output_text", "text": "Hello from the mock upstream.", "annotations": [] }]
}
],
"usage": {
"input_tokens": 12,
"output_tokens": 7,
"total_tokens": 19,
"input_tokens_details": { "cached_tokens": 0 },
"output_tokens_details": { "reasoning_tokens": 0 }
}
}

With "stream": true, the response arrives as SSE events. Each event has an event: line and a data: line. Common events:

  • response.created: generation started.
  • response.output_text.delta: a new piece of text, in delta.
  • response.completed: generation finished; response holds the full result, including usage.
Example stream (test instance)
event: response.created
data: {"type":"response.created","sequence_number":0,"response":{"id":"resp_mock123","object":"response","created_at":1790609867,"status":"in_progress","model":"gpt-6-sol","output":[],"usage":null}}
event: response.output_text.delta
data: {"type":"response.output_text.delta","sequence_number":1,"item_id":"msg_mock123","output_index":0,"content_index":0,"delta":"Hello"}
event: response.output_text.delta
data: {"type":"response.output_text.delta","sequence_number":2,"item_id":"msg_mock123","output_index":0,"content_index":0,"delta":" from the mock upstream."}
event: response.completed
data: {"type":"response.completed","sequence_number":3,"response":{"id":"resp_mock123","object":"response","created_at":1790609867,"status":"completed","model":"gpt-6-sol","output":[{"type":"message","id":"msg_mock123","status":"completed","role":"assistant","content":[{"type":"output_text","text":"Hello from the mock upstream.","annotations":[]}]}],"usage":{"input_tokens":12,"output_tokens":7,"total_tokens":19,"input_tokens_details":{"cached_tokens":0},"output_tokens_details":{"reasoning_tokens":0}}}}
POST https://noviahub.com/v1/responses/compact

This matches OpenAI’s conversation compaction endpoint: it condenses a long conversation history so the conversation can continue. It is available only for models labelled openai-response-compact.

  • model is required. Only these fields are forwarded upstream: model, input, instructions, previous_response_id, parallel_tool_calls, prompt_cache_key, prompt_cache_options, prompt_cache_retention.
  • tools, reasoning and text are recognised but not forwarded; service_tier is removed by default.
  • No streaming; the response always comes back in one piece.
  • The body is returned from the upstream as is, with id, object, created_at, output and usage (plus error on failure).
  • Billed by tokens, like a normal chat request.
POST https://noviahub.com/v1/alpha/search

This is the standalone web search endpoint used by the Codex command-line tool; most applications don’t need it. It is available only for models labelled openai-alpha-search.

  • Only model is required. The body is forwarded upstream as is, and the response is returned as is.
  • No streaming.
  • Billing is different: the upstream reports no token usage, so NoviaHub charges one web-search tool call per request, with zero tokens. The price is set by the platform.
  • Calling it for an unsupported model returns an error, such as channel does not support /v1/alpha/search.
HTTP status Cause
400 Missing model, or the body isn’t valid JSON.
401 The key is invalid, disabled, expired or out of quota.
403 The key isn’t allowed to use this model, the IP isn’t on the allow list, or the account balance is too low.
503 Wrong model ID, or no channel is currently available for the model.

See Errors and troubleshooting for details.