API
Protocol conversion
Which endpoints a model supports, how NoviaHub converts between the OpenAI, Anthropic and Gemini protocols, and what gets lost on the way.
Models from different vendors speak different “languages”: Claude uses the Anthropic Messages protocol, Gemini uses the Gemini protocol, and many others use the OpenAI protocol. NoviaHub translates in between: whichever protocol you call with, it converts the request into a format the model understands and converts the reply back into your protocol.
That lets you call models from different vendors with the same code. Translation has limits, though. This page explains which combinations work and what information is lost in conversion.
Start with the model’s endpoint labels
Section titled “Start with the model’s endpoint labels”Open a model on the Models & Pricing page to see its endpoint labels. They tell you which endpoints can call this model:
| Label | Endpoint |
|---|---|
| Chat | POST /v1/chat/completions |
| Response | POST /v1/responses |
openai-response-compact |
POST /v1/responses/compact |
openai-alpha-search |
POST /v1/alpha/search |
| Anthropic | POST /v1/messages |
| Gemini | POST /v1beta/models/{model}:generateContent |
| Image | POST /v1/images/generations |
The same information is in the supported_endpoint_types field returned by GET /v1/models.
The rule is simple: only use the endpoints listed in the labels. The labels are generated automatically from the service behind the model, and each listed endpoint has a working implementation. Some combinations outside the labels also work (in our tests, calling a Claude model labelled only Anthropic and Chat through /v1/responses succeeded), but we don’t guarantee them and they may change.
Conversions we tested
Section titled “Conversions we tested”We ran these on a local test instance built from the same version as production, with simulated upstream services. All returned 200, in the format of the endpoint that was called:
| Endpoint you call | Model type | Result |
|---|---|---|
Chat (/v1/chat/completions) |
Claude model | Works; Chat format returned |
| Chat | Gemini model | Works; Chat format returned |
Anthropic (/v1/messages) |
OpenAI-protocol model (such as deepseek-v4-flash) | Works; Anthropic format returned, streaming included |
| Anthropic | Gemini model | Works; Anthropic format returned |
Gemini (/v1beta/...:generateContent) |
OpenAI-protocol model | Works; Gemini format returned |
Response (/v1/responses) |
Claude model | Works; Responses format returned (not in the labels, see above) |
What conversion loses or changes
Section titled “What conversion loses or changes”Extra fields in the response
Section titled “Extra fields in the response”After conversion, usage may contain fields beyond the protocol’s standard ones, such as billing_usage, usage_semantic, usage_source and claude_cache_creation_5_m_tokens. The gateway uses them for accounting; you can ignore them.
Claude thinking signatures
Section titled “Claude thinking signatures”- When you call a Claude model through Chat Completions, its thinking comes back in
reasoning_content, but the signature attached to Claude’s thinking blocks is not kept. - The other way round,
thinkingblocks returned when you call a non-Claude model through/v1/messageshave no signature.
If your program must send signed thinking blocks back to Claude unchanged (for example in multi-turn conversations that combine extended thinking with tool use), call Claude models directly through /v1/messages.
Calling non-Claude models through /v1/messages
Section titled “Calling non-Claude models through /v1/messages”- A
systemwritten as an array of text blocks is merged into one string, andcache_controlon those blocks is lost. - All tools are treated as function tools.
- Send images as base64 (
source.typeset tobase64).
Calling Claude models through Chat Completions
Section titled “Calling Claude models through Chat Completions”- The Anthropic protocol requires
max_tokens. If you leave it out, the gateway fills in a default (built-in default 8192; the platform may change it). To control output length, sendmax_tokensormax_completion_tokensyourself. stopbecomesstop_sequences,developermessages are merged intosystem, andweb_search_optionsbecomes Claude’s web search tool.
Token counts in streamed responses
Section titled “Token counts in streamed responses”When you stream a non-Claude model through /v1/messages, input_tokens in the opening message_start event is the gateway’s estimate. The usage in the final message_delta event is authoritative.
Recommendations
Section titled “Recommendations”- For new code, prefer the protocol listed first in the model’s labels. That is usually the vendor’s own protocol and exposes the most features.
- To call models from several vendors with one code path, Chat Completions is the easiest: almost every model is labelled Chat.
- Before relying on advanced features such as thinking, caching or tool calling, try them on a small request first.