Skip to content

API

Protocol conversion

Which endpoints a model supports, how NoviaHub converts between the OpenAI, Anthropic and Gemini protocols, and what gets lost on the way.

Models from different vendors speak different “languages”: Claude uses the Anthropic Messages protocol, Gemini uses the Gemini protocol, and many others use the OpenAI protocol. NoviaHub translates in between: whichever protocol you call with, it converts the request into a format the model understands and converts the reply back into your protocol.

That lets you call models from different vendors with the same code. Translation has limits, though. This page explains which combinations work and what information is lost in conversion.

Open a model on the Models & Pricing page to see its endpoint labels. They tell you which endpoints can call this model:

Label Endpoint
Chat POST /v1/chat/completions
Response POST /v1/responses
openai-response-compact POST /v1/responses/compact
openai-alpha-search POST /v1/alpha/search
Anthropic POST /v1/messages
Gemini POST /v1beta/models/{model}:generateContent
Image POST /v1/images/generations

The same information is in the supported_endpoint_types field returned by GET /v1/models.

The rule is simple: only use the endpoints listed in the labels. The labels are generated automatically from the service behind the model, and each listed endpoint has a working implementation. Some combinations outside the labels also work (in our tests, calling a Claude model labelled only Anthropic and Chat through /v1/responses succeeded), but we don’t guarantee them and they may change.

We ran these on a local test instance built from the same version as production, with simulated upstream services. All returned 200, in the format of the endpoint that was called:

Endpoint you call Model type Result
Chat (/v1/chat/completions) Claude model Works; Chat format returned
Chat Gemini model Works; Chat format returned
Anthropic (/v1/messages) OpenAI-protocol model (such as deepseek-v4-flash) Works; Anthropic format returned, streaming included
Anthropic Gemini model Works; Anthropic format returned
Gemini (/v1beta/...:generateContent) OpenAI-protocol model Works; Gemini format returned
Response (/v1/responses) Claude model Works; Responses format returned (not in the labels, see above)

After conversion, usage may contain fields beyond the protocol’s standard ones, such as billing_usage, usage_semantic, usage_source and claude_cache_creation_5_m_tokens. The gateway uses them for accounting; you can ignore them.

  • When you call a Claude model through Chat Completions, its thinking comes back in reasoning_content, but the signature attached to Claude’s thinking blocks is not kept.
  • The other way round, thinking blocks returned when you call a non-Claude model through /v1/messages have no signature.

If your program must send signed thinking blocks back to Claude unchanged (for example in multi-turn conversations that combine extended thinking with tool use), call Claude models directly through /v1/messages.

Calling non-Claude models through /v1/messages

Section titled “Calling non-Claude models through /v1/messages”
  • A system written as an array of text blocks is merged into one string, and cache_control on those blocks is lost.
  • All tools are treated as function tools.
  • Send images as base64 (source.type set to base64).

Calling Claude models through Chat Completions

Section titled “Calling Claude models through Chat Completions”
  • The Anthropic protocol requires max_tokens. If you leave it out, the gateway fills in a default (built-in default 8192; the platform may change it). To control output length, send max_tokens or max_completion_tokens yourself.
  • stop becomes stop_sequences, developer messages are merged into system, and web_search_options becomes Claude’s web search tool.

When you stream a non-Claude model through /v1/messages, input_tokens in the opening message_start event is the gateway’s estimate. The usage in the final message_delta event is authoritative.

  • For new code, prefer the protocol listed first in the model’s labels. That is usually the vendor’s own protocol and exposes the most features.
  • To call models from several vendors with one code path, Chat Completions is the easiest: almost every model is labelled Chat.
  • Before relying on advanced features such as thinking, caching or tool calling, try them on a small request first.