Skip to main content
The gateway speaks three protocols to you and three to your providers, and every pairing works. The protocol you call is chosen by the URL, the protocol used upstream is chosen by the provider’s compatibility setting in the console. So an OpenAI SDK pointed at https://api.nemu.cc/v1 can use a Claude model, and the Anthropic SDK pointed at https://api.nemu.cc can use a GPT model, with the same gateway key. The model name is always the gateway name, such as anthropic/claude-opus-5-5.

Reasoning and thinking

Ask for reasoning the way your client already does and it is translated for the model behind it. An effort of max or xhigh is lowered to the highest level the model accepts. A model that cannot reason ignores the request instead of failing it, and says so in the compatibility notes described below. What comes back depends on the protocol.
  • Chat completions return the reasoning text in message.reasoning_content (and delta.reasoning_content while streaming). From Claude models they also return message.thinking_blocks, which carry Anthropic’s signature. Send those blocks back unchanged on the assistant message of your next request and Claude keeps its earlier thinking across tool calls. They are used only when reasoning is on for that request.
  • Responses return a reasoning output item with a summary.
  • Anthropic Messages return thinking blocks, and redacted_thinking blocks when the model hides its reasoning.
Many Claude models return an empty thinking text by default. The signature still comes back and still works for the next turn.

Usage and caching

Every protocol reports the same counts in its own fields. For Anthropic providers, cache_control on a message part, a system part or a tool is passed through, including the ttl of 5m or 1h. If you send none, the gateway places cache breakpoints on the system prompt, the tools and the end of the conversation for you. If you send any, it adds none, so you stay under Anthropic’s limit of four.

Structured output

response_format with json_schema, and the Responses text.format, become Anthropic’s output_config.format when the provider is Anthropic. json_object has no Anthropic equivalent, so the gateway adds an instruction to the system prompt asking for a single JSON object.

Files and images

Images work in every direction, as base64 or as a URL. PDFs and text files sent as a chat file part with file_data, or a Responses input_file, become Anthropic document blocks. A file_id is not supported because the gateway has no file store.

Parameters

Non-streaming responses list every adjustment in metadata.compatibility_notes, for example top_p_dropped: temperature is set or reasoning_effort_dropped: model does not support reasoning. Hosted tools such as web search, code execution and computer use are not translated between protocols. They work when an Anthropic client talks to an Anthropic provider, where the request is passed through, and are dropped otherwise. For that same pairing the anthropic-beta header you send is forwarded upstream.

Responses that remember

previous_response_id works with every provider. The gateway stores the conversation behind a response for 24 hours, scoped to your workspace, and prepends it to the next request. Send store: false to keep nothing. A response larger than 512 KB is not stored.
GET /v1/responses/{id} returns a stored response and DELETE /v1/responses/{id} removes it. An unknown or expired id answers 404 with the code previous_response_not_found. When the provider is itself an OpenAI Responses provider, the upstream keeps the state and the id is the upstream’s, so the gateway stores nothing for it.