> ## Documentation Index
> Fetch the complete documentation index at: https://docs.nemu.cc/llms.txt
> Use this file to discover all available pages before exploring further.

# Any model through any API

> Call Claude from an OpenAI client, GPT from an Anthropic client, and what is translated on the way

The gateway speaks three protocols to you and three to your providers, and every
pairing works. The protocol you call is chosen by the URL, the protocol used
upstream is chosen by the provider's compatibility setting in the
[console](/console/providers).

| You call | Provider compatibility |
| - | - |
| `POST /v1/chat/completions` (OpenAI Chat) | OpenAI compatible, Anthropic, OpenAI Responses |
| `POST /v1/responses` (OpenAI Responses) | OpenAI compatible, Anthropic, OpenAI Responses |
| `POST /v1/messages` or `/anthropic/v1/messages` (Anthropic Messages) | OpenAI compatible, Anthropic, OpenAI Responses |

So an OpenAI SDK pointed at `https://api.nemu.cc/v1` can use a Claude model, and
the Anthropic SDK pointed at `https://api.nemu.cc` can use a GPT model, with the
same gateway key. The model name is always the gateway name, such as
`anthropic/claude-opus-5-5`.

## Reasoning and thinking

Ask for reasoning the way your client already does and it is translated for the
model behind it.

| Client | How you ask |
| - | - |
| OpenAI Chat | `reasoning_effort` |
| OpenAI Responses | `reasoning.effort` |
| Anthropic Messages | `thinking` with `output_config.effort`, or `thinking.budget_tokens` |

An effort of `max` or `xhigh` is lowered to the highest level the model accepts.
A model that cannot reason ignores the request instead of failing it, and says
so in the compatibility notes described below.

What comes back depends on the protocol.

* **Chat completions** return the reasoning text in `message.reasoning_content`
  (and `delta.reasoning_content` while streaming). From Claude models they also
  return `message.thinking_blocks`, which carry Anthropic's signature. Send
  those blocks back unchanged on the assistant message of your next request and
  Claude keeps its earlier thinking across tool calls. They are used only when
  reasoning is on for that request.
* **Responses** return a `reasoning` output item with a summary.
* **Anthropic Messages** return `thinking` blocks, and `redacted_thinking`
  blocks when the model hides its reasoning.

Many Claude models return an empty thinking text by default. The signature still
comes back and still works for the next turn.

## Usage and caching

Every protocol reports the same counts in its own fields.

| Count | Chat | Responses | Anthropic |
| - | - | - | - |
| Prompt tokens, cached included | `usage.prompt_tokens` | `usage.input_tokens` | `input_tokens` plus the two cache fields |
| Served from cache | `usage.prompt_tokens_details.cached_tokens` | `usage.input_tokens_details.cached_tokens` | `cache_read_input_tokens` |
| Written to cache | `usage.cache_creation_input_tokens` | not reported | `cache_creation_input_tokens` |
| Reasoning tokens | `usage.completion_tokens_details.reasoning_tokens` | `usage.output_tokens_details.reasoning_tokens` | `output_tokens_details.thinking_tokens` |

For Anthropic providers, `cache_control` on a message part, a system part or a
tool is passed through, including the `ttl` of `5m` or `1h`. If you send none,
the gateway places cache breakpoints on the system prompt, the tools and the end
of the conversation for you. If you send any, it adds none, so you stay under
Anthropic's limit of four.

## Structured output

`response_format` with `json_schema`, and the Responses `text.format`, become
Anthropic's `output_config.format` when the provider is Anthropic. `json_object`
has no Anthropic equivalent, so the gateway adds an instruction to the system
prompt asking for a single JSON object.

## Files and images

Images work in every direction, as base64 or as a URL. PDFs and text files sent
as a chat `file` part with `file_data`, or a Responses `input_file`, become
Anthropic `document` blocks. A `file_id` is not supported because the gateway has
no file store.

## Parameters

| Parameter | What happens |
| - | - |
| `max_tokens`, `max_completion_tokens`, `max_output_tokens` | All three mean the same thing and are mapped to the field upstream expects. Providers on `api.openai.com` receive `max_completion_tokens`. |
| `temperature` above 1 | Lowered to 1 for Anthropic providers. |
| `top_p` together with `temperature` | `top_p` is dropped for Anthropic providers, which refuse both. |
| `stop`, `stop_sequences` | Converted. OpenAI providers accept at most four. |
| `user`, `metadata.user_id` | Converted. |
| `parallel_tool_calls`, `disable_parallel_tool_use` | Converted. |
| Forced tool choice with thinking on | Relaxed to `auto`, because Anthropic refuses forcing a tool while thinking. |
| Forced tool choice on a model that refuses it | Some of the newest Claude models reject `required` or a named tool outright. The gateway retries once with `auto` and an instruction in the system prompt to call the tool, which is what Anthropic recommends for those models. The model almost always complies, but it is no longer guaranteed. |
| `n` above 1, `logprobs`, penalties, `seed`, `logit_bias` | Not supported by Anthropic providers and dropped. |
| `developer` messages | Treated as system messages. |

Non-streaming responses list every adjustment in `metadata.compatibility_notes`,
for example `top_p_dropped: temperature is set` or
`reasoning_effort_dropped: model does not support reasoning`.

Hosted tools such as web search, code execution and computer use are not
translated between protocols. They work when an Anthropic client talks to an
Anthropic provider, where the request is passed through, and are dropped
otherwise. For that same pairing the `anthropic-beta` header you send is
forwarded upstream.

## Responses that remember

`previous_response_id` works with every provider. The gateway stores the
conversation behind a response for 24 hours, scoped to your workspace, and
prepends it to the next request. Send `store: false` to keep nothing. A response
larger than 512 KB is not stored.

```bash theme={null}
curl https://api.nemu.cc/v1/responses \
  -H "Authorization: Bearer $NEMU_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"anthropic/claude-opus-5-5","previous_response_id":"resp_...","input":"And in Rome?"}'
```

`GET /v1/responses/{id}` returns a stored response and `DELETE /v1/responses/{id}`
removes it. An unknown or expired id answers 404 with the code
`previous_response_not_found`.

When the provider is itself an OpenAI Responses provider, the upstream keeps the
state and the id is the upstream's, so the gateway stores nothing for it.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.