openai/gpt-4o as usual. Once that model has cost
more than €50 in the last seven days, they run on openai/gpt-4o-mini instead,
and go back to normal at the start of the next week.
Nothing in your application changes. It keeps sending openai/gpt-4o.
The graph
Four nodes.Building it
-
Open Workflows in the console and create one. Give it a name you will
recognise later, such as
gpt-4o cost cap. -
Add a Model trigger and set its model to
openai/gpt-4o. This is the name your application sends. -
Add a Metric check and connect the trigger to it. Add one condition:
metric
cost, operatorgreater_than, threshold50, window7 days. Leave inheritance on automatic, so it measures the triggering model. -
Add a Model selection set to
openai/gpt-4o-mini. Drag from the metric check’s success output to it. Success means the condition passed, which here means you have gone over budget. -
Add a second Model selection set to
openai/gpt-4o. Drag from the metric check’s failure output to it. This is the normal path. - Apply changes, then enable the workflow from the list.
Confirming it works
Send a request foropenai/gpt-4o and open Metrics. The recorded model is
the one that actually ran, so under budget you see openai/gpt-4o and over it
you see openai/gpt-4o-mini.
If both show the original model, the workflow is not routing. The usual causes
are that it is not enabled, that the trigger name does not exactly match what
your client sends, or that a node in the graph failed validation. See the
warning about unfinished nodes in the node reference.
Why the second selection node
You could leave the failure output unconnected and the request would run onopenai/gpt-4o anyway, because a workflow that reaches no selection node
changes nothing.
Wiring it explicitly is still worth it. It makes the normal path visible on the
canvas, and it gives you one place to change if the default model ever moves.
Variations
Swap the metric forerror_rate over 1 hour to fail away from a model that has
started returning errors. Swap it for latency_avg to shed load onto something
faster under pressure. Add an Image detection node before the selection to
send only requests carrying an image to a model that can see.