A request becomes an ordered list of candidates: each model in models (or the single model), and for each model the endpoints able to serve it, ordered by your provider preferences, their measured health and their price. Up to three attempts are made down that list, and the first that answers serves the request.
What moves on to the next candidate
- A provider error, 5xx, or a 429 rate limit from the provider.
- A provider that does not answer in time. Each phase has its own limit: the first byte, the gap between streamed chunks, and the whole response.
- An endpoint that cannot carry the request, such as a content type that provider does not accept.
What does not
- A request error every endpoint would refuse, such as malformed messages or an invalid parameter. It returns 400 at once, and no other endpoint is tried.
- A failure after the first byte of a response has been sent. Text already delivered cannot be taken back, so the request is not restarted on another provider.
- A refusal by your own limits or guardrails, which is final.
- Your client disconnecting. The request stops, and nothing is retried on your behalf.
Health
Every attempt feeds the endpoint’s measured health, and health ranks the candidates, so an endpoint that is failing now is tried later. A disconnect by your client is not counted against the provider.
Seeing what happened
One row per attempt made for a request, in order, with each provider’s outcome. A request that failed over twice shows three rows.