CLOUDFLARE has publicly released its AI Gateway Auto Router in public beta. The feature automatically routes each user request to an appropriate model, selecting the one that is capable enough for the task without requiring end users to choose models themselves. In internal testing with Cloudflare’s OpenCode harness, Auto Router delivered competitive performance to frontier models while cutting costs, achieving an approximate 30% saving compared with using only models such as OpenAI’s Sol or Anthropic’s Opus.
How it works is described in two stages. First, the gateway builds a pool of candidate models that can handle the request, filtering out those that do not support the required format or execution mode and accounting for credentials, billing rules, access controls and spend limits.
The system then analyses the most recent conversation and runs a compact, two‑stage scoring process: a multi‑head classifier assigns the task to one of 14 categories and rates the request on four dimensions (complexity, ambiguity, stakes, and context dependence). A separate scoring matrix combines these signals with model benchmarks to estimate fit, then a cost-aware utility is calculated as: utility = expected quality − adaptive cost penalty.
This means cheaper models can win on straightforward tasks, while more demanding tasks may justify higher‑cost options; cache read/write costs are also factored in across turns and sessions.
Cloudflare reports that Auto Router’s results for general knowledge work were comparable to state‑of‑the‑art daily drivers, with cloudflare/auto showing about 80% of the cost of Sol and 35% of Opus, depending on task and context. The project is in beta and the team plans to expand model support, introduce capacity‑aware filtering, and enhance decision modelling.