Think of a big mailroom with one street address. Mail for accounting, mail for legal, mail for the warehouse — it all arrives at the same loading dock, and a sorter reads each label and sends it to the right department. The senders outside never need to know where any department actually sits inside the building.
Gateway Routing is that mailroom for your services. Clients send everything to one endpoint, and the gateway reads each request and forwards it to whichever backend should handle it — by path, by hostname, or by version.
The problem
When clients call services directly, they have to know your internal layout: which host serves /orders, where /users lives, which host has the new v2 API. That knowledge gets baked into every mobile app and browser you ship. The moment you split a service in two, rename one, or move part of it somewhere new, every client pointing at the old address breaks until it's updated — and you can't update apps already on people's phones. DNS can't save you either: it maps a hostname to servers but never sees the path, so it can't send /returns one way and /orders another.
Versioning makes it worse. Rolling out a new release means somehow getting clients to point at the new endpoint, with no clean way to send just a fraction of traffic to it first to make sure it's healthy.
Below, the Orders team splits returns out into a service of its own. Predict what happens to an app that's already installed, then flip to Gateway routing and make the same split.
How it works
The gateway publishes a single, stable endpoint and keeps the map of what lives where. When a request arrives, it inspects the path, host, or headers and forwards it to the matching backend: /orders/* goes to the orders service, /users/* to the users service, a v2 header or path prefix to the new version. The routing rules live entirely in the gateway's config, so clients stay blissfully unaware of the topology behind the door.
Because you control routing centrally, you can change it without redeploying a single client. Move a service and you just update one rule. Want a canary release? Send 10% of traffic to v2 and 90% to v1, then shift the dial as confidence grows.
Step through a canary below. Before the bad news arrives, predict how much damage a buggy v2 can do; then flip to All at once to see the same bug released without one.
Routing rules are a deployment superpower. Because the gateway decides where traffic goes, you get blue-green and canary releases almost for free: stand up the new version alongside the old, route a trickle of traffic to it, watch the metrics, then ramp up or roll back by editing one rule — no client ever notices.
You split /returns out of the Orders service into a new Returns service. Clients reach everything through a routing gateway. What's the smallest change that keeps already-installed apps working?
Canary by user, not by request. If every request is routed independently with 10% odds, one person can bounce between v1 and v2 within a single session: a cart saved by v2 and read back by v1, or a page that changes shape on every click. Most gateways can split on a stable key such as a user id, a cookie, or a header, so each person sees one version consistently while the overall share stays at 10%.
When to use it
Routing is worth it as soon as you have more than a couple of backend services that external clients would otherwise need to address individually, or when you want to evolve your service layout and versions without breaking clients. It's a foundational job of an API gateway and the natural companion to gateway aggregation and gateway offloading — route the request, combine what's needed, handle the shared concerns.
It's distinct from load balancing: a load balancer spreads requests across identical servers, while routing sends them to different services based on what the request says. For a single-service app there's nothing to route, so skip it — but the moment your backend becomes a collection of services, a routing front door is what keeps clients sane.
Your canary sends 10% of /checkout traffic to v2, and v2 fails 30% of the requests it gets. Roughly what share of all checkout requests fail?