A user clicks "Export my report" and then stares at a spinner for twenty seconds while the server crunches numbers and renders a PDF. Their request thread is held hostage the whole time, and if a handful of people click at once, the web server runs out of threads and everyone starts waiting — even people who just wanted to load the home page.
The Web-Queue-Worker style fixes this by getting slow work off the request path entirely. The web tier hands the heavy job to a queue and answers the user immediately; a separate worker does the actual grinding in the background.
The problem
When a web front end does slow work inline — sending a confirmation email, resizing an uploaded image, generating a report — it ties up the request thread for the entire duration of that work. The user waits, the connection stays open, and a thread that could be serving other requests is stuck doing one slow thing.
Under load this falls apart quickly. A web server has a fixed pool of request threads, and a burst of expensive requests can occupy every one of them. Requests that need only milliseconds of work — the home page, a login — now wait behind twenty-second jobs, latency climbs for everyone, and the whole site can look down because heavy background work and fast page loads are competing for the same threads.
Step through a Monday-morning rush below and predict what happens to a visitor who only wants the home page. Then flip the switch to replay the same rush with the slow work on a queue.
How it works
The style splits the application into three cooperating pieces. The web front end handles incoming HTTP requests and is built to respond fast — when it hits something slow, it doesn't do the work itself, it drops a message describing the job onto a queue and returns right away. The queue buffers those background jobs, and a separate worker process pulls them off and does the heavy lifting asynchronously.
The web tier and the worker typically share data stores — a database, blob storage — so the worker can read the inputs and write back results that the front end can later show the user.
The queue does more than pass work along: it keeps each job safe while a worker processes it. Most queues don't delete a message when a worker receives it. They hide it for a while — often called a visibility timeout or a lock — and delete it only when the worker reports that it's done. Step through one export below. The worker crashes halfway through, and you'll predict what happens to the job.
That promise is called at-least-once delivery: a job is never silently lost, but it can run more than once. So every worker handler should be idempotent — running it twice must leave things the same as running it once. Write outputs under the job's ID so a rerun overwrites instead of duplicating, check whether a side effect already happened before repeating it, and pass an idempotency key to services that accept one.
Set the hide timeout longer than your slowest job. If a healthy worker takes longer than the timeout, the queue assumes it died and hands the same job to another worker while the first is still busy. Now two workers run one job at once. Size the timeout from real job durations, or have long jobs extend their lock as they go.
The queue is the seam that makes everything else possible. It lets the web tier say "I've accepted this work" without saying "I've finished it," which is precisely what frees the request thread to move on to the next user.
Why the queue matters
The queue isn't just a hand-off; it's a shock absorber. This is queue-based load leveling in action: when traffic spikes, jobs pile up in the queue instead of overwhelming the worker, and the worker drains them at a steady, sustainable rate. The web tier never has to slow down just because the worker is busy.
It also lets the two halves scale independently. If the backlog grows, you add more workers reading from the same queue — that's competing consumers — so processing throughput becomes a dial you turn separately from your web capacity.
Your export endpoint now enqueues a job and returns 202. A newsletter sends 500 people to click Export within a minute. What grows?
Living with asynchronous results
There's a catch you have to design for: because the web tier responds before the job is done, the result is asynchronous. The user gets an immediate "we're working on it," not the finished output, so you need a way to deliver the result later.
Common approaches are to have the worker write a status record the front end can poll, to email or notify the user when the job completes, or to push an update over a websocket. None of this is hard, but it is extra plumbing that an inline, synchronous design wouldn't need.
Because the web and worker share the same data stores and often the same codebase, it's easy for them to drift into a single tangled monolith — business logic smeared across both, deployed together, impossible to change in isolation. Keep the boundary deliberate: the contract between them should be the queue message, not a shared pile of internal functions.
After every deploy, a few customers receive two copies of the same export email. What's the most likely cause?
When to use it
Web-Queue-Worker is a natural starting style for simple cloud applications: a web app with some background processing, no microservices sprawl, just two clearly separated halves connected by a queue. It maps cleanly onto managed cloud services and is one of the most common serverless shapes — an HTTP function for the web tier, a queue, and a queue-triggered function for the worker.
Reach for it whenever you have user-facing requests that shouldn't wait on slow work. Outgrow it when the worker starts doing many unrelated kinds of jobs, or when independent teams need to own and deploy pieces separately — at that point you're looking at splitting into proper services rather than one web-plus-worker pair.