Imagine renting a whole delivery van for every single parcel — one van for a letter, another van for a small box, a third for a postcard. Each van costs the same fixed amount whether it's full or nearly empty, and most of them roll out almost empty. The obvious fix is to load many parcels into one van. That's the idea behind Compute Resource Consolidation.
The problem
In the cloud it's tempting to give every small task its own compute instance — one VM or container for the image thumbnailer, another for the nightly report job, another for the email sender. It feels clean, but each instance carries fixed overhead: the OS, the runtime, monitoring agents, and a minimum size you pay for whether or not the task is busy. Since most of these tasks are bursty or light, you end up with a fleet of machines sitting at 5% utilization while the bill reflects 100% provisioning. You're paying for capacity nobody is using, and managing a sprawl of instances to boot.
Step through the bill for ten such services below. Each VM costs, say, $70 a month. Predict how much of the $700 actually pays for work before you look.
How it works
Instead of one instance per task, consolidate compatible tasks onto a shared set of compute units. A single instance (or small pool) hosts several cooperating tasks together, so the fixed overhead of that machine is amortized across all of them and average utilization climbs. The image worker, the report job, and the email sender share the same node, each taking a slice of its CPU and memory. You provision for the combined demand rather than summing each task's worst case in isolation — and because small tasks are rarely all busy at the same moment, that's usually far less hardware.
Step through it below: predict what each node's CPU meter reads once the ten services share two nodes. Then a bad deploy sends one service into a busy loop, and you'll see the price of sharing. Once you've seen it, switch on per-service limits and replay the same bad deploy.
Group tasks that get along. The cost win comes from sharing, and sharing works best for tasks with compatible scaling, lifecycle, and trust: they can be deployed, scaled, and patched together, and none of them needs a security boundary from the others. Keep anything that needs strong isolation — as the bulkhead pattern would dictate — on its own instance.
You packed 12 services onto 3 nodes with no resource limits. One service starts leaking memory, and soon services all over its node are being killed and restarted. What would have contained it?
Sharing safely
A shared node needs rules, the way a shared flat needs a rota. Three are standard, and container platforms and orchestrators support all of them:
- Limits cap what one task can take: no more than this much CPU, no more than this much memory. A runaway task hits its own ceiling and slows down, or is restarted, instead of starving its neighbors.
- Reservations (often called requests) guarantee the minimum a task needs, so the scheduler only places it on a node with that much room to spare.
- Quotas cap a whole team or namespace, so one group can't fill the shared pool with its own workloads.
Then watch the node as a whole. Consolidation turns ten quiet machines into a couple of busy ones, so a node that fails now takes several tasks with it. Run at least two, keep some headroom for bursts, and alert on sustained high CPU or memory before your users notice.
Don't size by the average. "Ten tasks at 5% each fit in half a machine" is only true if they're busy at different times. Ten report jobs that each average 5% over the week, but do all of it flat out between 9:00 and 17:00 on Monday, need ten machines' worth of CPU for those eight hours, all at once. Look at when each task is busy, consolidate the ones whose peaks don't line up, and size each node for the combined peak.
When to use it
Consolidation pays off when you have many small, light, or bursty tasks whose individual instances would mostly sit idle, and where those tasks have similar scaling and lifecycle needs and a comparable trust level. It complements elastic scaling and competing consumers by making each shared node do real work, and it's a manual cousin of serverless, which consolidates for you behind the scenes.
Don't consolidate tasks that genuinely conflict: those with very different scaling curves, hard isolation or security boundaries, or peaks that land at the same time and would force you to provision for the worst of all worlds. When in doubt, isolate what must be isolated and pool the rest.
Ten report jobs each average 5% CPU over the week on their own VMs, but all of that work happens on Monday between 9:00 and 17:00, when each one runs flat out. You plan to put all ten on one VM, since 10 × 5% = 50%. What's wrong with the plan?