When something happens in a system — an order is placed, a user signs up, a file is uploaded — several other parts often need to react. The order should be charged, a confirmation email sent, inventory updated, and analytics recorded.
Publish/subscribe (Pub/Sub) is a messaging pattern that lets the part where the event happened announce it once, and lets every interested part react independently — without any of them having to know about each other.
The problem: one service calling many others
The naive approach is to have the order service call each downstream service itself: save the order, then call the email service, then inventory, then analytics. Every call is synchronous, so the customer's wait is the sum of all of them, and the slowest collaborator sets the pace for checkout. If the mail server takes three seconds to send a receipt, every customer waits three seconds to see “Order placed” — and if the email service is down, placing an order can fail outright.
Worse, the order service now knows about every collaborator. The day you want to add a fraud check, you open the order service, add a call and redeploy it. This is tight coupling: the producer of the event carries the burden of knowing every consumer, and the list of consumers leaks into code that has nothing to do with consuming.
Step through one order below. When the mail server slows down, you'll predict what that does to checkout before you see it.
How it works
Pub/Sub puts a broker in the middle. Instead of calling anyone, the order service publishes a single message — say OrderPlaced — to a named topic on the broker. It doesn't address any recipient; it hands the message over and returns.
Each interested service has a subscription to that topic. Think of a subscription as a private queue the broker keeps for one subscriber. When a message arrives, the broker drops a copy into every subscription, and each subscriber works through its own queue at its own pace. The publisher never learns who the subscribers are or how many there are; the subscribers never know who sent the message. That mutual ignorance is the whole point — it's what decouples them — and the separate queues are what stop one slow subscriber from holding up the rest.
Step through a few orders below. Partway through, the email service goes down: predict what happens before you look.
Adding a consumer is now free for the producer. To introduce fraud detection, you deploy a new service and give it a subscription to the OrderPlaced topic. The order service doesn't change — it's still publishing the same single message. This is the payoff of decoupling: you extend the system by adding subscribers at the edges, not by editing the code at the center.
Your order service publishes OrderPlaced to a topic. A teammate now needs every new order to award loyalty points. What's the smallest change?
Delivery guarantees and semantics
One detail trips up almost everyone: copies are made per subscription, not per service. It stops being academic the day you scale a subscriber. Below, analytics falls behind and you run three instances of it. Predict how many orders each instance receives, then flip between giving each instance its own subscription and letting all three share one.
Decoupling buys flexibility, but it forces you to think about what "delivery" actually promises:
- Fan-out vs. consumer groups — fan-out means every subscription gets its own copy of every message: email, inventory and analytics each see all orders. A consumer group is several instances of one service sharing a subscription, so each message goes to only one of them. That's how you scale a single subscriber horizontally — the competing consumers pattern, much like load balancing spreads work across servers. Give each instance its own subscription by mistake and every message is processed once per instance.
- At-least-once delivery — most brokers guarantee a message arrives at least once, but to survive crashes and retries they may deliver it more than once. Consumers must therefore be idempotent: processing the same message twice should have the same effect as processing it once (e.g. key the email send on the order ID so a duplicate is a no-op).
- Ordering — within a single topic you often can't assume messages arrive in the order they were published, especially once delivery is parallelized. If order matters, you usually need a partition or ordering key so related messages travel the same path.
The trade-offs
Pub/Sub is not free:
- Eventual consistency — because subscribers process asynchronously, there's a window where the order exists but the confirmation email hasn't sent and inventory hasn't updated. The system converges to a correct state, but not instantly.
- Harder debugging and observability — there's no single call stack to follow. A request fans out into independent flows, so tracing "what happened to order #123" means correlating logs across several services and the broker. Distributed tracing and correlation IDs become essential rather than optional.
- Duplicates and ordering land on you — the at-least-once and unordered semantics above mean every consumer has to be written defensively, which is real engineering effort.
- The broker is critical infrastructure — you've concentrated all communication through one component. If the broker is down or backed up, nothing gets delivered, so it has to be highly available, monitored, and capacity-planned with care.
Asynchronous does not mean fire-and-forget-and-stop-caring. Messages can pile up faster than subscribers can drain them, and a consumer that keeps failing will see the same message redelivered forever. Watch consumer lag (how far behind subscribers are) and route messages that fail repeatedly to a dead-letter queue so one poison message can't block the rest.
Customers sometimes get the same receipt twice. The email subscriber runs as a single instance and has no obvious bug. What's the likely cause and fix?
When to reach for it
Pub/Sub shines when a single event has many independent reactions, when you want to add or remove consumers without touching the producer, and when those reactions can happen asynchronously rather than blocking the original request. Event-driven architectures, decoupling microservices, and broadcasting state changes are all natural fits.
It's the wrong tool when the caller genuinely needs an immediate answer — a synchronous request/response is simpler and clearer there. And like caching, it's a trade: you give up the simplicity of a direct call and strict consistency in exchange for decoupling, scalability, and resilience to slow or failing consumers.