TECH
Kafka and Pulsar Diverge on Rebalance Cost at Scale
Kafka and Pulsar both reassign work when brokers join or leave a cluster. They do it in ways that produce different failure shapes. Kafka's consumer group protocol stops every member during a rebalance, and that pause scales with the number of consumers. Pulsar transfers ownership of a topic partition as a metadata operation, and consumers reconnect to the new owner without a group-wide barrier. The divergence is not cosmetic. It decides which metric pages you at 3 a.m. and how long your lag spike lasts.
Rebalance Cost Emerges as the Hidden Scaling Tax
Every streaming platform has to answer one question: who consumes which partition right now? Kafka answers it through a group coordinator that negotiates among consumers. Pulsar answers it through a broker registry that hands out ownership of topic partitions. Both answers become expensive as the cluster grows, but the expense lands in different places.
Kafka's cost grows with consumer group membership. A group of 200 consumers pays a coordination tax on every membership change, because each member must revoke partitions, rejoin, and receive a new assignment. Pulsar's cost grows with topic count and metadata churn, because ownership is tracked per partition rather than per group.
The practical consequence is that a Kafka cluster with many small groups behaves differently from one with a few enormous groups. The second case is where rebalance windows become visible in dashboards. A related piece on this site about service mesh sidecar overhead makes a similar point: fixed costs per hop compound when you have many hops.
Kafka's Eager Rebalance Protocol in Practice
Kafka's original consumer group design, shaped in part by Neha Narkhede's work on the project, uses an eager protocol. The group coordinator bumps the generation, every consumer revokes its partitions, and the group renegotiates from scratch. During that window, no consumer in the group is processing.
Group size drives the duration. Each additional consumer adds a round trip to the join and sync phases. Large groups can see rebalance windows in the tens of seconds, and that estimate is generous when network latency or a slow member stalls the barrier. The coordinator waits for every member, so the slowest participant sets the floor.
Two mitigations reduce the blast radius without changing the protocol. Static membership lets a consumer rejoin under the same identity after a restart, which avoids a generation bump for brief disconnects. The cooperative sticky assignor, introduced later, revokes only the partitions that need to move rather than all of them. Both help. Neither removes the group-wide coordination step.
Pulsar's Ownership Handoff Avoids the Stop
Pulsar assigns ownership of topic partitions to brokers, not to consumer groups. When a broker leaves, the ownership registry updates and a new broker takes over the partition. Consumers connected to the old owner receive a redirect and reconnect to the new one. There is no group-wide barrier because there is no group-level generation to bump.
Managed offerings expose this as a tier rather than a protocol detail. DataStax Astra Streaming, built on Apache Pulsar, sells the ownership model as part of its managed messaging service. The operational pitch is that broker restarts do not cascade into consumer group pauses.
The tradeoff is more moving parts in the metadata layer. Ownership state has to be consistent across brokers, and a slow metadata store becomes a new failure mode. Pulsar pushes complexity into the registry; Kafka pushes it into the group coordinator. Neither is free, and the bill arrives differently. A related piece on Postgres logical replication ordering shows the same pattern in a different system: the coordination layer is where surprises live.
Where the Divergence Bites at Scale
Rolling broker restarts make the split obvious. In Kafka, each restart can trigger a rebalance if membership changes, and a rolling restart across a large cluster produces a sequence of pauses. In Pulsar, each restart triggers an ownership transfer, and consumers reconnect without waiting for peers.
The first visible symptom in either system is consumer lag. In Kafka, lag spikes during the rebalance window and recovers after the group stabilizes. In Pulsar, lag spikes only for consumers whose partition moved, and the spike is shorter because there is no barrier to clear.
Kafka's rebalance cost scales with group size, not partition count. A group of 50 consumers on 500 partitions pays less per rebalance than a group of 500 consumers on 50 partitions. Pulsar's cost scales with topic count and metadata churn, so a cluster with tens of thousands of short-lived topics pays more than one with a few thousand long-lived ones. The shapes are different, and so are the tuning knobs.
Quantifying the Pause: What Metrics Actually Show
Rebalance duration is not a single number. It is a distribution with a long tail, and the tail is what hurts. During a controlled restart of a 200-consumer group, the median pause might land in the low single-digit seconds, but the 99th percentile can stretch into the tens of seconds when a slow member or a network hiccup delays the join phase. That tail is where consumer lag accumulates and where downstream jobs miss their SLAs.
Pulsar's ownership transfer latency shows a different distribution. Because there is no group-wide barrier, the transfer time is dominated by the metadata store round trip and the consumer's reconnect handshake. In healthy clusters, that is often sub-second, but it can spike if the metadata store is under load or if the new owner is already saturated. The key difference is blast radius: a slow transfer affects only the partitions that moved, not every consumer in a group.
These numbers are not universal. They depend on network topology, broker load, and client version. The point is to measure your own distribution, not to assume that one system is always faster. A team that benchmarks only the median will miss the tail that actually pages them.
Choosing Between the Two Models
Kafka wins when consumer groups stay small and stable. A team running a handful of groups with modest membership will rarely notice the rebalance tax, and the ecosystem around Kafka is deeper. The protocol's age is an advantage here: the failure modes are documented, and the tooling reflects years of production use.
Pulsar wins when topics are many and long-lived, and when broker restarts are frequent. The ownership model absorbs churn that would otherwise become group-wide pauses. Teams that run multi-tenant clusters with thousands of topics tend to feel this difference first.
Static membership buys time but not a different protocol. It reduces rebalances from brief disconnects, and it does nothing for a genuine membership change or a broker that stays down. Benchmark with your own restart cadence, not vendor slides. A cluster that restarts monthly and a cluster that restarts daily will produce different answers from the same benchmark.
Practical Moves for Either Stack
Measure rebalance duration during a planned broker restart, not during a quiet period. The number you want is the gap between the first revoke and the last assignment, captured from coordinator logs or consumer metrics. If you cannot find that number, you are guessing about your own system.
Cap consumer group size and shard by key where possible. A group of 200 consumers is a coordination liability; two groups of 100 on separate key ranges is often the same workload with a smaller blast radius. This is the single highest-leverage change for Kafka users.
Enable cooperative rebalancing before scaling group membership, not after. Retrofitting it during an incident is worse than adopting it during a normal sprint. The cooperative sticky assignor changes which partitions move, and that changes how much of the group pauses.
Track ownership transfer latency as a first-class metric in Pulsar, alongside consumer lag. The transfer is the operation that matters, and it is easy to miss because it does not page anyone by default. Add a dashboard panel before you need it.
Revisit the choice when restart frequency changes. A deployment pipeline that moves from weekly to daily restarts changes the economics of both models. The architecture decision that was correct at weekly cadence may not survive the new one. A related piece on DPUs replacing SmartNICs makes the same argument about hardware choices outliving their justification.