From a60053fd42b63f606389025a03165f286afcb001 Mon Sep 17 00:00:00 2001 From: Sam Barker Date: Mon, 21 Sep 2026 14:02:39 +1200 Subject: [PATCH 1/4] Gemini first draft. Signed-off-by: Sam Barker --- _drafts/routing-post1.md | 139 +++++++++++++++++++++++++++++++++++++++ 1 file changed, 139 insertions(+) create mode 100644 _drafts/routing-post1.md diff --git a/_drafts/routing-post1.md b/_drafts/routing-post1.md new file mode 100644 index 0000000..35d8cd0 --- /dev/null +++ b/_drafts/routing-post1.md @@ -0,0 +1,139 @@ +--- +layout: post +title: "Kroxylicious release 0.24.0" +date: 2026-09-04 15:00:00 +1200 +author: "Sam Barker" +author_url: "https://github.com/sambarker" +# noinspection YAMLSchemaValidation +categories: blog kroxylicious-proxy +tags: [ "routing" ] +--- + +# Layer 7 Kafka Routing: A Pattern Taxonomy for the Consolidation Minefield + +If you've ever tried to migrate a live Kafka workload between clusters, consolidate regional deployments, or shift cloud providers, you already know the bottom line: Kafka clients are remarkably opinionated about network topology. They don't just talk to a virtual endpoint; they demand exact broker metadata, explicit partition assignments, and direct TCP connections to specific physical nodes. + +When you try to reshape that underlying physical infrastructure without breaking application teams, you usually end up picking which operational headache you dislike the least. + +## Pick your poison: Dual writes, MirrorMaker, or scheduled downtime + +Before looking at proxy-level patterns, it helps to review the standard tools people use when trying to move or consolidate Kafka traffic—and why they so often result in late-night incident reviews. + +* **Application-level dual writes:** You ask application teams to update their producer code to write to both Cluster A and Cluster B simultaneously. In architectural diagrams, this looks clean. In production, network blips cause asymmetric failures, message ordering drifts instantly, and handling duplicate delivery becomes the application's problem. Worse, getting twenty product teams to deploy matching code changes on the same timeline is an exercise in cat-herding. +* **Replication pipelines (MirrorMaker 2, etc.):** Running an intermediary replication cluster works reasonably well for asynchronous backup, but relying on it for live consolidation adds latency, doubles your storage and network bills, and forces consumers to deal with offset translation. You're essentially running twice as much hardware to move bytes you already owned. +* **Hard cutovers and maintenance windows:** You schedule a Sunday 2:00 AM window, drain topic queues, update DNS records, and restart clients. This is conceptually simple right up until a legacy service ignores DNS TTLs, holds onto stale socket connections indefinitely, and drops messages silently when the old brokers finally go dark. + +Intercepting traffic at Layer 7—the Kafka wire protocol itself—offers an alternative. By placing a proxy like Kroxylicious between clients and brokers, we can manipulate metadata and route requests on the fly. + +To be clear: introducing an L7 proxy adds a hop, consumes CPU, and gives you another piece of infrastructure to manage. If a simple DNS CNAME flip actually solves your problem, do that instead. But when you need to decouple physical cluster topologies from what clients see, proxying gives you control back. + +--- + +## Where does the proxy run? Forward, Reverse, and Sidecars + +Choosing a routing pattern is only half the battle. You also have to decide where the proxy tier physically lives, who owns it, and whether it acts as a forward proxy for egress or a reverse proxy for ingress. + +``` + ┌─────────────────────────────────────────────────────────┐ + │ FORWARD PROXY │ + │ (Client-Side Egress) │ + │ ┌───────────────────────┐ ┌───────────────────────┐ │ + │ │ Client App + Pod │ │ Client Cluster │ │ + │ │ Sidecar Proxy │ │ Centralized Gateway │ │ + │ └───────────┬───────────┘ └───────────┬───────────┘ │ + └──────────────┼───────────────────────────┼──────────────┘ + │ │ + NETWORK BOUNDARY / TRANSIT / VPC PEERING │ + │ │ + ┌──────────────┼───────────────────────────┼──────────────┐ + │ ▼ ▼ │ + │ ┌───────────────────────┐ ┌───────────────────────┐ │ + │ │ Broker Cluster │ │ Broker Node + │ │ + │ │ Ingress Gateway │ │ Broker Sidecar Proxy │ │ + │ └───────────────────────┘ └───────────────────────┘ │ + │ REVERSE PROXY │ + │ (Broker-Side Ingress) │ + └└─────────────────────────────────────────────────────────┘ + +``` + +### The Forward Proxy Model (Client-Side Egress) + +Lives in the client's network boundary and is managed by application or client-platform teams. + +* **Client Cluster Gateway:** A shared proxy fleet inside the client Kubernetes cluster or VPC. Applications point to a local gateway service, and the proxy handles cross-cluster egress across network boundaries. +* **App Pod Sidecar:** Co-located inside the application pod as an egress proxy. Crucially, the proxy does *not* flatten or hide the Kafka cluster model—the client driver still receives metadata mapped to local endpoints (e.g., port ranges on `localhost`) and maintains individual TCP sockets per broker. You get isolated blast radius per pod, but running hundreds of Netty/JVM proxy containers across a microservice fleet levies a noticeable baseline memory tax. + +### The Reverse Proxy Model (Broker-Side Ingress) + +Lives in the Kafka cluster's network boundary and is managed by the central infrastructure/platform team. + +* **Broker Cluster Ingress Gateway:** A shared ingress fleet sitting in front of physical Kafka brokers. Provides a unified front door for incoming client connections and shields underlying cluster topology, though a gateway outage impacts all incoming traffic. +* **Broker Node Sidecar:** Co-located on the actual physical broker hardware (one proxy instance per broker node). While it eliminates an internal network hop on paper, almost nobody does this in production. You risk L7 proxy bugs or memory spikes starving the underlying broker JVM or page cache, and a broker-side proxy loses most of its cross-cluster routing superpowers anyway. + +--- + +## The four L7 traffic patterns + +With deployment boundaries established, here are the four architectural patterns we use to manipulate Kafka traffic at Layer 7. + +``` +┌─────────────────────────────────────────────────────────────────┐ +│ Kroxylicious L7 Proxy │ +├─────────────────┬─────────────────┬──────────────┬──────────────┤ +│ Cluster Aliasing│ Union Clusters │ Union Topics │ Topology- │ +│ │ │ │ Aware │ +│ [Virtual A] │ [Virtual Hub] │ [Logical T] │ [AZ-a Client]│ +│ │ │ ┌─────┴─────┐ │ ┌─────┴───┐ │ │ │ +│ ▼ │ ▼ ▼ │ ▼ ▼ │ ▼ │ +│ [Physical A/B] │ [Phys 1] [Phys 2]│ [P0-1] [P2-3]│ [Broker-a] │ +└─────────────────┴─────────────────┴──────────────┴──────────────┤ + │ + ▼ │ +┌─────────────────────────────────────────────────────────────────┐ +│ Physical Kafka Clusters │ +└─────────────────────────────────────────────────────────────────┘ + +``` + +### 1. Cluster Aliasing + +**What it is:** Mapping a static virtual cluster endpoint to actual physical backend clusters, allowing you to switch the target backend at the proxy layer without reconfiguring or restarting clients. + +**How it works:** Clients connect to `kafka-virtual.company.internal`. The proxy inspects incoming Kafka frames (`Produce`, `Fetch`, `Metadata`) and maps them to physical brokers in `cluster-blue`. When you want to migrate to `cluster-green`, you update the proxy's routing target. The proxy handles the connection handoff to the new brokers under the hood. + +**Real-world caveat:** Swapping backend targets at the proxy layer solves client connectivity, but it doesn't magically sync topic data or offset state between backends. If you flip the pointer without state replication, your consumers will hit offset mismatches. (We'll cover how we pair this with byte-level replication and KIP-1279 in Post 2). + +### 2. Union Clusters + +**What it is:** Exposing multiple distinct backend physical clusters through a single virtual cluster endpoint. + +**How it works:** The proxy intercepts `Metadata` requests and synthesizes a single, unified cluster layout for the client. To an incoming producer or consumer, it looks like one massive cluster. Behind the scenes, the proxy routes requests for `orders-*` topics to a high-throughput physical cluster, while `analytics-*` topics head to a cheaper, storage-optimized cluster. + +**Real-world caveat:** Namespace collisions will ruin your day. If `orders-v1` exists on two backend clusters, the proxy has to decide which physical cluster wins or enforce explicit topic-prefix rules. Also, we haven't benchmarked proxy metadata synthesis overhead at tens of thousands of topics across dozens of physical backends yet, so expect memory usage to scale with metadata volume. + +### 3. Union Topics + +**What it is:** Presenting a single logical Kafka topic to clients while sharding its underlying partitions across multiple physical clusters. + +**How it works:** A client asks for metadata for topic `events`, which appears to have 32 partitions. The proxy returns metadata where partitions 0–15 point to physical brokers in Cluster A, and partitions 16–31 point to physical brokers in Cluster B. When a client sends a `ProduceRequest` for partition 20, the proxy routes those specific frames to Cluster B. + +**Real-world caveat:** Transaction coordinator boundaries stop working cleanly here. If your applications rely on multi-topic transactions or read-committed isolation levels across partitions, splitting a topic across physical cluster boundaries breaks those guarantees today. Treat this pattern as a fit for simple, un-keyed, or independently partition-keyed workloads until proxy transaction handling matures. + +### 4. Topology-Aware Routing + +**What it is:** Directing client traffic dynamically based on client metadata, network topology, or locality attributes (like cloud availability zones). + +**How it works:** Kafka cross-AZ data transfer fees are a recurring budget surprise for infra teams. With topology-aware routing, the proxy reads client IP blocks or rack attributes and routes fetch requests to brokers or read-replicas inside the same availability zone, cutting down on inter-zone bandwidth costs. + +**Real-world caveat:** Local routing savings disappear if your proxy fleet is deployed inefficiently. If a client in `us-east-1a` sends frames to a proxy instance running in `us-east-1b`, which then forwards bytes to a broker in `us-east-1a`, you've just doubled your cross-AZ costs instead of eliminating them. Proxy placement must match client topology. + +--- + +## What's next + +This taxonomy gives us a common vocabulary for describing how L7 proxies reshape Kafka traffic and where they fit into physical deployment topologies. Over the coming weeks leading up to Current in San Francisco, we're going to dive deeper into the actual implementations. + +Next week in Post 2, we'll take a close look at **Active-Passive DR with KIP-1279 & Cluster Aliasing**, breaking down how Virtual Cluster Keys swap backend targets without dropping client sockets, and showing the code behind our live demo. + +If you're playing with Kafka traffic routing, building custom extensions, or just want to tell us where our architecture assumptions are wrong, drop by our [GitHub](https://github.com/kroxylicious/kroxylicious?utm_source=gemini), join us on [Slack](https://www.google.com/search?q=https://kroxylicious.slack.com&utm_source=gemini), or find us on [Bluesky](https://bsky.app/profile/kroxylicious.io?utm_source=gemini). From 75737d141898bd71ae37adbcd051c8fe74a1cebf Mon Sep 17 00:00:00 2001 From: Sam Barker Date: Mon, 21 Sep 2026 14:23:43 +1200 Subject: [PATCH 2/4] Bob's rewrite Signed-off-by: Sam Barker --- _drafts/routing-post1.md | 129 ++++++++++----------------------------- 1 file changed, 32 insertions(+), 97 deletions(-) diff --git a/_drafts/routing-post1.md b/_drafts/routing-post1.md index 35d8cd0..d01dea5 100644 --- a/_drafts/routing-post1.md +++ b/_drafts/routing-post1.md @@ -1,6 +1,6 @@ --- layout: post -title: "Kroxylicious release 0.24.0" +title: "These are not the brokers you are connecting to" date: 2026-09-04 15:00:00 +1200 author: "Sam Barker" author_url: "https://github.com/sambarker" @@ -9,131 +9,66 @@ categories: blog kroxylicious-proxy tags: [ "routing" ] --- -# Layer 7 Kafka Routing: A Pattern Taxonomy for the Consolidation Minefield +Back in May, Tom wrote about [a proof of concept for routing]({% post_url 2026-05-21-topic-routing %}) — Kafka clients producing and consuming from topics scattered across multiple clusters, with no idea that's what they're doing. The proxy handles the fan-out, the response merging, the session state. The client just sees topics. -If you've ever tried to migrate a live Kafka workload between clusters, consolidate regional deployments, or shift cloud providers, you already know the bottom line: Kafka clients are remarkably opinionated about network topology. They don't just talk to a virtual endpoint; they demand exact broker metadata, explicit partition assignments, and direct TCP connections to specific physical nodes. +That POC has been quietly maturing. We're now in the process of turning it into something you can actually run in production, and in a few weeks we'll be talking about it at [Current in San Francisco](https://current.confluent.io/). The talk is called *"These are not the brokers you are connecting to"*, which felt like the honest description of what the proxy is doing. -When you try to reshape that underlying physical infrastructure without breaking application teams, you usually end up picking which operational headache you dislike the least. +This post is the first in a short series leading up to that talk. The goal here is to establish some vocabulary — I've been thinking about the routing design space and I'm convinced there are a handful of distinct patterns worth naming separately, because each one carries different tradeoffs and breaks down in different ways. The names are mine and I'll happily accept better ones, but I think having *something* to call them makes it easier to reason about the problems without conflating them. -## Pick your poison: Dual writes, MirrorMaker, or scheduled downtime +The deeper dives on each pattern will follow in subsequent posts. -Before looking at proxy-level patterns, it helps to review the standard tools people use when trying to move or consolidate Kafka traffic—and why they so often result in late-night incident reviews. +## Pick your poison -* **Application-level dual writes:** You ask application teams to update their producer code to write to both Cluster A and Cluster B simultaneously. In architectural diagrams, this looks clean. In production, network blips cause asymmetric failures, message ordering drifts instantly, and handling duplicate delivery becomes the application's problem. Worse, getting twenty product teams to deploy matching code changes on the same timeline is an exercise in cat-herding. -* **Replication pipelines (MirrorMaker 2, etc.):** Running an intermediary replication cluster works reasonably well for asynchronous backup, but relying on it for live consolidation adds latency, doubles your storage and network bills, and forces consumers to deal with offset translation. You're essentially running twice as much hardware to move bytes you already owned. -* **Hard cutovers and maintenance windows:** You schedule a Sunday 2:00 AM window, drain topic queues, update DNS records, and restart clients. This is conceptually simple right up until a legacy service ignores DNS TTLs, holds onto stale socket connections indefinitely, and drops messages silently when the old brokers finally go dark. +If you've tried to migrate a live Kafka workload between clusters, consolidate regional deployments, or shift cloud providers, you know the bottom line: Kafka clients are remarkably opinionated about where their brokers live. They don't just talk to a virtual endpoint — they demand exact broker metadata, explicit partition assignments, and direct TCP connections to specific physical nodes. You can't just update a load balancer and go home. -Intercepting traffic at Layer 7—the Kafka wire protocol itself—offers an alternative. By placing a proxy like Kroxylicious between clients and brokers, we can manipulate metadata and route requests on the fly. +So when you need to reshape the underlying infrastructure without breaking application teams, you usually end up picking which operational headache you dislike the least. -To be clear: introducing an L7 proxy adds a hop, consumes CPU, and gives you another piece of infrastructure to manage. If a simple DNS CNAME flip actually solves your problem, do that instead. But when you need to decouple physical cluster topologies from what clients see, proxying gives you control back. +**Application-level dual writes** look clean in architecture diagrams. In production, network blips cause asymmetric failures, message ordering drifts, and duplicate delivery becomes every application team's problem. Getting twenty teams to deploy matching code changes on the same timeline is an exercise in cat-herding that usually ends with someone's Friday afternoon becoming someone else's Saturday morning. ---- - -## Where does the proxy run? Forward, Reverse, and Sidecars - -Choosing a routing pattern is only half the battle. You also have to decide where the proxy tier physically lives, who owns it, and whether it acts as a forward proxy for egress or a reverse proxy for ingress. - -``` - ┌─────────────────────────────────────────────────────────┐ - │ FORWARD PROXY │ - │ (Client-Side Egress) │ - │ ┌───────────────────────┐ ┌───────────────────────┐ │ - │ │ Client App + Pod │ │ Client Cluster │ │ - │ │ Sidecar Proxy │ │ Centralized Gateway │ │ - │ └───────────┬───────────┘ └───────────┬───────────┘ │ - └──────────────┼───────────────────────────┼──────────────┘ - │ │ - NETWORK BOUNDARY / TRANSIT / VPC PEERING │ - │ │ - ┌──────────────┼───────────────────────────┼──────────────┐ - │ ▼ ▼ │ - │ ┌───────────────────────┐ ┌───────────────────────┐ │ - │ │ Broker Cluster │ │ Broker Node + │ │ - │ │ Ingress Gateway │ │ Broker Sidecar Proxy │ │ - │ └───────────────────────┘ └───────────────────────┘ │ - │ REVERSE PROXY │ - │ (Broker-Side Ingress) │ - └└─────────────────────────────────────────────────────────┘ +**Replication pipelines** (MirrorMaker 2, et al.) work fine for asynchronous backup, but for live consolidation you're adding latency, doubling storage and network costs, and asking consumers to deal with offset translation. You're running twice as much hardware to move bytes you already owned. -``` +**Maintenance windows** are conceptually simple right up until a legacy service ignores DNS TTLs, holds onto stale sockets, and drops messages silently when the old brokers finally go dark. Sunday 2am is a fine time for this to happen. -### The Forward Proxy Model (Client-Side Egress) +## What a Layer 7 proxy buys you -Lives in the client's network boundary and is managed by application or client-platform teams. +Intercepting at Layer 7 — the Kafka wire protocol itself — is the alternative. The proxy sits between clients and brokers, inspects and rewrites Kafka frames, and presents whatever cluster topology it wants to clients regardless of what's actually behind it. Clients connect to a virtual cluster. What they get told about that cluster is up to us. -* **Client Cluster Gateway:** A shared proxy fleet inside the client Kubernetes cluster or VPC. Applications point to a local gateway service, and the proxy handles cross-cluster egress across network boundaries. -* **App Pod Sidecar:** Co-located inside the application pod as an egress proxy. Crucially, the proxy does *not* flatten or hide the Kafka cluster model—the client driver still receives metadata mapped to local endpoints (e.g., port ranges on `localhost`) and maintains individual TCP sockets per broker. You get isolated blast radius per pod, but running hundreds of Netty/JVM proxy containers across a microservice fleet levies a noticeable baseline memory tax. +To be clear about the tradeoffs: an L7 proxy adds a network hop, consumes CPU for frame inspection and rewriting, and gives you another piece of infrastructure to keep alive. If a DNS CNAME swap actually solves your problem, do that instead. But when you need to decouple what clients see from how your physical infrastructure is actually laid out, this is the lever. -### The Reverse Proxy Model (Broker-Side Ingress) - -Lives in the Kafka cluster's network boundary and is managed by the central infrastructure/platform team. - -* **Broker Cluster Ingress Gateway:** A shared ingress fleet sitting in front of physical Kafka brokers. Provides a unified front door for incoming client connections and shields underlying cluster topology, though a gateway outage impacts all incoming traffic. -* **Broker Node Sidecar:** Co-located on the actual physical broker hardware (one proxy instance per broker node). While it eliminates an internal network hop on paper, almost nobody does this in production. You risk L7 proxy bugs or memory spikes starving the underlying broker JVM or page cache, and a broker-side proxy loses most of its cross-cluster routing superpowers anyway. - ---- +## Four patterns (working names) -## The four L7 traffic patterns +After reviewing the [routing design proposal](https://github.com/kroxylicious/design/pull/70) and thinking through the problem space, I think there are four distinct patterns worth naming. They're not mutually exclusive and they compose, but each one has its own failure modes, so it helps to be able to talk about them separately. -With deployment boundaries established, here are the four architectural patterns we use to manipulate Kafka traffic at Layer 7. +### Cluster aliasing -``` -┌─────────────────────────────────────────────────────────────────┐ -│ Kroxylicious L7 Proxy │ -├─────────────────┬─────────────────┬──────────────┬──────────────┤ -│ Cluster Aliasing│ Union Clusters │ Union Topics │ Topology- │ -│ │ │ │ Aware │ -│ [Virtual A] │ [Virtual Hub] │ [Logical T] │ [AZ-a Client]│ -│ │ │ ┌─────┴─────┐ │ ┌─────┴───┐ │ │ │ -│ ▼ │ ▼ ▼ │ ▼ ▼ │ ▼ │ -│ [Physical A/B] │ [Phys 1] [Phys 2]│ [P0-1] [P2-3]│ [Broker-a] │ -└─────────────────┴─────────────────┴──────────────┴──────────────┤ - │ - ▼ │ -┌─────────────────────────────────────────────────────────────────┐ -│ Physical Kafka Clusters │ -└─────────────────────────────────────────────────────────────────┘ +A static virtual cluster endpoint maps to one physical backend, but the mapping can be changed at the proxy layer without touching clients. -``` +Clients connect to `kafka-virtual.company.internal`. The proxy maps their connections to `cluster-blue`. When you want to migrate to `cluster-green`, you update the proxy config. Clients never reconnect, never need new bootstrap addresses, never need to know a migration happened. -### 1. Cluster Aliasing +The catch is that this only solves connectivity. Flipping the pointer without replicating state means consumers hit offset mismatches on the new cluster. The routing layer isn't magic — it's just decoupling one specific piece of the problem. In Post 2 we'll look at pairing this with KIP-1279 to handle the state problem, which makes this pattern actually useful for DR rather than just theoretically interesting. -**What it is:** Mapping a static virtual cluster endpoint to actual physical backend clusters, allowing you to switch the target backend at the proxy layer without reconfiguring or restarting clients. +### Union clusters -**How it works:** Clients connect to `kafka-virtual.company.internal`. The proxy inspects incoming Kafka frames (`Produce`, `Fetch`, `Metadata`) and maps them to physical brokers in `cluster-blue`. When you want to migrate to `cluster-green`, you update the proxy's routing target. The proxy handles the connection handoff to the new brokers under the hood. +Multiple distinct physical clusters exposed through a single virtual cluster endpoint. The proxy synthesises a unified metadata view — to a producer or consumer it looks like one cluster. Behind the scenes, `orders-*` topics route to a high-throughput cluster, `analytics-*` topics go to a cheaper storage-optimised one. -**Real-world caveat:** Swapping backend targets at the proxy layer solves client connectivity, but it doesn't magically sync topic data or offset state between backends. If you flip the pointer without state replication, your consumers will hit offset mismatches. (We'll cover how we pair this with byte-level replication and KIP-1279 in Post 2). +The obvious failure mode is namespace collisions. If `orders-v1` exists on two backends, the proxy has to pick a winner or enforce prefix rules. We haven't pushed the metadata synthesis hard enough at scale yet to characterise the memory behaviour at tens of thousands of topics across many backends — worth flagging if that's your situation. -### 2. Union Clusters +### Union topics -**What it is:** Exposing multiple distinct backend physical clusters through a single virtual cluster endpoint. +A single logical topic whose partitions are sharded across multiple physical clusters. A client asks for metadata for `events` with 32 partitions; partitions 0–15 live on Cluster A, 16–31 on Cluster B. A `ProduceRequest` for partition 20 gets routed to Cluster B transparently. -**How it works:** The proxy intercepts `Metadata` requests and synthesizes a single, unified cluster layout for the client. To an incoming producer or consumer, it looks like one massive cluster. Behind the scenes, the proxy routes requests for `orders-*` topics to a high-throughput physical cluster, while `analytics-*` topics head to a cheaper, storage-optimized cluster. +The hard limit today: transaction coordinator boundaries break across physical clusters. If you're relying on multi-topic transactions or read-committed isolation, splitting a topic across cluster boundaries removes those guarantees. This pattern is a good fit for simple, independently partition-keyed workloads. For anything transactional, wait. -**Real-world caveat:** Namespace collisions will ruin your day. If `orders-v1` exists on two backend clusters, the proxy has to decide which physical cluster wins or enforce explicit topic-prefix rules. Also, we haven't benchmarked proxy metadata synthesis overhead at tens of thousands of topics across dozens of physical backends yet, so expect memory usage to scale with metadata volume. +### Topology-aware routing -### 3. Union Topics +Traffic directed dynamically based on client metadata or locality — the main use case being cross-AZ bandwidth costs, which have a habit of appearing as a surprise line item in cloud bills. The proxy reads rack attributes or client IP blocks and routes fetch requests to brokers inside the same availability zone. -**What it is:** Presenting a single logical Kafka topic to clients while sharding its underlying partitions across multiple physical clusters. - -**How it works:** A client asks for metadata for topic `events`, which appears to have 32 partitions. The proxy returns metadata where partitions 0–15 point to physical brokers in Cluster A, and partitions 16–31 point to physical brokers in Cluster B. When a client sends a `ProduceRequest` for partition 20, the proxy routes those specific frames to Cluster B. - -**Real-world caveat:** Transaction coordinator boundaries stop working cleanly here. If your applications rely on multi-topic transactions or read-committed isolation levels across partitions, splitting a topic across physical cluster boundaries breaks those guarantees today. Treat this pattern as a fit for simple, un-keyed, or independently partition-keyed workloads until proxy transaction handling matures. - -### 4. Topology-Aware Routing - -**What it is:** Directing client traffic dynamically based on client metadata, network topology, or locality attributes (like cloud availability zones). - -**How it works:** Kafka cross-AZ data transfer fees are a recurring budget surprise for infra teams. With topology-aware routing, the proxy reads client IP blocks or rack attributes and routes fetch requests to brokers or read-replicas inside the same availability zone, cutting down on inter-zone bandwidth costs. - -**Real-world caveat:** Local routing savings disappear if your proxy fleet is deployed inefficiently. If a client in `us-east-1a` sends frames to a proxy instance running in `us-east-1b`, which then forwards bytes to a broker in `us-east-1a`, you've just doubled your cross-AZ costs instead of eliminating them. Proxy placement must match client topology. +The failure mode here is proxy placement. A client in `us-east-1a` hitting a proxy in `us-east-1b` that forwards to a broker in `us-east-1a` doubles your cross-AZ cost instead of eliminating it. The proxy fleet has to be co-located with clients for the locality savings to materialise. --- -## What's next - -This taxonomy gives us a common vocabulary for describing how L7 proxies reshape Kafka traffic and where they fit into physical deployment topologies. Over the coming weeks leading up to Current in San Francisco, we're going to dive deeper into the actual implementations. +These four patterns give us a working vocabulary for the rest of the series. They're not all at the same maturity level — cluster aliasing and union topics are closest to production-ready, topology-aware routing is further out — and the posts that follow will be honest about where the code is versus where the design points. -Next week in Post 2, we'll take a close look at **Active-Passive DR with KIP-1279 & Cluster Aliasing**, breaking down how Virtual Cluster Keys swap backend targets without dropping client sockets, and showing the code behind our live demo. +Next up: **Active-Passive DR with cluster aliasing and KIP-1279**. Fair warning: this one is going to live as a POC on a branch rather than shipped code, but it's what we're demoing at Current and the mechanics are worth understanding on their own terms. -If you're playing with Kafka traffic routing, building custom extensions, or just want to tell us where our architecture assumptions are wrong, drop by our [GitHub](https://github.com/kroxylicious/kroxylicious?utm_source=gemini), join us on [Slack](https://www.google.com/search?q=https://kroxylicious.slack.com&utm_source=gemini), or find us on [Bluesky](https://bsky.app/profile/kroxylicious.io?utm_source=gemini). +If you're thinking about Kafka topology problems, have a use case that doesn't fit cleanly into any of these patterns, or want to tell us where the vocabulary breaks down — find us on [GitHub](https://github.com/kroxylicious/kroxylicious), [Slack](https://kroxylicious.slack.com), or [Bluesky](https://bsky.app/profile/kroxylicious.io). From dd141aeeda81c4083deda270272a766745aa7039 Mon Sep 17 00:00:00 2001 From: Sam Barker Date: Thu, 24 Sep 2026 11:52:50 +1200 Subject: [PATCH 3/4] Rewrite for voice and detail Signed-off-by: Sam Barker --- _drafts/routing-post1.md | 59 ++++++++++++++++++++++++---------------- 1 file changed, 35 insertions(+), 24 deletions(-) diff --git a/_drafts/routing-post1.md b/_drafts/routing-post1.md index d01dea5..2cf5175 100644 --- a/_drafts/routing-post1.md +++ b/_drafts/routing-post1.md @@ -11,64 +11,75 @@ tags: [ "routing" ] Back in May, Tom wrote about [a proof of concept for routing]({% post_url 2026-05-21-topic-routing %}) — Kafka clients producing and consuming from topics scattered across multiple clusters, with no idea that's what they're doing. The proxy handles the fan-out, the response merging, the session state. The client just sees topics. -That POC has been quietly maturing. We're now in the process of turning it into something you can actually run in production, and in a few weeks we'll be talking about it at [Current in San Francisco](https://current.confluent.io/). The talk is called *"These are not the brokers you are connecting to"*, which felt like the honest description of what the proxy is doing. +That POC has been quietly maturing. We're now in the process of turning it into something you can actually run in production, and in a few weeks we'll be talking about it at [Current in San Francisco](https://current.confluent.io/san-francisco/sessions#session-SESS-165). In a session called *"These are not the brokers you are connecting to"* — my employer does not condone lying, but the proxy will quite happily tell your clients whatever you need it to. -This post is the first in a short series leading up to that talk. The goal here is to establish some vocabulary — I've been thinking about the routing design space and I'm convinced there are a handful of distinct patterns worth naming separately, because each one carries different tradeoffs and breaks down in different ways. The names are mine and I'll happily accept better ones, but I think having *something* to call them makes it easier to reason about the problems without conflating them. +This post is the first in a short series leading up to that talk. The goal here is to establish some vocabulary and get into the gory details I won't have time for onstage. I've been thinking about the design space routing enables and I'm convinced there are a handful of distinct patterns worth naming separately. There are no hard boundaries between the patterns, they all leverage the same underlying infrastructure, but they are distinct because they trade off different aspects of the problem space. The names are mine and while I'll happily accept better ones, we need something to start a conversation (plus I'm right :wink: ). The deeper dives on each pattern will follow in subsequent posts. ## Pick your poison -If you've tried to migrate a live Kafka workload between clusters, consolidate regional deployments, or shift cloud providers, you know the bottom line: Kafka clients are remarkably opinionated about where their brokers live. They don't just talk to a virtual endpoint — they demand exact broker metadata, explicit partition assignments, and direct TCP connections to specific physical nodes. You can't just update a load balancer and go home. +If you've tried to migrate a live Kafka workload between clusters, consolidate regional deployments, or shift cloud providers, you will have run into the same core problem: Kafka is not HTTP-based (stop rolling your eyes at the back). Kafka clients are genuinely indifferent to geography — they'll talk to a broker in the same pod just as happily as one in Timbuktu. What they care about, deeply and lastingly, is *which* broker they're talking to. Broker 1 is Broker 1. You can't quietly swap it out, move it between availability zones, or retire it without a lot of fuss. The load balancer doesn't get a say. Most proxies can only manage Kafka as a layer 4 protocol — and layer 4 is BOOOOORRRRING (I will not listen to arguments from the network engineer in the corner). So when you need to reshape the underlying infrastructure without breaking application teams, you usually end up picking which operational headache you dislike the least. -**Application-level dual writes** look clean in architecture diagrams. In production, network blips cause asymmetric failures, message ordering drifts, and duplicate delivery becomes every application team's problem. Getting twenty teams to deploy matching code changes on the same timeline is an exercise in cat-herding that usually ends with someone's Friday afternoon becoming someone else's Saturday morning. +**Application-level dual writes** look clean in architecture diagrams. In production, network blips cause asymmetric failures, message ordering drifts, and duplicate delivery becomes every application team's problem. Getting twenty teams to deploy matching code changes on the same timeline is an exercise in cat-herding that usually ends with someone's Friday afternoon becoming someone else's out of hours page. -**Replication pipelines** (MirrorMaker 2, et al.) work fine for asynchronous backup, but for live consolidation you're adding latency, doubling storage and network costs, and asking consumers to deal with offset translation. You're running twice as much hardware to move bytes you already owned. +**Replication pipelines** (MirrorMaker 2, et al.) work great for asynchronous backup, but for live consolidation knowing you are in sync is very difficult question, not to mention doubling storage and network costs, and asking consumers to deal with offset translation. You're running twice as much hardware to move bytes you already owned, fine if the migration ever actually ends. -**Maintenance windows** are conceptually simple right up until a legacy service ignores DNS TTLs, holds onto stale sockets, and drops messages silently when the old brokers finally go dark. Sunday 2am is a fine time for this to happen. +**Maintenance windows** are conceptually simple right up until a legacy service ignores DNS TTLs, holds onto stale sockets, and drops messages silently when the old brokers finally go dark. Sunday 2am is when the core reporting pipeline splutters to a halt. ## What a Layer 7 proxy buys you -Intercepting at Layer 7 — the Kafka wire protocol itself — is the alternative. The proxy sits between clients and brokers, inspects and rewrites Kafka frames, and presents whatever cluster topology it wants to clients regardless of what's actually behind it. Clients connect to a virtual cluster. What they get told about that cluster is up to us. +Jumping a few layers up the OSI stack to Layer 7 — the Kafka wire protocol itself — changes what's possible. Rather than blindly forwarding bytes, the proxy can inspect, rewrite, and route individual Kafka frames. It knows a `Metadata` response from a `Produce` request. It can answer the client's question about where the brokers are with whatever answer it likes. -To be clear about the tradeoffs: an L7 proxy adds a network hop, consumes CPU for frame inspection and rewriting, and gives you another piece of infrastructure to keep alive. If a DNS CNAME swap actually solves your problem, do that instead. But when you need to decouple what clients see from how your physical infrastructure is actually laid out, this is the lever. +Kroxylicious has used the concept of a Virtual Kafka Cluster (VKC) from the start — it's the endpoint clients connect to. What building routing support gave us was a breakthrough — the VKC is more than a networking construct. It's a stable, client-facing identity that we fully control. The brokers behind it can change. The topology behind it can change. The client doesn't need to know. -## Four patterns (working names) +That shift — from "the VKC is where you point your bootstrap servers" to "the VKC is the contract between the client and whatever we've decided is behind it" — is what opens up the design space this series is exploring. The posts that follow get into the mechanics of each pattern and are honest about where the implementation is today versus where the design points. The talk covers the highlights; the blog is where the details live. + +We get it, adding a proxy is operational overhead and we wish as much as you do swapping a CNAME from site A to site B was enough — but as far as Kafka clients are concerned, that's a divorce and why we are all here. + +## Name them we must + +After reviewing the [routing design proposal](https://github.com/kroxylicious/design/pull/70) and thinking through the problem space, I think there are multiple distinct deployment patterns worth naming. As with all patterns the boundaries are fuzzy and often more than one applies at once. The thing that convinces me they're real is that they have descriptive power — each one has its own tradeoffs, failure modes, and a distinct place on the spectrum from the clear(ish) waters of cluster aliasing to the mangrove swamp of virtual topics. Like all good abstractions, they're useful. + +If you're reaching for the [Enterprise Integration Patterns](https://www.enterpriseintegrationpatterns.com/patterns/messaging/MessageRoutingIntro.html) book right now — well done for being as old as I am, the rest of you just Googled it. It describes what happens to individual messages, which is interesting, but I'm talking about whole streams. Those patterns are what makes all this possible under the hood; what I'm describing sits a layer or two higher. + +The patterns do build on each other in terms of what the proxy has to understand about the Kafka protocol — each one introduces a new class of work the router stage has to own. But the complexity doesn't compound when you compose them, because the routing pipeline is a DAG: each stage is narrowly scoped to its own job and doesn't need to know what the stages around it are doing. A topology-aware routing stage doesn't care whether the topic it's directing traffic to is a virtual topic or a plain one. A virtual topic stage doesn't care whether the cluster it's dispatching to is an alias or a physical cluster. You pay the cost of each stage once, and only if you need it. -After reviewing the [routing design proposal](https://github.com/kroxylicious/design/pull/70) and thinking through the problem space, I think there are four distinct patterns worth naming. They're not mutually exclusive and they compose, but each one has its own failure modes, so it helps to be able to talk about them separately. ### Cluster aliasing -A static virtual cluster endpoint maps to one physical backend, but the mapping can be changed at the proxy layer without touching clients. +Until recently a VKC provided a 1:1 mapping between the thing the client addressed and the thing the proxy dialled. The routing API breaks that constraint — one VKC can now connect to one or more physical backends, and the target isn't fixed. The VKC is stable; what's behind it doesn't have to be. + +*Boom.* A semi truck just drove through the wall of your DC and into your primary cage. What now? Kroxylicious already sits in the client cage — you tell it `VKC_ALPHA` now targets `blue-cluster` instead of `green-cluster`. Done. -Clients connect to `kafka-virtual.company.internal`. The proxy maps their connections to `cluster-blue`. When you want to migrate to `cluster-green`, you update the proxy config. Clients never reconnect, never need new bootstrap addresses, never need to know a migration happened. +This is the Kafka DR dream. So what's new? You could always update a config file and restart. The difference is the failover can now happen in flight — a router that listens to cluster health checks, or one that responds to being paged, can flip the switch without dropping client connections. That was always the easy part though. The hard part is knowing it's safe to do so. *KIP-1279: Cluster Mirroring enters stage left.* By mirroring `green` to `blue` we can trust that when we flip the switch, offset state is intact and it's safe to do so. -The catch is that this only solves connectivity. Flipping the pointer without replicating state means consumers hit offset mismatches on the new cluster. The routing layer isn't magic — it's just decoupling one specific piece of the problem. In Post 2 we'll look at pairing this with KIP-1279 to handle the state problem, which makes this pattern actually useful for DR rather than just theoretically interesting. +Aliasing is more powerful than a dead man's switch though. Consider: you're chasing an issue in the reporting pipeline that only shows up in the sixth hour of the run and nobody can figure out which record trips it up. You've been there, right? What you really want is access to the live data with a debugger. Alas, this job is stuffed full of Personally Identifiable Information — so you're out of luck. You've been there and got *that* t-shirt. What now? You could build a duplication router that shadows production traffic to a staging cluster, piping it through a redaction filter on the way so the PII becomes gibberish. The client never knows. Post 2 gets into the detail. + +Before routing, Kroxylicious was a pipe with opinions — it could inspect and rewrite the stream, but it was inherently connection-oriented. One client socket, one backend, straight through. Cluster aliasing breaks the static part of that: the backend can now change. But it's still 1:1. ### Union clusters -Multiple distinct physical clusters exposed through a single virtual cluster endpoint. The proxy synthesises a unified metadata view — to a producer or consumer it looks like one cluster. Behind the scenes, `orders-*` topics route to a high-throughput cluster, `analytics-*` topics go to a cheaper storage-optimised one. +Kroxylicious has always been like a cycle courier with opinions — it inspects, rewrites, and has strong views about what is appropriate to carry, but the destination was fixed when the client dropped off the parcel. Routing turns it into a sorting office with standing orders. You write *George Street* on the parcel and drop it at the counter. The sorting office decides whether you mean George Street, Edinburgh or George Street, Dunedin — a city Scottish settlers named after Edinburgh, gave the same street names, and promptly scrambled the layout. (Guess why I know.) The sender doesn't need to know both exist. The brokers behind a union cluster work the same way: they still have to exist somewhere, we're not making them up, but the client only ever sees the one address it was given. Because the proxy controls what gets returned in a metadata response, you decide which brokers are visible, under what names, and what maps where. And we don't need a distributed consensus layer to do it — the Kafka community just spent considerable effort going from two consensus systems down to one with KRaft. Nobody wants a third. -The obvious failure mode is namespace collisions. If `orders-v1` exists on two backends, the proxy has to pick a winner or enforce prefix rules. We haven't pushed the metadata synthesis hard enough at scale yet to characterise the memory behaviour at tens of thousands of topics across many backends — worth flagging if that's your situation. +You look after a gaggle of Kafka clusters. You know how they got there. You're not proud of all of them. You can't herd them — geese are worse than cats, they honk back — but you don't have to admit to anyone else how many geese there actually are. -### Union topics +Point your application teams at one VKC and the proxy stitches the backing clusters together into a single cluster view. The client asks for metadata and gets back a broker list. It has no idea that list was assembled from three physical clusters, one of which is on-prem and nobody's quite sure what breed it is. You implement the routing logic — the proxy gives you the framework to dispatch requests to the right place. -A single logical topic whose partitions are sharded across multiple physical clusters. A client asks for metadata for `events` with 32 partitions; partitions 0–15 live on Cluster A, 16–31 on Cluster B. A `ProduceRequest` for partition 20 gets routed to Cluster B transparently. +Two things to keep in mind. Consumer group co-ordination stays pinned to physical clusters — that's almost always fine, because a group's offset state is tied to the topic-partitions it's consuming and those live on one backend, but it's worth knowing the seam is there. And transactions don't cross cluster boundaries: a transaction that writes to topics on two different backing clusters is two independent transactions whether your code believes that or not. If cross-topic atomicity matters, those topics need to share a physical cluster. -The hard limit today: transaction coordinator boundaries break across physical clusters. If you're relying on multi-topic transactions or read-committed isolation, splitting a topic across cluster boundaries removes those guarantees. This pattern is a good fit for simple, independently partition-keyed workloads. For anything transactional, wait. +### Virtual topics -### Topology-aware routing +A single logical topic whose partitions are composed from physical topics on multiple clusters — which don't even need to share a name. A client asks for metadata for `topic-x` and gets back 32 partitions; partitions 0–15 are sourced from `topic-x` on `us-east`, 16–31 from `topic-eu` on `eu-west`. A `ProduceRequest` for partition 20 gets routed to `eu-west` transparently, with partition numbers rewritten to match the physical layout. The client sees one topic. It has no idea. -Traffic directed dynamically based on client metadata or locality — the main use case being cross-AZ bandwidth costs, which have a habit of appearing as a surprise line item in cloud bills. The proxy reads rack attributes or client IP blocks and routes fetch requests to brokers inside the same availability zone. +The motivating example here is one union clusters can't solve. Your reporting pipeline in `us-east` is hardcoded to read from `topic-x`. It has always read from `topic-x`. It will continue to read from `topic-x`. The problem is on the producer side: your business is growing in Europe, and sending every EU event across the Atlantic to land on the `us-east` cluster is expensive, slow, and fragile. The EU team provisions `topic-eu` on a cluster local to them, sized for their producer load. The router presents both physical topics as one logical `topic-x` to every client. EU producers write locally. The reporting pipeline reads the full partition space without a config change. Nobody crosses the Atlantic unnecessarily. -The failure mode here is proxy placement. A client in `us-east-1a` hitting a proxy in `us-east-1b` that forwards to a broker in `us-east-1a` doubles your cross-AZ cost instead of eliminating it. The proxy fleet has to be co-located with clients for the locality savings to materialise. +What the proxy has to do here is more demanding than in any of the previous patterns: partition numbers must be rewritten consistently across every Kafka RPC that mentions them — `Metadata`, `Produce`, `Fetch`, `OffsetCommit`, `OffsetFetch`, `ListOffsets`, group coordinator lookups. That's the essential complexity of the pattern, not incidental overhead. It's a lot of surfaces to get right. The operational contract also shifts: both physical topics have to honour the mapping, and the router config is now the source of truth for what `topic-x` means. Transactions don't cross cluster boundaries here either — same caveat as union clusters, but worth repeating because with virtual topics it's easier to forget the seam is there. --- -These four patterns give us a working vocabulary for the rest of the series. They're not all at the same maturity level — cluster aliasing and union topics are closest to production-ready, topology-aware routing is further out — and the posts that follow will be honest about where the code is versus where the design points. - -Next up: **Active-Passive DR with cluster aliasing and KIP-1279**. Fair warning: this one is going to live as a POC on a branch rather than shipped code, but it's what we're demoing at Current and the mechanics are worth understanding on their own terms. +These three patterns give us a working vocabulary for the rest of the series. The posts that follow take each one in turn — what it can do, what it can't, and why. Some of the limits are ours and will close over time. Some are the Kafka protocol's: transactions have no external coordinator, and no amount of clever routing changes that. And some would require the proxy to grow a consensus layer of its own — which is a whole different animal, and not one we're planning to domesticate any time soon. If you're thinking about Kafka topology problems, have a use case that doesn't fit cleanly into any of these patterns, or want to tell us where the vocabulary breaks down — find us on [GitHub](https://github.com/kroxylicious/kroxylicious), [Slack](https://kroxylicious.slack.com), or [Bluesky](https://bsky.app/profile/kroxylicious.io). From 94d86838c94776c275451a3c6e729108df79fbc5 Mon Sep 17 00:00:00 2001 From: Sam Barker Date: Thu, 24 Sep 2026 12:31:17 +1200 Subject: [PATCH 4/4] clarify a few points from gemini. Signed-off-by: Sam Barker --- _drafts/routing-post1.md | 6 ++++-- 1 file changed, 4 insertions(+), 2 deletions(-) diff --git a/_drafts/routing-post1.md b/_drafts/routing-post1.md index 2cf5175..da8a5bf 100644 --- a/_drafts/routing-post1.md +++ b/_drafts/routing-post1.md @@ -37,7 +37,9 @@ Kroxylicious has used the concept of a Virtual Kafka Cluster (VKC) from the star That shift — from "the VKC is where you point your bootstrap servers" to "the VKC is the contract between the client and whatever we've decided is behind it" — is what opens up the design space this series is exploring. The posts that follow get into the mechanics of each pattern and are honest about where the implementation is today versus where the design points. The talk covers the highlights; the blog is where the details live. -We get it, adding a proxy is operational overhead and we wish as much as you do swapping a CNAME from site A to site B was enough — but as far as Kafka clients are concerned, that's a divorce and why we are all here. +Kroxylicious swims happily on either side of the network boundary — helping traffic egress from the client domain (forward proxy territory) or ingress into the broker domain (reverse proxy territory). We prefer egress and ingress: they describe traffic flows from the user's perspective. As much as I pour my time and energy into Kroxylicious, it's always the supporting cast — Kafka is the leading light here, and the vocabulary should reflect that. Some of the patterns that follow land more naturally on one side than the other, and I'll call that out as we go. + +We get it, adding a proxy is operational overhead and we wish as much as you do swapping a CNAME from site A to site B was enough — but as far as Kafka clients are concerned, a CNAME flip is an unexpected divorce. That's why we're all here. ## Name them we must @@ -54,7 +56,7 @@ Until recently a VKC provided a 1:1 mapping between the thing the client address *Boom.* A semi truck just drove through the wall of your DC and into your primary cage. What now? Kroxylicious already sits in the client cage — you tell it `VKC_ALPHA` now targets `blue-cluster` instead of `green-cluster`. Done. -This is the Kafka DR dream. So what's new? You could always update a config file and restart. The difference is the failover can now happen in flight — a router that listens to cluster health checks, or one that responds to being paged, can flip the switch without dropping client connections. That was always the easy part though. The hard part is knowing it's safe to do so. *KIP-1279: Cluster Mirroring enters stage left.* By mirroring `green` to `blue` we can trust that when we flip the switch, offset state is intact and it's safe to do so. +This is the Kafka DR dream. So what's new? You could always update a config file and restart. The difference is the target swap can happen without restarting the proxy or touching a single application config — a router that listens to cluster health checks, or one that responds to being paged, can flip the switch while traffic is flowing. Today that means the client connection resets and the client driver does what it was always going to do: reconnect and retry against the new target. No app code changes, no pod restarts. Getting to a genuinely seamless in-flight handover — no reset, no retry — is somewhere we would love to go, but we have other roads to travel first. That was always the easy part though. The hard part is knowing it's safe to do so. *KIP-1279: Cluster Mirroring enters stage left.* By mirroring `green` to `blue` we can trust that when we flip the switch, offset state is intact and it's safe to do so. Aliasing is more powerful than a dead man's switch though. Consider: you're chasing an issue in the reporting pipeline that only shows up in the sixth hour of the run and nobody can figure out which record trips it up. You've been there, right? What you really want is access to the live data with a debugger. Alas, this job is stuffed full of Personally Identifiable Information — so you're out of luck. You've been there and got *that* t-shirt. What now? You could build a duplication router that shadows production traffic to a staging cluster, piping it through a redaction filter on the way so the PII becomes gibberish. The client never knows. Post 2 gets into the detail.