Most Kafka posts are written from the greenfield side: new services, Kafka Streams, a clean topology. This one is from the other side. A long-lived Java system at the core of gas operations, which had to start consuming data from new event-driven platforms upstream.

I spent most of 2025 on it. Here’s what actually happened, including the parts that went wrong.

The setup

  • Consumer: a Java application that has been running operations for many years. It had to be moved to Java 21 and Jakarta EE first, before any Kafka code could land.
  • Producers: three upstream platforms publishing forecasts, schedules and actuals as Avro, owned by other teams.
  • Platform: a managed Kafka service, first reached through an HTTP REST proxy, later with the native client.
  • Constraint: this data feeds operational planning. Losing a message is bad. Processing it twice can be just as bad.

Lesson 1: Build one consumer, not three

The first feed was the obvious prototype. When the second and third feeds arrived, I didn’t copy the consumer. I generalized it into one service with pluggable processors per feed and a settings UI for topics, so onboarding a new feed became configuration plus a mapper.

That paid off twice: when we switched transport (lesson 4) and when we added our own offset handling (lesson 5). Both changes were made in one place.

Lesson 2: Your unit tests will pass. Run it for four hours.

Everything was green. Unit tests, integration tests, manual tests on the test environment. Then I wrote a long-running test scenario to watch the REST proxy’s behavior over time, and after roughly four hours offset commits started misbehaving.

No test that runs for minutes will ever find that. Since then, a soak test that runs for hours is part of my definition of done for any consumer.

Lesson 3: Replays are a conversation, not a bug ticket

After the first beta, we saw messages being processed again. Chasing it involved:

  • a quick fix for consumer group conflicts, so we could keep testing
  • a written root-cause analysis of the reprocessing
  • triage with a single topic and minimal fetch sizes to isolate the behavior
  • a small shell script that reproduced the consumer’s behavior outside the application
  • aligning retention times with the platform team, because “how long can we replay from?” turned out to be an organizational question as much as a technical one
  • a configurable deduplicator as a safety net

The lesson: when two teams share a broker, replay behavior is a contract. Write it down together.

Lesson 4: REST proxy first, native client when you understand the limits

The REST proxy made sense at the start: no native client dependencies in a legacy build, and it fit existing firewall rules. It also added a layer we didn’t control for commits and consumer lifecycle.

Moving to the native client meant:

  • Avro deserialization and schema registry handling in-process
  • serializer dependency conflicts in an old build (days, not hours)
  • agreeing on partition keys and key schemas with a producer team
  • a new round of network and security reviews

I kept both clients behind one interface, so either can run. That turned out to be useful during rollout.

Lesson 5: Stop outsourcing your position in the stream

After enough surprises, the fix wasn’t another setting. I wrote an offset manager: we persist where we are ourselves, reconcile it with the broker on startup, and seek explicitly on both the REST proxy and the native client.

It’s more code to own. But in a system where “did we process this?” has operational consequences, knowing the answer from your own database beats inferring it from broker state.

The integration went live in October 2025.

By the numbers

DurationFeb 2025 to Dec 2025, 11 months
Effort~185 working days touched the Kafka integration
Upstream feeds3, one generalized consumer
TransportREST proxy first, native client added in July 2025, both behind one interface
Time until the long-running bug appeared~4 hours
First reprocessing seen to own offset manager~3.5 months (late May to early Sep 2025)
Releases shipped with it2 major releases
Go-liveOctober 2025

What I’d tell someone starting this today

  1. Generalize the consumer after the second feed, not the fifth.
  2. Soak test for hours before you trust any commit behavior.
  3. Agree on retention, replay and partition keys with the producer team in writing.
  4. Start with the simplest transport, but keep it behind an interface.
  5. If duplicates matter, own your offsets.

None of this is exotic. It’s just what you learn when Kafka meets a system that was never designed for it, and when “processed twice” is a real problem.