Data Engineering & Real-Time Event StreamsPlaybook4 min readUpdated September 2026

Where Real-Time Pipeline Costs Actually Come From

Most real-time pipeline cost comes from broker storage, cross-availability-zone network traffic, and replication settings, not from stream processor compute. Teams rarely revisit a replication factor or retention window after setup, so check those before reaching for a bigger fix like re-architecting the whole pipeline.

Here's where the money actually goes, in roughly descending order of how often it's the biggest lever, and what to check before you reach for a bigger fix like re-architecting the whole pipeline.

Retention windows that are longer than anyone actually needs

Every topic has a retention window, and the default your team picked during setup (often seven or thirty days, whatever the platform suggested) tends to stick around long after anyone checks whether it's actually needed. Say your fraud-detection topic only needs 48 hours of history for reprocessing, but it's still configured to retain 30 days; that's storage cost for data nobody reads after the first two days.

Audit retention topic by topic against what actually consumes old data: reprocessing jobs, backfills, or compliance requirements. Most topics need far less than their current setting, and cutting retention is usually the single fastest win because it doesn't touch application code at all.

A replication factor set once and never revisited

Replication factor trades storage cost directly for durability, and most teams set it once during initial setup and never ask whether every topic needs the same level of protection. A topic holding reconstructable, low-value telemetry doesn't need the same replication factor as a topic holding financial transactions.

Tier your topics by how bad it would actually be to lose them, and apply replication factor accordingly. This is a smaller lever than retention for most pipelines, but it compounds: storage cost scales directly with replication, so a topic replicated more heavily than it needs pays that multiplier on its full retention window too.

Over-provisioned consumers running for a traffic pattern that's gone

Consumer instance counts tend to get set for a peak traffic event (a launch, a seasonal spike) and never scaled back down afterward. Check your consumer group sizing against actual current partition count and throughput, not against the traffic pattern that justified the original sizing months ago.

If your traffic has meaningful daily or weekly seasonality, autoscaling consumer groups against a lag-based metric captures real savings without the manual babysitting of adjusting instance counts by hand every time traffic shifts.

Cross-zone and cross-region network traffic hiding in the bill

Producers, brokers, and consumers that aren't co-located in the same availability zone generate cross-zone network charges on every message, and this cost is easy to miss because it shows up as a general networking line item rather than something clearly tied to the pipeline. If your consumers are spread across zones for availability but your producers are concentrated in one, you may be paying a network tax on nearly every message.

This tradeoff is real: co-locating everything in one zone for cost savings reduces your resilience to a zone outage. Decide deliberately which topics need cross-zone resilience and which don't, rather than defaulting every topic to the same setup.

Where a sales tool like Pipedrive or Close doesn't help

Sales and pipeline CRM tools like Pipedrive and Close manage deal flow and customer outreach; they have no visibility into broker storage, replication settings, or cross-zone network charges, and they aren't built to. If cloud infrastructure cost is the actual problem, the fix lives in your cloud provider's billing breakdown and your streaming platform's own configuration, not in a sales tool.

The practical first step is pulling a cost breakdown by service (storage, network, compute) for the pipeline specifically, since that view usually reveals which of the four levers above is actually worth fixing first.

Compaction and tiered storage instead of just shrinking retention

If a topic genuinely needs long retention for compliance or replay reasons, shrinking the window isn't an option, but tiered storage often still is. Many streaming platforms can offload older segments to cheaper object storage automatically while keeping recent data on fast local disks, so you keep the retention window without paying premium storage prices for the tail end of it.

Log compaction is a different lever: for topics that only care about the latest value per key (a changelog of current account balances, say, rather than every historical update), compaction can shrink storage dramatically compared with keeping every message for the full retention period. Check whether your highest-storage topics actually need full history or just the latest state per key before assuming retention is the only variable available.

A quarterly review beats a one-time cleanup

A single cost cleanup buys you savings for a quarter or two, then traffic patterns shift, a new topic gets added with default settings, and the same waste creeps back in. Put a recurring review on the calendar, tied to the same cadence as your broader cloud cost reviews, rather than treating this as a one-time project.

The review doesn't need to be long: pull the cost breakdown by service, check it against the last quarter's numbers, and flag anything that grew faster than traffic did. Most of the time that fifteen-minute check catches a problem before it becomes a line item big enough to need a dedicated audit.

In each cost review, check these items:

  • Retention windows on every topic, compared with what reprocessing jobs, backfills, or compliance rules actually read.
  • Replication factor, tiered by how costly it would be to lose each topic's data.
  • Consumer instance counts, compared with current partition count and throughput rather than a past traffic peak.
  • Cross-zone network charges between producers, brokers, and consumers that aren't co-located.
  • Whether tiered storage can move older segments to cheaper object storage while keeping the retention you need.
Executive Capability Standard

What Good Looks Like

Pipeline cost is under control when retention, replication factor, and consumer sizing are each set per topic based on actual need, not on defaults inherited from setup.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Pull a cost breakdown by service (storage, network, compute) for the pipeline so you know which lever actually matters before changing anything.
2. Do Manually:Audit retention and replication factor topic by topic against what genuinely still reads old data.
3. Delegate:Assign a platform engineer to own quarterly cost review for the pipeline, tied to the same cadence as your broader cloud cost reviews.
4. Automate:Set up lag-based autoscaling for consumer groups so instance count tracks real traffic instead of a manually set peak.
5. Buy:Bring in a cloud cost or FinOps specialist for a one-time deep audit if the pipeline's bill has grown faster than its traffic has.

How to Get Started

Frequently Asked Questions

What's usually the fastest cost win on a real-time pipeline?

Cutting retention windows on topics where nothing actually reads data past the first day or two. It requires no code changes, takes effect immediately, and on pipelines that have been running for a while with default settings, it's frequently the single largest line item nobody had actually checked.

Should every topic have the same replication factor?

No. Match replication factor to how costly it would be to lose that topic's data. Reconstructable telemetry or logs can usually run leaner than financial or transactional topics, and since replication cost multiplies across the full retention window, this tiering compounds into a meaningful saving on high-volume, low-value topics.

How do we know if our consumers are over-provisioned?

Compare current consumer instance count against actual partition count and average lag over the past few weeks, not against whatever traffic spike justified the original sizing. If lag stays near zero with room to spare during normal traffic, you're very likely paying for headroom you no longer need.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides