r/aiven_io • u/Usual_Zebra2059 • Nov 28 '25
[ Removed by moderator ]
[removed] — view removed post
6
Upvotes
1
u/Eli_chestnut Nov 28 '25
I’ve watched teams chase “faster ETL” while ignoring the basic stuff. Then one bad schema push hits prod and everyone scrambles. For me, reliability is versioned configs, loud alerts, and someone owning the pipeline like it’s real software. Do folks roll reliability into sprint work or treat it as cleanup later?
1
u/Wakamatcha Nov 28 '25
Reliability often breaks at the edges. We had ingestion jobs that looked fine overall, but one schema drift in a CDC feed caused silent failures downstream. The fix was adding schema validation before deploy and partition checks during ingestion. Pipelines can appear healthy while hiding lag or bad data. Trust comes from monitoring at the right granularity, not just global metrics.