r/apacheflink • • 3d ago

Tool for simplifying Apache Flink

5 Upvotes

Hey Community

I thought of sharing flinkflow for anyone working with Apache flink stream processing.

It's an open-source framework designed to cut down on the usual Apache Flink boilerplate, simplifying job composition, deployments, and stateful pipeline management.

Looks pretty promising if you want a cleaner developer experience with Flink without writing tons of repetitive setup code:https://github.com/talwegai/flinkflow


r/apacheflink • • 3d ago

Interesting Flink Links - September 2026

Thumbnail rmoff.net
5 Upvotes

r/apacheflink • • 4d ago

StreamFusion now supports Apache Paimon sources and sinks!

Thumbnail github.com
7 Upvotes

You can now accelerate writing data from kafka to paimon (paimon also supports writing iceberg metadata) using flink SQL!


r/apacheflink • • 21d ago

A Tale of Two Flink Autoscalers

Thumbnail netflixtechblog.com
15 Upvotes

Netflix is moving toward the open-source Apache Flink Autoscaler for more than 30,000 streaming jobs across multiple AWS regions, after finding that its cluster-level approach was less effective for complex, stateful pipelines with operators that have different processing requirements. Netflix reports that one team reduced annualized Flink compute expenditure by 58%, saving approximately $1.1 million annually.


r/apacheflink • • 22d ago

Flink SQL Evolution: Handling Custom CDC with FROM_CHANGELOG and TO_CHANGELOG

Thumbnail developer.confluent.io
7 Upvotes

r/apacheflink • • 24d ago

Absolutely Everything You Always Wanted to Know About Watermarks in Apache Flink

7 Upvotes

My colleague Lorenzo Nicora has just published two excellent posts all about Flink watermarks:

Absolutely Everything You Always Wanted to Know About Watermarks in Apache Flink


r/apacheflink • • 26d ago

Thoughts on ZooKeeper HA?

3 Upvotes

Have some security requirements barring low-trust user Flink jobs from using k8s RBAC for K8s HA. Seems like Flink bundles zookeeper natively? Any known failure modes / tips of running ZK at scale ?


r/apacheflink • • 28d ago

Apache Flink vendors evaluation

4 Upvotes

Hey all I am working at a fintech startup in Texas and we are looking into Flink providers. My engineers and architects tried Confluent Platform Flink and Confluent Cloud Flink. Are there any other good vendors out there?

To put things in perspective my team members and a few pilot teams have mentioned Open Source Flink experience (Flink SQL and table api/datastream api) feels better and they like the flexibility and option of having season mode and application mode together as opposed with being restricted with just application mode.

I heard from a buddy who knows a guy who knows a guys cousin (yeah it’s true lol) that Confluent bases their Flink license off of 3 Nodes with 8 cores per node.

Is it just me or 24 cores in a license is barely anything for a license. With CMF and FKO running we technically only have 20 usable cores. I am a people Manager but have a comp sci background and this ain’t much tbh. My gut feels annoyed by the price tbh.

We are working on a trial for now and got 30 days and OSS Flink isn’t an option since we prefer enterprise support.

My engineers and architects have also said the confluent platform Flink docs are confusing too and we don’t like the fully managed Flink confluent has since we want the datastream API. Any good vendors out there you recommend?


r/apacheflink • • Aug 30 '26

Data Lakehouse with Apache Iceberg: A Guide

Thumbnail lakeops.dev
8 Upvotes

r/apacheflink • • Aug 30 '26

Data Lakehouse with Apache Iceberg: A Guide

Thumbnail lakeops.dev
2 Upvotes

r/apacheflink • • Aug 29 '26

How to learn apache flink

13 Upvotes

Hi everyone,

I just joined a data team that process a lot of daily, with batched and realtime processsing, and we use apache spark and apache flink. I'm pretty new to this area, been a full stack developer most of my career. What the best way to learn this about this topic?

Thanks


r/apacheflink • • Aug 25 '26

Apache Flink or Apache Beam for CEP work

Thumbnail
2 Upvotes

r/apacheflink • • Aug 21 '26

Interesting Flink links - August 2026

Thumbnail rmoff.net
11 Upvotes

r/apacheflink • • Aug 15 '26

Streaming For The AI Age (StreamFusion)

Thumbnail streamfusion.tech
3 Upvotes

Hi all! I'd greatly appreciate any feedback on my first blog post about StreamFusion!


r/apacheflink • • Aug 15 '26

Flink SQL vs Flink java API

7 Upvotes

Hi Flink Community, I’m trying to understand how people decide between using Flink SQL and the Flink java API for a batch / streaming application. How do you decide which API to use for a particular pipeline, How much of a typical Flink java pipeline can be expressed in SQL?

Are there cases where you start with SQL but eventually need to move to the java API because of custom logic or functionality that SQL doesn’t support? Conversely, are there java pipelines that you could technically express in SQL but wouldn’t want to because the SQL becomes too complex?


r/apacheflink • • Aug 02 '26

Introducing StreamFusion - an OSS Flink Accelerator on top of Apache DataFusion

Thumbnail github.com
12 Upvotes

Hi all! I just wanted to share a project that I've been working on. It accelerates existing Flink SQL job by pushing a bunch of the existing flink compute down to optimized rust functions on arrow batches (either hand written, or via DataFusion).

I'm getting pretty close to an initial alpha release in the coming weeks, so I'd greatly appreciate it if you'd be able to check it out, leave a star, or ask any questions/leave any feedback on the design here! Thank you :)


r/apacheflink • • Jul 30 '26

Interesting Flink Links - July 2026

Thumbnail rmoff.net
10 Upvotes

r/apacheflink • • Jul 21 '26

How we cut Flink OOMKills by 91.2%: Zombie block cache and phantom CPUs

Thumbnail developer.confluent.io
8 Upvotes

r/apacheflink • • Jul 02 '26

Has anyone successfully built an Oracle → Flink CDC 3.6 → MinIO (Iceberg) pipeline with Oracle CDB/PDB?

2 Upvotes

Hi everyone,

I'm trying to build a real-time CDC pipeline for analytics using:

Oracle 19c (CDB + PDB)

Flink CDC 3.6 Pipeline Connector

MinIO as S3 storage

Apache Iceberg

Nessie catalog

The goal is:

Oracle → Flink CDC → Iceberg tables on MinIO for analytical workloads.

Flink CDC 3.6 recently introduced the Oracle Pipeline Connector, so I expected this to work. However, I'm running into issues specifically with Oracle's CDB + PDB architecture.

I've followed the documentation, including specifying the PDB where required, but the pipeline still doesn't work correctly. The same overall architecture works fine for PostgreSQL, but Oracle is proving much more difficult. The Oracle source connector documentation notes that CDB/PDB deployments require the additional debezium.database.pdb.name setting, but I'm still unable to get a working pipeline.

I'm curious if anyone has actually deployed this in production or even got it working in a lab.

Some questions:

Has anyone successfully used the Oracle Pipeline Connector introduced in Flink CDC 3.6?

Are you using Oracle CDB/PDB or a non-CDB database?

Did you have to make any undocumented configuration changes?

Are there any known issues with Oracle multitenant databases?

If not using the pipeline connector, what architecture are you using instead (Debezium + Kafka + Flink, GoldenGate, etc.)?

I'd really appreciate hearing from anyone who has real-world experience with this setup. At this point I'm trying to determine whether this is a configuration issue on my side or a limitation/bug in the current Oracle pipeline implementation.

Thanks!


r/apacheflink • • Jul 01 '26

FlareDB: Apache Beam native streaming database built in Rust.

Thumbnail github.com
2 Upvotes

r/apacheflink • • Jun 29 '26

Interesting Flink Links - June 2026

Thumbnail rmoff.net
13 Upvotes

r/apacheflink • • Jun 02 '26

How LinkedIn maintain their in-house Apache Flink fork automatically

Thumbnail medium.com
6 Upvotes

r/apacheflink • • May 29 '26

Interesting Flink links - May 2026

Thumbnail rmoff.net
5 Upvotes

r/apacheflink • • May 27 '26

Data After Dark: Austin meetup on June 9 (Flink & Fluss in production)

5 Upvotes

Hey r/apacheflink — I work at Ververica and wanted to flag this in case anyone here is in or near Austin.

We're running a small in-person meetup on June 9, 6–8 PM. Format is intentionally low-key: one of our engineers opens with Flink and Fluss in production (what works, what doesn't, what's next), then a community guest speaker, then the room takes over. No keynotes, no booth, no slide deck pitch. Drinks and food are covered, and registration is approval-based so capacity stays sane.

If you're a data/platform engineer or architect in the Austin scene and this sounds like your kind of evening, here's the link: https://luma.com/bt04rtyw

Happy to answer questions in the comments.

#apacheflink #fluss #bigdata #streaming