Enhance your career, get your certificate as a Data Streaming Engineer | Get your Certificate
Stream MQTT sensor data into Kafka: a walkthrough of Mosquitto, the Confluent MQTT Connector, HiveMQ, Zilla, AWS IoT Core & Kepware.
A deep dive into deduplicating streams with Flink SQL, weighing state growth, changelog modes, late events, and latency, and why a custom PTF can beat clever SQL.
A framework for reviewing Flink SQL solutions: state growth, append-only vs. updating output, late events, latency, and determinism before going to production.
How the new FROM_CHANGELOG and TO_CHANGELOG built-in functions let Flink SQL read custom CDC formats and turn updating streams back into append-only ones.
Confluent Cloud for Apache Flink watermarks deep-dive. Differences and extensions to Apache Flink: default watermarks, idleness, alignment, late events handling, watermark observability.
Deep-dive reference on Apache Flink watermarks: generation, propagation, idle partition detection, alignment, and why lateness is non-deterministic.
How Hardwood detects fixed-length lists in Parquet's encoded definition and repetition levels, bypassing Dremel reconstruction for a 1.1×–3.9× parse speed-up.
Learn how to use Queues for Apache Kafka (KIP-932) with the Quarkus framework, including share group configuration, explicit message acknowledgment, and cooperative consumption from a single partition.
How we traced a native-memory leak in Flink's RocksDB block cache and a jemalloc fragmentation bug, cutting OOMKilled Task Managers by 91.2% — plus what running the investigation with Claude taught us.