Small File Problem in Data Lakehouse - Apache Iceberg Compaction

Small files are a common problem in Lakehouse tables such as Apache Iceberg, particularly with streaming workloads that continuously commit data in small batches.

In this episode, Dipankar explores how Apache Iceberg compaction helps address the small file problem.

He goes over:
✅ How small files are created
✅ Why they can affect query performance
✅ How Iceberg's Rewrite Data Files operation works
✅ How the default bin-packing strategy rewrites small files
✅ Key parameters that control grouping, concurrency, output file size, and commit behavior
✅ How Cloudera Lakehouse Optimizer can automate and operationalize compaction across multiple Iceberg tables

Episode 1: https://youtu.be/oW6dNIxWKwk
Episode 2: https://youtu.be/6dM5dVVGGos
Episode 3: https://youtu.be/PWD84gK5_O0
Episode 4: https://youtu.be/Y61BpYDumsk

🔔 Subscribe to stay ahead in enterprise AI, open architectures, and data strategy: https://www.youtube.com/channel/UCXY5wm6HlBL_Y_8SDxJNR0g

Join the Cloudera Community to learn more! 👉https://community.cloudera.com
Explore the Full Series: 👉 Full Playlist: https://www.youtube.com/playlist

Chapters:

00:00 Solving Iceberg Small Files

01:05 What Causes Iceberg Small Files

01:58 How Iceberg Compaction Works

08:56 Target File Size Optimization

12:44 Balancing Performance & Costs

14:15 [Demo] Iceberg Compaction Policy

#lakehouse #ApacheIceberg #DataEngineering #DataLakehouse #Cloudera #BigData