1 min readfrom Towards Data Science

Avoiding Entity Key Drift in a Data Lake: Step 1, Normalization

Avoiding Entity Key Drift in a Data Lake: Step 1, Normalization

This is the opening piece of a four-part deep dive series, on building a high-frequency streaming pipeline against a live public API. The data source is openSenseMap, a citizen-science IoT network used for climate research, mostly in Germany. A live public API is what makes it useful: it produces data-quality problems and edge cases that clean sample datasets never show. This article focuses on step-1: Normalization, later pieces cover matching algorithms, adaptive polling and noise filtering, and a vendor-agnostic Apache Iceberg pipeline with Terraform that runs locally in Docker and moves to AWS or GCP with minimal change.

The post Avoiding Entity Key Drift in a Data Lake: Step 1, Normalization appeared first on Towards Data Science.

Want to read more?

Check out the full article on the original site

View original article

Tagged with

#Data Lake
#Entity Key Drift
#Normalization
#Streaming Pipeline
#Live Public API
#openSenseMap
#IoT
#Climate Research
#Data Quality
#Edge Cases
#Apache Iceberg
#Terraform
#Docker
#AWS
#GCP
#Adaptive Polling
#Noise Filtering
#Matching Algorithms
#Vendor-Agnostic
#Citizen-Science