Data System Design Interview Glossary
The core concepts and tradeoffs used in data-platform architecture interviews.
Systems in Practice · Library
Pipelines, lakehouses, streaming, and reliable data operations.
14 items · newest first
The core concepts and tradeoffs used in data-platform architecture interviews.
A study guide to ingestion, enrichment, quality, governance, and operations at scale.
A structured reference to the stages and tradeoffs in the Nemotron pretraining data pipeline.
Examining the data preparation behind the Nemotron 3 Super pretraining corpus.
The connected systems that acquire, validate, enrich, serve, and monitor production ML data.
How batch and streaming paths combine into consistent views for machine learning.
Designing the ingestion, serving, and monitoring systems behind production feature stores.
Turning immutable dataset manifests into WebDataset shards and efficient training loaders.
Precomputation, provenance, and observability turn a search demo into dependable data infrastructure.
The next design questions for a multimodal lakehouse: quality, deduplication, and evaluation.
Moving from single-machine ETL to distributed pipelines without losing correctness or control.
Content-addressed storage, dataset versioning, and deduplication in a multimodal lakehouse.
A 12-stage pipeline for turning text, images, video, and audio into traceable training data.
Orchestrating a batch analytics pipeline with Airflow, AWS EMR, and Redshift.
No writing matches these filters. Try another term or clear the filters.