Study Guide: Data Operations Architecture at Scale
A study guide to ingestion, enrichment, quality, governance, and operations at scale.
Systems in Practice · Library
Training data, feature stores, retrieval, and model evaluation.
17 items · newest first
A study guide to ingestion, enrichment, quality, governance, and operations at scale.
A reference to requirements, architecture, and operational concepts for agentic systems.
A structured reference to the stages and tradeoffs in the Nemotron pretraining data pipeline.
Examining the data preparation behind the Nemotron 3 Super pretraining corpus.
The connected systems that acquire, validate, enrich, serve, and monitor production ML data.
Building a product-quality feature store, from raw reviews to serving and drift monitoring.
How batch and streaming paths combine into consistent views for machine learning.
Designing the ingestion, serving, and monitoring systems behind production feature stores.
Turning immutable dataset manifests into WebDataset shards and efficient training loaders.
Using evaluation feedback to decide whether a new dataset version deserves promotion.
The next design questions for a multimodal lakehouse: quality, deduplication, and evaluation.
A 12-stage pipeline for turning text, images, video, and audio into traceable training data.
Finding unusual electricity consumption patterns with time-series analysis and machine learning.
Building and serving a computer-vision model for recognizing amenities in property photos.
Comparing forecasting approaches for retail product demand across stores and states.
Examining seasonality and time-series patterns in retail sales before building forecasts.
Exploring the M5 retail dataset through product demand, prices, stores, and sales patterns.
No writing matches these filters. Try another term or clear the filters.