Professional snapshot
Srujan Jabbireddy · Senior Data Engineer
I design data platforms, lakehouses, streaming systems, and ML data infrastructure with an emphasis on correctness, measurable efficiency, and safe evolution.
Experience themes
Lakehouse architecture and CDC
Redesigned change-data paths around durable object storage and Apache Iceberg, reducing warehouse compute while preserving safe replay, sequence-aware state, and migration controls for downstream consumers.
Streaming ingestion and recovery
Built an AWS-native path using EventBridge, Kinesis, S3, and Snowpipe so event producers, durable storage, and analytical consumers could fail and recover independently.
Incremental query-path design
Replaced full-history processing with bounded correction windows and layouts aligned to investigation patterns, improving freshness without introducing streaming complexity the workflow did not need.
Traceable multimodal datasets
Designed a multimodal pipeline with stable content identity, quality gates, reusable GPU workers, vector search, dataset manifests, provenance, and training-oriented materialization.
Core capabilities
Systems I work on
- Batch and streaming data pipelines
- Change data capture and replay safety
- Lakehouse tables and query optimization
- Data quality, lineage, and reconciliation
- Distributed preprocessing and ML data loading
- Dataset identity, versioning, and provenance
Selected tools
Python · SQL · Spark · Ray · Airflow · AWS DMS · Kinesis · S3 · Glue · Apache Iceberg · Snowflake · LanceDB · FAISS · PyTorch · WebDataset
Evidence, not keyword lists