Srujan Jabbireddy · Senior Data Engineer

I design data platforms, lakehouses, streaming systems, and ML data infrastructure with an emphasis on correctness, measurable efficiency, and safe evolution.

$126Kannual platform savings
74%less data scanned
5d → 1dpipeline freshness
96%fewer schema incidents

Experience themes

Data platform modernization

Lakehouse architecture and CDC

Redesigned change-data paths around durable object storage and Apache Iceberg, reducing warehouse compute while preserving safe replay, sequence-aware state, and migration controls for downstream consumers.

Event infrastructure

Streaming ingestion and recovery

Built an AWS-native path using EventBridge, Kinesis, S3, and Snowpipe so event producers, durable storage, and analytical consumers could fail and recover independently.

Analytical processing

Incremental query-path design

Replaced full-history processing with bounded correction windows and layouts aligned to investigation patterns, improving freshness without introducing streaming complexity the workflow did not need.

ML data infrastructure

Traceable multimodal datasets

Designed a multimodal pipeline with stable content identity, quality gates, reusable GPU workers, vector search, dataset manifests, provenance, and training-oriented materialization.

Core capabilities

Systems I work on

  • Batch and streaming data pipelines
  • Change data capture and replay safety
  • Lakehouse tables and query optimization
  • Data quality, lineage, and reconciliation
  • Distributed preprocessing and ML data loading
  • Dataset identity, versioning, and provenance

Selected tools

Python · SQL · Spark · Ray · Airflow · AWS DMS · Kinesis · S3 · Glue · Apache Iceberg · Snowflake · LanceDB · FAISS · PyTorch · WebDataset

Evidence, not keyword lists

See the architecture, decisions, and results.

Explore selected work