Senior Data Engineer Data Lakehouse Infrastructure
Design, implement, and scale a modern data lakehouse for complex analytical and real-time workloads, including data modeling, ingestion, metadata management, query optimization, governance, automation, and observability.
Responsibilities
- Architect and scale a high-performance data lakehouse on GCP
- Design, build, and optimize distributed query engines such as Trino, Spark, or Snowflake
- Implement metadata management using open table formats and compatible catalogs
- Develop and orchestrate ETL and ELT pipelines with Airflow, Spark, and GCP-native tools
- Build streaming and batch data pipelines using Dataflow and Kafka
- Optimize query performance and data modeling for analytical workloads
- Automate operational tasks including cluster scaling and self-service infrastructure
- Implement observability and data discovery frameworks for governance
Requirements
- 5+ years of experience in data or software engineering focused on distributed data systems
- Experience building and scaling data platforms on GCP
- Strong command of query engines such as Trino, Presto, Spark, or Snowflake
- Experience with Apache Hudi, Iceberg, Delta Lake, or similar table formats
- Strong Python programming and SQL or SparkSQL skills
- Hands-on experience with Airflow and GCP-native orchestration and streaming services
Benefits
- Equity plan participation