Senior Data Engineer Data Lakehouse Infrastructure

Design, implement, and scale a modern data lakehouse for complex analytical and real-time workloads, including data modeling, ingestion, metadata management, query optimization, governance, automation, and observability.

Responsibilities

  • Architect and scale a high-performance data lakehouse on GCP
  • Design, build, and optimize distributed query engines such as Trino, Spark, or Snowflake
  • Implement metadata management using open table formats and compatible catalogs
  • Develop and orchestrate ETL and ELT pipelines with Airflow, Spark, and GCP-native tools
  • Build streaming and batch data pipelines using Dataflow and Kafka
  • Optimize query performance and data modeling for analytical workloads
  • Automate operational tasks including cluster scaling and self-service infrastructure
  • Implement observability and data discovery frameworks for governance

Requirements

  • 5+ years of experience in data or software engineering focused on distributed data systems
  • Experience building and scaling data platforms on GCP
  • Strong command of query engines such as Trino, Presto, Spark, or Snowflake
  • Experience with Apache Hudi, Iceberg, Delta Lake, or similar table formats
  • Strong Python programming and SQL or SparkSQL skills
  • Hands-on experience with Airflow and GCP-native orchestration and streaming services

Benefits

  • Equity plan participation

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available