Senior Data Engineer
Summary
Senior data engineer designing and developing autonomous ETL/ELT pipelines using Python, PySpark, SQL Server, and Apache Airflow; managing lakehouse architecture on Kubernetes/Docker in on-premise environments; and mentoring junior engineers.
Job Description:
- Design and develop autonomous, production-grade ETL/ELT data pipelines using Python and PySpark that ingest, transform, and deliver high-quality data while maintaining integrity and performance standards.
- Implement and manage flexible lakehouse architecture across raw, curated, and consumption layers, including data partitioning, cataloging, and metadata management.
- Deploy and manage data pipelines using Kubernetes and Docker to ensure scalability, reliability, and efficient resource utilization in on-premise environments.
- Leverage strong SQL Server expertise to design optimal data models, write complex queries, and perform query optimization across the data platform.
- Establish and maintain robust CI/CD practices for data pipeline deployment, including automated testing, version control, and continuous monitoring.
- Enforce security, governance, and role-based access controls across all data layers while ensuring compliance and auditability.
- Mentor junior engineers, conduct code reviews, and establish best practices across the team.
- Collaborate with Data Scientists, Business Analysts, and stakeholders to deliver datasets aligned with operational and analytical needs.
- Provide L3 support and expert consultation for complex data challenges; evaluate and recommend new tools and practices to improve agility and performance.
Requirements:
Must-have qualifications:
- 8+ years IT experience; 5+ years hands-on data engineering or data pipeline development.
- Expert-level SQL proficiency with strong expertise in SQL Server, including query optimization, indexing, and performance tuning.
- Advanced Python programming skills for data processing, automation, and production-grade pipeline development.
- Kubernetes expertise – Design, deploy, and manage containerized data pipelines in on-premise environments.
- Strong data modeling expertise – Both relational and non-relational concepts.
- Proven experience with flexible lakehouse/data lake architecture – Multi-layer data lakes, partitioning strategies, and metadata management, Iceberg tables and optimization.
- CI/CD and DevOps practices – Setting up CI/CD pipelines, Git, automated testing, and infrastructure-as-code tools.
- ETL/ELT orchestration experience – Apache Airflow or similar tools for scheduling and monitoring batch and real-time jobs.
- Hands-on experience with at least one NoSQL database (MongoDB, Cassandra, etc.).
- Hands-on experience with Apache Spark and PySpark for distributed data processing and performance optimization.
- Data security and governance – Role-based access control, data masking, and compliance sframeworks.
- Proven ability to work autonomously on complex projects while maintaining high code quality standards.
- Excellent problem-solving, communication, and cross-functional collaboration skills.
- Bachelor's degree in Computer Science, IT, Engineering, or related field with demonstrated continuous learning ethos.
Preferred qualifications:
- Experience with on-premise data virtualization or logical data warehouse concepts.
- Understanding of data mesh or data fabric architecture patterns.
- Real-time streaming technologies (Kafka, Apache Flink).
- Metadata management and data lineage tools.
- Experience mentoring junior engineers or leading technical initiatives.
- Agile delivery methodologies and product-oriented data architecture.
Other Professional Skills and Mind-set:
- Autonomous Work Ethic – Work independently on complex problems while proactively seeking collaboration.
- Continuous Learning – Committed to staying current with data engineering trends and best practices.