PySpark Data Engineer
Summary
On-site PySpark Data Engineer in Jaipur building and maintaining big data applications, contributing to architecture and performance tuning (executor sizing, partitions, code optimization) while interfacing with business users on requirements and troubleshooting.
Roles and
Responsibilities:
- Responsible
for developing and maintaining applications with PySpark
- Contribute
to the overall design and architecture of the application developed and
deployed.
- Performance
Tuning wrt to executor sizing and other environmental parameters, code
optimization, partitions tuning, etc
- Interact
with business users to understand requirements and troubleshoot issues.
- Implement
Projects based on functional specifications.
Must-Have
Skills:
- Relevant
Experience: 3-6 Years
- SQL -
Mandatory
- Python -
Mandatory
- SparkSQL -
Mandatory
- PySpark -
Mandatory
- Hive -
Mandatory
- HDFS and
Spark - Mandatory
- Scala -
Advantage
- Apache
Airflow - Advantage
Requirements
3-6 Years of Experience
Must Have: PySpark/Spark, Python, SQL, Knowledge on Hadoop ecosystem
Good to have: Airflow, Scala