Data Scientist
Design, build, and optimize data pipelines and ETL workflows in Snowflake using Snowpark, Streams/Tasks, and Snowpipe. Develop scalable data models for user 360 views, churn prediction, and recommendation engine inputs; integrate diverse data sources; implement CI/CD and data quality checks; mentor junior engineers; support machine learning feature productionization; and establish governance, lineage, metadata standards, and streaming architecture practices.
Responsibilities
- Design, build, and optimize data pipelines and ETL workflows in Snowflake using Snowpark, Streams/Tasks, and Snowpipe
- Develop scalable data models supporting user 360 views, churn prediction, and recommendation engine inputs
- Lead integration across MySQL, BigQuery, Redis, Kafka, GCP Storage, and API Gateway
- Implement CI/CD for data pipelines using Git, dbt, and automated testing
- Define data quality checks and auditing pipelines for ingestion and transformation layers
- Mentor and guide junior data engineers on data modeling, performance tuning, and Snowflake best practices
- Partner with Data Science, ML, and Backend teams to productionize machine learning features in Snowflake
- Ensure compliance, privacy, and governance of user data with Legal, Security, and Infrastructure teams
- Translate business requirements into technical specifications with stakeholders
- Tune algorithm performance and establish partitioning, clustering, and materialized views
- Build dashboards and monitors for pipeline health, job success, and data latency
- Establish naming conventions, data lineage, and metadata standards
- Lead code reviews, enforce documentation standards, and manage schema versioning
- Contribute to the company’s data mesh and streaming architecture vision
Requirements
- 5+ years of experience in a Data Scientist role, including 3+ years with Spark
- Strong SQL and Python skills with ETL/ELT experience at scale
- Deep understanding of algorithm performance tuning, query optimization, and warehouse orchestration
- Experience with Airflow, Prefect, dbt, or similar orchestration tools
- Understanding of Kimball, Data Vault, or hybrid data modeling
- Proficiency with Kafka, GCP, or AWS for real-time or batch ingestion
- Familiarity with API-based data integration and microservice architectures
- Experience leading machine learning teams or deploying ML feature pipelines
- Background in ad-tech, gaming, or e-commerce recommendation systems
- Familiarity with data contracts and feature stores such as Feast or Tecton
- Experience managing small data engineering teams and setting technical direction
- Strong ownership, autonomy, cross-functional communication, problem-solving, and mentoring skills
Benefits
- Medical, dental, and vision insurance
- PTO
- Personalized career roadmap
- Professional development through training and educational opportunities