Synthetic Data Engineer (AI Data/Training)
We are seeking a talented and innovative Synthetic Data Engineer. In this role, you will design and implement domain-specific synthetic data generation pipelines, ensuring high-quality data management for training loops. Your expertise will drive the success of data processing and model training within the organization.
Responsibilities:
- Design domain-specific synthetic data generation (SDG) pipelines via self-instruct and constitutional prompting.
- Implement automated quality scoring and de-duplication systems.
- Manage data pipelines that feed directly into SFT and DPO training loops.
Qualifications:
- Proven experience building large-scale data pipelines (Airflow, Spark, Ray).
- Deep knowledge of prompt engineering for data generation.
- Familiarity with dataset distillation and bias mitigation.
As published by greenhouse
First Name, Last Name, Email, Phone, Resume/CV, Cover Letter
- Preferred First Name optional
- Website optional
- LinkedIn Profile optional
- Notice Period
- Current Annual Salary (with Currency)
- Expected Annual Salary (with Currency)
- Working Location choose any
- Do you have any Web3 experience? choose one
- Web3 Vertical Experience choose any
- Any personal experience in Web3 (e.g. side project, personal investment) if no professional experience. written answer