Senior Data Engineer

Summary

Senior Data Engineer at Luma AI designing and scaling distributed data infrastructure for multimodal foundation models and Physical AI pipelines, ingesting 100M+ media items/day across multi-cloud GPU environments using Python, PyTorch, and Ray.

Intro to the Role and Team

We are expanding our Singapore Infrastructure & Data Systems Hub to power Luma's next-generation multimodal foundation models and Physical AI pipelines. As a Senior Data Engineer, you will lead the design and scaling of the core data engines feeding our AI models, collaborating directly with our global research and infrastructure teams and setting technical direction for high-throughput data systems, distributed storage, and multimodal dataset processing.

What You'll Own

  • High-Throughput Data Infrastructure: Architect, scale, and operate distributed pipelines capable of ingesting and filtering 100M+ multimodal media items per day.
  • Physical AI & Synthetic Data Engines: Design end-to-end hardware and software data pipelines for egocentric video, Physical AI, and robotics datasets, and build domain-specific synthetic data platforms.
  • Reliability & Storage: Own SRE operations, unified access control, and distributed storage systems across multi-cloud and GPU environments (PyTorch, Ray).
  • Systems Performance: Identify and resolve bottlenecks across network, storage, and compute layers to maximize GPU training cluster utilization; drive engineering best practices and mentor junior engineers.

What You'll Bring

Basic Requirements:

  • Education: Bachelor's or Master's degree in Computer Science, Computer Engineering, Physics, Mathematics, or a related quantitative field.
  • Experience: 5+ years building and operating large-scale data systems in production.
  • Core Technical Stack: Strong proficiency in Python plus systems languages or frameworks for high-concurrency workloads (Scala, Spark, C++, or Rust).
  • Domain Expertise: Deep hands-on experience in at least two of the following: distributed data pipelines and vector storage at scale (100M+ items/day); large-scale web data acquisition and network infrastructure; multimodal, Physical AI, or synthetic/robotics dataset engineering; SRE practices and distributed systems management.
  • Leadership: Track record of driving technical direction autonomously and raising the bar for engineering quality.

Nice-to-Haves:

  • Experience building or maintaining data loaders for large-scale GPU training frameworks (PyTorch, Ray).
  • Familiarity with multimodal vector databases or egocentric video dataset processing.
  • Prior experience scaling globally distributed teams in high-growth AI startups.

About Luma: Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world. We believe multimodality is critical for intelligence — the next step beyond language models comes from vision. Luma is an equal opportunity employer.

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available