Data Engineer

Summary

Data Engineer at Luma AI building high-throughput distributed data pipelines for multimodal foundation models and Physical AI, processing video/robotics datasets and synthetic data on multi-cloud GPU infrastructure (Python, PyTorch, Ray).

Intro to the Role and Team

We are expanding our Singapore Infrastructure & Data Systems Hub to power Luma's next-generation multimodal foundation models and Physical AI pipelines. As a Data Engineer, you will join an agile engineering group that builds and operates the core data engines feeding our AI models, working with our global research and infrastructure teams on large-scale data acquisition, distributed storage, and multimodal dataset processing.

What You'll Own

  • High-Throughput Data Pipelines: Build and operate automated distributed pipelines for ingesting and filtering large volumes of multimodal media data.
  • Dataset Engineering: Develop data processing workflows for video, Physical AI, and robotics datasets, as well as synthetic data generation.
  • Infrastructure Operations: Support distributed storage systems and data infrastructure across multi-cloud and GPU environments (PyTorch, Ray).
  • Performance Optimization: Diagnose and resolve network and storage bottlenecks to keep data flowing efficiently to GPU training clusters.

What You'll Bring

Basic Requirements:

  • Education: Bachelor's degree in Computer Science, Computer Engineering, Physics, Mathematics, or a related quantitative field.
  • Core Technical Stack: Proficiency in Python and at least one other language or framework for heavy data workloads (Scala, Spark, C++, or Rust).
  • Domain Expertise: Hands-on experience in at least one of the following: distributed data pipelines at scale; large-scale web data acquisition; multimodal or synthetic dataset engineering; cloud infrastructure and distributed systems.
  • Engineering Mindset: Ownership and attention to detail, with the ability to debug issues across network, storage, and compute layers.

Nice-to-Haves:

  • Experience with data loaders for GPU training frameworks (PyTorch, Ray).
  • Familiarity with vector databases or video dataset processing.
  • Prior experience in high-growth AI startups.

About Luma: Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world. We believe multimodality is critical for intelligence — the next step beyond language models comes from vision. Luma is an equal opportunity employer.

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available