Data Infra Engineer - role plate

Data Infra Engineer

FIG. 02.01 - Role plate
Location
San Francisco
Employment type
Full-time
Salary
$200K-$300K (San Francisco)
Department
Infra
Posted
Aug 2026
NO. 01

Role overview

Build the pipelines and data infrastructure behind training our AI models.

You will work directly with researchers and engineers to make collecting, processing, curating, and feeding large datasets into training systems fast, reliable, and scalable.

NO. 02

What you'll do

  • Build large-scale ingestion, processing, and preprocessing pipelines.
  • Build systems for dataset cleaning, deduplication, transformation, versioning, and validation.
  • Feed data efficiently into distributed multi-node / multi-GPU training.
  • Build and operate storage and high-throughput data systems.
  • Develop internal tooling for dataset creation, curation, and monitoring.
  • Improve reliability, reproducibility, and throughput across the data stack.
NO. 03

Who we're looking for

  • Ideally 4-6 years of professional experience.
  • A degree in Computer Science, Computer Engineering, or a related highly technical field.
  • Strong software and distributed-systems fundamentals.
  • Experience building infrastructure directly supporting ML training.
  • Able to own critical infrastructure end-to-end.
NO. 04

What you bring

  • Experience with large-scale data pipelines, distributed systems, and ML training infrastructure.
  • Experience with object/distributed storage, preprocessing, and high-throughput data loading.
  • Familiarity with Ray, Spark, or equivalent distributed frameworks.
  • Experience with model-training infrastructure on AWS or bare-metal clusters.
  • Strong understanding of reliability, lineage, reproducibility, and data quality.
NO. 05

The background we value

Experience in AI labs, foundation-model companies, robotics, autonomous driving, computer vision, generative video/3D, or other companies training models on large proprietary datasets.

NO. 06

The attitude we value

Own the data layer: Build simple, reliable systems that remove bottlenecks and let researchers iterate faster as datasets and training workloads scale.

NO. 07

What we bring to the table

  • Competitive compensation: $200K-$300K in San Francisco, based on the candidate and level of experience.
  • Visa sponsorship and relocation support: we sponsor work visas and support your move, so joining us is straightforward wherever you are coming from.
  • A hyper-focused, high-agency team working on defining the next chapter of Physical AI models.
  • Large-scale access to GPU and compute infrastructure, giving you the resources to experiment, train, deploy, and ship without compute becoming the bottleneck.
  • Unlimited Claude and ChatGPT usage to work with the best AI-native tools available.
  • Direct access to researchers training frontier spatial AI models and ownership over infrastructure that sits on the critical path of model development.
NO. 08

Location and commitment

San Francisco · Full-time