
ML Platform Engineer / Infrastructure Engineer
- Location
- San Francisco
- Employment type
- Full-time
- Salary
- $200K-$300K (San Francisco)
- Department
- Infra
- Posted
- Aug 2026
Role overview
Build the infrastructure platform used by research, data, training, inference, and product teams across Speridlabs.
You will help define how our cloud, compute, security, clusters, and developer infrastructure are organized as the company scales.
What you'll do
- Build and evolve our AWS cloud and compute infrastructure.
- Design cloud accounts, environments, permissions, IAM, and governance.
- Build secure infrastructure for training, data, inference, and product workloads.
- Operate clusters and distributed compute environments.
- Build orchestration, scheduling, observability, and infrastructure automation.
- Create infrastructure that stays simple and scalable as the company and infrastructure team grow.
Who we're looking for
- Deep experience building and operating production infrastructure.
- Strong knowledge of AWS; Azure experience is valuable.
- Experience with compute-intensive or distributed workloads.
- Strong understanding of security, networking, reliability, and cloud governance.
- Able to design infrastructure for an organization, not just operate individual systems.
What you bring
- Strong AWS, cloud architecture, IAM, networking, and governance experience.
- Experience with infrastructure-as-code, automation, and observability.
- Experience with Docker, Kubernetes, SLURM, or similar orchestration and scheduling systems.
- Experience building infrastructure for ML training, data, and inference workloads.
- Nice to have: experience with bare-metal, GPU or HPC infrastructure, as well as NVIDIA Run:ai, KAI Scheduler, or similar GPU scheduling platforms.
The background we value
Experience at cloud or compute infrastructure providers, HPC environments, AI research organizations, or startups training or deploying AI models at scale.
The attitude we value
Build for scale: Create infrastructure and operating principles that work today and can be inherited by a growing infrastructure team tomorrow - without complexity exploding as the organization grows.
What we bring to the table
- Competitive compensation: $200K-$300K in San Francisco, based on the candidate and level of experience.
- Visa sponsorship and relocation support: we sponsor work visas and support your move, so joining us is straightforward wherever you are coming from.
- A hyper-focused, high-agency team working on defining the next chapter of Physical AI models.
- Large-scale access to GPU and compute infrastructure, giving you the resources to experiment, train, deploy, and ship without compute becoming the bottleneck.
- Unlimited Claude and ChatGPT usage to work with the best AI-native tools available.
- Ownership over the infrastructure foundations of the company and the opportunity to help shape the future infrastructure team as Speridlabs scales.
Location and commitment
San Francisco · Full-time