All open positions

Transparent Search Group

ML Infrastructure Engineer

Posted today
Redwood City, California, United States$220K–$350KRobotics

About the role

Company: Confidential - Series A robotics company
Location: Redwood City, CA (in office 5 days per week)
Compensation: $220,000 - $350,000 + competitive equity
Employment Type: Full-time
Visa Sponsorship: Visa transfers (OPT, H-1B transfer)

About the Company

A well-funded Series A robotics company building general-purpose robots powered by its own embodied AI foundation model, already deployed at customer sites in hospitality and food service.

The Role

Our client is hiring an ML Infrastructure Engineer to own training infrastructure end to end and turn a multi-cloud GPU fleet into a world-class training engine for massive multimodal models. You will be the connective tissue between researchers and compute, and your work directly speeds the path from model to deployed robot.

What You Will Do

  • Architect and scale distributed training across large GPU clusters, implementing sharding, activation checkpointing and memory optimization (ZeRO, FSDP).
  • Build researcher-friendly tooling and job scheduling (Kubernetes, SLURM) with fast iteration, automated retries and failure recovery.
  • Design high-throughput pipelines that ingest terabytes of multimodal robot data (video, proprioception, 3D signals) so GPUs never starve.
  • Build low-latency inference pipelines for real-time robot control using quantization, distillation and compilation (TensorRT, Triton).
  • Profile GPU utilization, I/O bottlenecks and memory fragmentation to maximize fleet performance.

What You Bring

  • 5-7+ years as an infrastructure engineer, including leading technical projects in HPC or ML infrastructure
  • Built and maintained ML or data infrastructure on a team with a high talent bar
  • Deep PyTorch experience and hands-on large-scale distributed training (FSDP, ZeRO, failure recovery)
  • GPU performance optimization and profiling (CUDA, NCCL, Triton)
  • Genuine interest in robotics and physical AI
  • Ability to work in the Redwood City office 5 days a week

Nice to Have

  • Robotics experience at startups or enterprise teams
  • Early-stage or founding infrastructure hire
  • Multimodal systems (video, audio, multimedia models)
  • DeepSpeed or Accelerate; model serving optimization and monitoring

Interview Process

Recruiter screen (30 min), system design (45 min), two coding rounds (45 and 30 min), behavioral.

Tech Stack

PyTorch, DeepSpeed, Accelerate, FSDP, Kubernetes, SLURM, GCP, AWS, TensorRT, Triton, NCCL, Docker, Python, CUDA

Apply for ML Infrastructure Engineer

Your application is reviewed by a TSG search consultant.