Transparent Search Group
ML Infrastructure Engineer
About the role
Company: Confidential - Series A robotics company
Location: Redwood City, CA (in office 5 days per week)
Compensation: $220,000 - $350,000 + competitive equity
Employment Type: Full-time
Visa Sponsorship: Visa transfers (OPT, H-1B transfer)
About the Company
A well-funded Series A robotics company building general-purpose robots powered by its own embodied AI foundation model, already deployed at customer sites in hospitality and food service.
The Role
Our client is hiring an ML Infrastructure Engineer to own training infrastructure end to end and turn a multi-cloud GPU fleet into a world-class training engine for massive multimodal models. You will be the connective tissue between researchers and compute, and your work directly speeds the path from model to deployed robot.
What You Will Do
- Architect and scale distributed training across large GPU clusters, implementing sharding, activation checkpointing and memory optimization (ZeRO, FSDP).
- Build researcher-friendly tooling and job scheduling (Kubernetes, SLURM) with fast iteration, automated retries and failure recovery.
- Design high-throughput pipelines that ingest terabytes of multimodal robot data (video, proprioception, 3D signals) so GPUs never starve.
- Build low-latency inference pipelines for real-time robot control using quantization, distillation and compilation (TensorRT, Triton).
- Profile GPU utilization, I/O bottlenecks and memory fragmentation to maximize fleet performance.
What You Bring
- 5-7+ years as an infrastructure engineer, including leading technical projects in HPC or ML infrastructure
- Built and maintained ML or data infrastructure on a team with a high talent bar
- Deep PyTorch experience and hands-on large-scale distributed training (FSDP, ZeRO, failure recovery)
- GPU performance optimization and profiling (CUDA, NCCL, Triton)
- Genuine interest in robotics and physical AI
- Ability to work in the Redwood City office 5 days a week
Nice to Have
- Robotics experience at startups or enterprise teams
- Early-stage or founding infrastructure hire
- Multimodal systems (video, audio, multimedia models)
- DeepSpeed or Accelerate; model serving optimization and monitoring
Interview Process
Recruiter screen (30 min), system design (45 min), two coding rounds (45 and 30 min), behavioral.
Tech Stack
PyTorch, DeepSpeed, Accelerate, FSDP, Kubernetes, SLURM, GCP, AWS, TensorRT, Triton, NCCL, Docker, Python, CUDA
Your application is reviewed by a TSG search consultant.
More like this
Transparent Search Group
Senior Forward Deployed Engineer
Redwood City, California, United States$215K–$290KView roleTransparent Search Group
Senior Frontend Engineer
Redwood City, California, United States$200K–$300KView roleTransparent Search Group
Forward Deployed AI Engineer, Post-Sales
Redwood City, California, United States$230K–$300KView roleTransparent Search Group
Senior/Staff Applied AI Engineer, Agent Harness
San Francisco, California, United States$200K–$300KView role