Transparent Search Group
Machine Learning Engineer - Infrastructure
About the role
Company: Confidential - Venture-backed AI research company
Location: San Francisco, CA (South Park office, in person 5 days per week; relocation provided)
Compensation: $200,000 - $400,000 + highly competitive early-stage equity
Employment Type: Full-time
Visa Sponsorship: Visa transfers; can sponsor visas
About the Company
A venture-backed research company building a physics foundation model to learn cause and effect, starting with weather, the most observed physical system on earth. The founders come from leading self-driving and AI research teams.
The Role
Our client is hiring infrastructure engineers to tackle the unsolved training and inference challenges of a physics foundation model. The work demands deep expertise in standing up distributed training clusters and optimizing performance for large models. If you have built large-scale ML infrastructure for language, vision, robotics or biology models and want to bet on a counterintuitive technical thesis, this is the role.
What You Will Do
- Design, deploy and maintain large distributed ML training and inference clusters.
- Build efficient, scalable end-to-end pipelines for petabyte-scale datasets and model training across the ML lifecycle.
- Research and test training approaches, including parallelization techniques and numerical-precision trade-offs across model scales.
- Analyze, profile and debug low-level GPU operations to optimize performance.
- Bring new ideas from current research into the stack.
What You Bring
- 2-10 years building large-scale ML infrastructure for core foundation models (not fine-tuning)
- Deep expertise optimizing large-scale training and inference workloads
- Proficiency with distributed training frameworks (FSDP, DeepSpeed)
- Knowledge of cloud platforms (GCP, AWS or Azure) and containers/orchestration (Kubernetes, Docker)
- Experience with distributed task management and scalable model serving architectures
- Strong grasp of monitoring, logging and observability for ML systems
- Ability to work in person in San Francisco 5 days a week
Nice to Have
- Experience at a science or physical AI company (self-driving, robotics, biology, climate/weather)
- Generalist experience across the ML lifecycle
- Low-level GPU performance optimization and debugging (CUDA, JAX)
Interview Process
Initial call (30 min), technical screen, onsite day.
Tech Stack
FSDP, DeepSpeed, NVIDIA GPUs, Python, C++, Linux, Kubernetes, Docker, GCP/AWS/Azure
Your application is reviewed by a TSG search consultant.
More like this
Transparent Search Group
Member of Technical Staff, Head of Quality
San Francisco, California, United States$180K–$280KView roleTransparent Search Group
Strategic Project Lead
San Francisco, California, United States$175K–$275KView roleTransparent Search Group
Senior Recruiter, Finance Experts
San Francisco, California, United States$150K–$200KView roleTransparent Search Group
Machine Learning Researcher
San Francisco, California, United States$200K–$200KView role