All open positions

Transparent Search Group

Machine Learning Engineer - Infrastructure

Posted today
San Francisco, California, United States$200K–$400KAI Research

About the role

Company: Confidential - Venture-backed AI research company
Location: San Francisco, CA (South Park office, in person 5 days per week; relocation provided)
Compensation: $200,000 - $400,000 + highly competitive early-stage equity
Employment Type: Full-time
Visa Sponsorship: Visa transfers; can sponsor visas

About the Company

A venture-backed research company building a physics foundation model to learn cause and effect, starting with weather, the most observed physical system on earth. The founders come from leading self-driving and AI research teams.

The Role

Our client is hiring infrastructure engineers to tackle the unsolved training and inference challenges of a physics foundation model. The work demands deep expertise in standing up distributed training clusters and optimizing performance for large models. If you have built large-scale ML infrastructure for language, vision, robotics or biology models and want to bet on a counterintuitive technical thesis, this is the role.

What You Will Do

  • Design, deploy and maintain large distributed ML training and inference clusters.
  • Build efficient, scalable end-to-end pipelines for petabyte-scale datasets and model training across the ML lifecycle.
  • Research and test training approaches, including parallelization techniques and numerical-precision trade-offs across model scales.
  • Analyze, profile and debug low-level GPU operations to optimize performance.
  • Bring new ideas from current research into the stack.

What You Bring

  • 2-10 years building large-scale ML infrastructure for core foundation models (not fine-tuning)
  • Deep expertise optimizing large-scale training and inference workloads
  • Proficiency with distributed training frameworks (FSDP, DeepSpeed)
  • Knowledge of cloud platforms (GCP, AWS or Azure) and containers/orchestration (Kubernetes, Docker)
  • Experience with distributed task management and scalable model serving architectures
  • Strong grasp of monitoring, logging and observability for ML systems
  • Ability to work in person in San Francisco 5 days a week

Nice to Have

  • Experience at a science or physical AI company (self-driving, robotics, biology, climate/weather)
  • Generalist experience across the ML lifecycle
  • Low-level GPU performance optimization and debugging (CUDA, JAX)

Interview Process

Initial call (30 min), technical screen, onsite day.

Tech Stack

FSDP, DeepSpeed, NVIDIA GPUs, Python, C++, Linux, Kubernetes, Docker, GCP/AWS/Azure

Apply for Machine Learning Engineer - Infrastructure

Your application is reviewed by a TSG search consultant.