Transparent Search Group
Member of Technical Staff, ML Systems
About the role
Company: Confidential - Seed-stage AI systems startup
Location: Menlo Park, CA (in office 5 days per week)
Compensation: $180,000 - $230,000 + competitive equity
Employment Type: Full-time
Visa Sponsorship: Visa transfers and new sponsorship (H-1B, TN)
About the Company
A seed-stage startup rebuilding the training and inference stack for video, image and world models, co-designing GPU kernels, distributed systems and models, with public benchmarks showing large speed and cost gains.
The Role
Our client is hiring ML systems engineers to own speed and efficiency across its stack: low-level kernels, distributed inference engines and multi-node training and serving systems. You report directly to the cofounder and CEO. The work sits below the application layer (kernels, runtimes and distributed engines for video and world models), so this is not an agents or RAG role. You will feel at home if you would rather make a video model ten times faster than train one.
What You Will Do
- Optimize GPU and system performance for training and inference on image, video and world-model workloads.
- Profile and remove bottlenecks at the kernel, memory, system and cluster level using Nsight and related tooling.
- Write low-level optimizations in CUDA and Triton on code paths that run in production.
- Build distributed inference and training engines for diffusion models across multiple GPUs and nodes.
- Own communication performance: NCCL, RDMA over InfiniBand or RoCE, and disaggregated serving.
- Build benchmarking and regression harnesses so performance gains hold in production.
What You Bring
- Authored core features in an inference or training framework (vLLM, SGLang, TensorRT-LLM, Megatron or similar), not only deployment or integration work
- 1+ year of full-time, hands-on ML systems or GPU performance work in your current or most recent role
- Kernel work on NVIDIA GPUs in CUDA, CUTLASS, Triton or PTX
- Inference or training performance experience: GPU kernels, runtimes, distributed execution
- Strong ML and CS fundamentals; degree in CS or a related quantitative field
- Ability to work in the Menlo Park office 5 days a week
Nice to Have
- Optimized diffusion, video, image or other multimodal workloads
- Public, verifiable kernel contributions to open-source ML systems
Interview Process
Hiring manager screen with the CEO (30 min), domain deep dive (60 min), system design (60 min), optional onsite.
Tech Stack
CUDA, Triton, PyTorch, Nsight, NCCL, RDMA
Your application is reviewed by a TSG search consultant.
More like this
Transparent Search Group
Member of Technical Staff
San Francisco, California, United States$200K–$300KView roleTransparent Search Group
AI Field Engineer - Enterprise
San Mateo, California, United States$176K–$224KView roleTransparent Search Group
Forward Deployed Engineer
San Francisco, California, United States$220K–$320KView roleTransparent Search Group
Software Engineer (C++ Systems)
San Francisco, California, United States$215K–$325KView role