All open positions

Transparent Search Group

Member of Technical Staff, Inference Systems

Posted today
Palo Alto, California, United States$230K–$350KAI Infrastructure

About the role

Company: Confidential - Seed-stage AI inference startup
Location: Palo Alto, CA (on-site 5 days per week)
Compensation: $230,000 - $350,000 + competitive equity
Employment Type: Full-time
Visa Sponsorship: Visa transfers and new sponsorship (H-1B, TN)

About the Company

A seed-stage startup building a next-generation AI inference platform from the ground up, engineered for performance with Rust at its core, founded by engineers with deep AI infrastructure experience.

The Role

Our client is hiring Members of Technical Staff (2+ years) to build a high-performance inference platform from scratch. You are a systems engineer who knows inference internals (attention, KV cache, batching, scheduling) and wants to own the whole stack rather than a narrow slice. You will join a small team building a new inference system in Rust, where you shape every design decision.

What You Will Do

  • Build a new inference runtime in Rust, owning batching, scheduling, request routing and the serving stack.
  • Design KV cache management, prefix caching and optimizations that reduce latency and cost per token.
  • Scale serving across GPUs and nodes.
  • Profile, benchmark and ship performance improvements across the inference pipeline.
  • Make core architecture decisions with the founding team.

What You Bring

  • 2+ years of systems engineering experience
  • Deep knowledge of inference internals: attention, KV cache, batching and scheduling
  • Experience inside inference engines such as vLLM, SGLang or TensorRT-LLM
  • Strong systems programming (Rust, C++ or similar)
  • Ability to work on-site in Palo Alto 5 days a week

Tech Stack

Rust, Python, PyTorch, C++, Go, vLLM, SGLang, TensorRT-LLM, CUDA, Triton, NCCL

Apply for Member of Technical Staff, Inference Systems

Your application is reviewed by a TSG search consultant.