All open positions

Transparent Search Group

Senior/Staff Applied AI Engineer, Agent Harness

Posted today
San Francisco, California, United States$200K–$300KAI / Security

About the role

Company: Confidential - Venture-backed AI security startup
Location: San Francisco, CA - in person, Monday to Friday
Compensation: $200,000 - $300,000 base + 1-2% equity
Employment Type: Full-time
Visa Sponsorship: Available (H-1B, O-1, OPT)

About the Company

A venture-backed San Francisco startup building AI coworkers for IT teams: security-first agents that act as governed identities with scoped, human-approved access in isolated environments.

The Role

This role builds the layer that turns model capability into systems that actually work for users. You will develop the core agent harness, including the execution loop, tool-use strategies, context construction and model-facing experimentation, and improve agent behavior across real customer workflows and long-horizon tasks.

A defining principle at the company: the creative step happens once, when a workflow is authored, and what runs afterward is deterministic, compiled, type-checked code rather than open-ended tool-chaining. You will own the boundary between what the model decides at runtime and what ships as code.

You will build and run evals against replicas of real customer environments, reliable enough to gate a release. You will analyze production failures and trace them to the right layer (model, prompt, tool contract, environment state or retry logic). And you will extend the computer-use agent to systems with no API, owning the guarantees that let it run in production: a fresh microVM per task, credentials injected at session start and never seen by the model, every action recorded, and any session stoppable mid-run.

What You Will Own

  • The core agent harness: execution loop, tool-use strategies, context construction and model-facing experimentation
  • Agent behavior across real customer workflows and long-horizon tasks
  • The boundary between runtime model decisions and deterministic, compiled, type-checked execution
  • Evals against replicas of real customer environments, reliable enough to gate releases
  • Production failure analysis and systematic robustness improvements, traced to the right layer
  • Extending computer use to systems with no API, with production guarantees (per-task microVM, model-blind credentials, full action recording, stoppable sessions)
  • Feedback loops and data systems that bring better real-task data into evaluation and training

What You Bring

  • 4+ years of experience
  • Experience with agent frameworks or tool-using LLM systems
  • Strong Python or TypeScript and comfort with modern AI tooling
  • Experience with model evaluation, fine-tuning or prompt design
  • The ability to own systems end to end and debug across the stack
  • A focus on systems and user outcomes, not just model metrics
  • You enjoy debugging messy, real-world failures and turning them into improvements
  • Ability to work in person in San Francisco, Monday to Friday, at startup intensity

Nice to Have

  • Built computer-use or browser-automation agents
  • Experience with virtualization and sandboxed execution environments, and scaling them
  • AI research published at top conferences
  • Shipped systems where correctness had to survive partial failure (idempotency, resumability, compensating actions)
  • Integrated enterprise SaaS APIs (identity providers, directory services, ticketing, ERPs)
  • Experience with large, messy datasets or production logs
  • Been an early engineer somewhere you also talked to customers

Benefits

Meals in office, health insurance, unlimited PTO.

Interview Process

Culture screen, second culture screen, technical interview, then a one-day on-site.

Apply for Senior/Staff Applied AI Engineer, Agent Harness

Your application is reviewed by a TSG search consultant.