SeniorStaff Applied AI Engineer, Agent Harness
San Francisco, CA - USA
Job Summary
Company: Init Intelligence
Location: San Francisco CA - in person Monday to Friday
Compensation: $200000 - $300000 base 1-2% equity
Employment Type: Full-time
Visa Sponsorship: Available (H-1B O-1 OPT)
Init is building AI coworkers for IT teams. Its security-focused product is an agent that registers as a governed identity in a customers directory requests scoped access for each task escalates to a human for approval runs in a fresh microVM with credentials the model never sees and can operate systems it was never given an API for.
Init was founded in 2026 by Isaiah De La Fuente and Sazzad Islam. The team comes from Stanford MIT and Delve among others.
This role builds the layer that turns model capability into systems that actually work for users. You will develop the core agent harness including the execution loop tool-use strategies context construction and model-facing experimentation and improve agent behavior across real customer workflows and long-horizon tasks.
A defining principle at Init: the creative step happens once when a workflow is authored and what runs afterward is deterministic compiled type-checked code rather than open-ended tool-chaining. You will own the boundary between what the model decides at runtime and what ships as code.
You will build and run evals against replicas of real customer environments reliable enough to gate a release. You will analyze production failures and trace them to the right layer (model prompt tool contract environment state or retry logic). And you will extend the computer-use agent to systems with no API owning the guarantees that let it run in production: a fresh microVM per task credentials injected at session start and never seen by the model every action recorded and any session stoppable mid-run.
- The core agent harness: execution loop tool-use strategies context construction and model-facing experimentation
- Agent behavior across real customer workflows and long-horizon tasks
- The boundary between runtime model decisions and deterministic compiled type-checked execution
- Evals against replicas of real customer environments reliable enough to gate releases
- Production failure analysis and systematic robustness improvements traced to the right layer
- Extending computer use to systems with no API with production guarantees (per-task microVM model-blind credentials full action recording stoppable sessions)
- Feedback loops and data systems that bring better real-task data into evaluation and training
- 4 years of experience
- Experience with agent frameworks or tool-using LLM systems
- Strong Python or TypeScript and comfort with modern AI tooling
- Experience with model evaluation fine-tuning or prompt design
- The ability to own systems end to end and debug across the stack
- A focus on systems and user outcomes not just model metrics
- You enjoy debugging messy real-world failures and turning them into improvements
- Ability to work in person in San Francisco Monday to Friday at startup intensity
- Built computer-use or browser-automation agents
- Experience with virtualization and sandboxed execution environments and scaling them
- AI research published at top conferences
- Shipped systems where correctness had to survive partial failure (idempotency resumability compensating actions)
- Integrated enterprise SaaS APIs (identity providers directory services ticketing ERPs)
- Experience with large messy datasets or production logs
- Been an early engineer somewhere you also talked to customers
Meals in office health insurance unlimited PTO.
Culture screen second culture screen technical interview then a one-day on-site.
Offered
$44000 - $66000
Net 60-90 day payout
- Senior/Staff agent harness
- Python/TS modern AI tooling
- Evals fine-tuning prompt design
- Exceptional signal of excellence in some facet: led an important team at a high-growth company top competition results (ICPC IOI) or elite achievement in a sport game or craft
- Top-tier schooling (Harvard Stanford Berkeley CMU Northeastern Georgia Tech UIUC UW Waterloo Oxford) for a CS/eng degree OR genuinely exceptional experience in lieu of it
- Worked at a great company on a relevant high-caliber team (team matters: agents/AI infra not ads)
- Known for good agents in production
- High agency: former founders big scope real ambiguity
- AI demos without production rigor; only implemented scoped features; depends on other teams for architecture infra or product decisions