Founding AI Systems Engineer
San Francisco, CA - USA
Job Summary
Agent systems ML systems distributed systems agent runtime evals
Build the learning layer that lets agents carry experience forward.
Location: San Francisco five days per week in person
Base salary: $180000$300000
Equity: 0.5%2.0% target initial ownership on a fully diluted basis
Relocation: Up to $15000 for candidates moving to the Bay Area
Benefits: $1000 monthly healthcare stipend initially plus flexible PTO
Visa: Sponsorship considered case by case
Travel: Up to 10%
Frontier models can reason use tools and complete increasingly complex work. But the systems around them still lose the most valuable part of doing the work: what decision was made what happened next which correction mattered and where that lesson should apply.
RAG retrieves information. LLABS is building systems that learn from experience and carry it forward. We connect decisions actions outcomes and feedback so useful experience can shape later work.
Models will keep changing. The experience an organization earns should not disappear with each new model or session. LLABS is building a durable learning layer that connects decisions actions outcomes and feedback so future agents can inherit reviewed experience and use it in the right context. That layer can become enduring infrastructure between frontier intelligence and the institutions putting it to work.
We are hiring a founding engineer to make that architecture reliable under real operating conditions. Your first mission is to make the runtime measurably reliable across state recovery evaluation and observability. From there you will deepen the experience-learning loop and turn successful operational work into reusable platform capabilities.
You will work directly with Brayden shape the runtime and evaluation architecture and help build the engineering team around it. The path from a research idea to production evidence is short: you can test a systems hypothesis see how it survives real operating constraints and turn what works into the foundation of the company.
LLABS has built an agent runtime tool system governed experience layer and customer-review surface. The current engineering substrate is primarily Python and TypeScript with API services a React/ product surface relational and cache/storage layers event-driven execution containerized cloud infrastructure and model-provider interfaces.
Those components serve one architectural bet: the model can change but useful operational experience should persist above it and remain scoped to the right user role workflow and permission boundary.
We call the learning architecture Causal Trajectory Learning. It keeps context decisions actions outcomes and feedback connected so reviewed experience can improve later work in the right scope. The next phase is to make that learning loop easier to evaluate operate recover and trust.
Agent lifecycle and recovery across request-scoped and long-running execution with identity and state preserved through pause resume failure and termination.
The experience and evaluation surfaces that connect decisions actions outcomes and feedback across sessions and model changes.
Observability root-cause diagnosis release gates rollback and recovery across models memory tools permissions runtime state and product behavior.
The technical path from a selected enterprise workflow into standard connectors tests telemetry and reusable platform capabilities.
Tenant isolation authentication secrets data boundaries and the operational controls required for consequential enterprise work.
Learning-system experiments expert acceptance criteria and the translation of evaluation results into research decisions.
Workflow mapping and technical reviews with the people responsible for the customer relationship commercial outcome and trust decisions.
You will be a founding technical lead reporting directly to founder Brayden Levangie. You will shape the runtime and evaluation architecture help hire the engineering team and have direct authority over the assigned technical path. Brayden built one of the earliest production RAG platforms before the category had a name architected the LLABS runtime and sold the companys first pilots.
Most of your time will go to the core platform and its research and evaluation loop. Selected customer workflows create the empirical pressure that turns promising ideas into dependable platform capabilities.
On-call is shared with the founder. Deployment work feeds reusable improvements into the common platform.
The process includes two focused conversations followed by a paid working session. We will discuss mutual fit and systems you have personally built then agree on a staging problem compensation and what strong work looks like.
What state must survive when an agent works for hours across files APIs browsers databases and subordinate agents
What should an agent preserve from a successful trajectory and what evidence shows that lesson generalizes beyond the original session or workflow
How do we find the component that owns the broken state when the UI says a job is running but execution has stopped somewhere else
What evaluation predicts whether a domain expert will accept the completed work
How should feedback change later behavior while preserving workflow and tenant boundaries
System: Important agent lifecycle and execution paths are observable repeatably tested and recoverable. Failures can be isolated to the right layer quickly.
Workflow: one selected workflow has a dependable LLABS-owned path to production acceptance with the critical permissions evaluations release checks and recovery behavior made explicit.
Compounding: at least one connector evaluation telemetry runtime or recovery capability learned from that workflow is reused in a second workflow.
Learning: a reviewed outcome improves later behavior in the intended workflow and permission boundary.
Team leverage: the engineer independently diagnoses and unblocks that path increasing founder leverage.
We care most about what you have personally made work.
You have built or operated a consequential distributed system ML platform developer platform agent runtime or other stateful production system.
You can show a failure you traced across several layers and explain the evidence that changed your diagnosis.
You have designed an evaluation experiment or observability system that changed a product or architecture decision.
You are comfortable moving between Python application code APIs queues databases containers infrastructure and a TypeScript product surface.
You treat authentication permissions isolation secrets recovery and release safety as product behavior.
You have worked directly with demanding users or customers and turned their reality into a general system.
You can disagree clearly simplify a system and stay hands-on when the failure is ambiguous.
You want research claims to survive production evidence and you want production evidence to change the research.
Research papers open-source work founding experience and unusual side projects can all be strong evidence.
LLABS exists to compress the distance between a breakthrough and the work it changes. Enterprise environments are our proving ground because they concentrate the conditions a learning system has to survive: old software long-running state strict permissions failure and expert judgment.
Research questions runtime behavior evaluation and deployment evidence therefore live in the same engineering loop. That loop connects experiments about experience representation with runtime failures expert-designed evaluations and reusable platform changes learned from deployment.
We want that learning infrastructure to help expert teams move faster on difficult scientific industrial and institutional problems with each new agent and model generation inheriting the experience earned before it.
This role sits at the center of that work. You will help build a system that survives real tools permissions failures corrections and users. Each experience should make it more capable.
Required Experience:
IC