Member of Technical Staff, Evals
San Francisco, CA - USA
Job Summary
Handshakes mission is to organize expert human knowledge to advance the AI economy. Handshake AI works directly with frontier labs on their most consequential data evaluation and post-training challenges building the systems that turn expert human knowledge into the data and evaluations that make frontier models better.
You will work alongside engineers researchers operators and builders from organizations including Scale AI Meta Google Amazon xAI Notion and Palantirand help build the systems that make expert human knowledge useful for advancing AI.
We are hiring a Member of Technical Staff Evals to help define how frontier AI systems are measured understood and improved. This is a broad high-ownership role for researchers who build.
You will partner with AI researchers domain experts and customers to develop new benchmarks reward and verifier systems agent-evaluation methodologies and data-quality techniques. You will work on questions at the center of frontier AI progress: what should be measured how to design evaluations that reflect real capability how to create high-signal feedback and how to build the environments and data systems that make those answers actionable.
Early members of the team will have unusual influence over our technical direction operating culture and the open-source software benchmarks and research products we build. We care more about demonstrated research capability technical judgment and a builders mindset than a specific title degree or career path.
Location: San Francisco preferred; we are open to exceptional candidates in other locations.
Design and build evaluation frameworks benchmarks and methodologies for frontier LLMs AI agents multimodal models and reinforcement-learning environments.
Develop reward models programmatic verifiers graders and other feedback systems that make model behavior measurable and improvable.
Research what makes evaluations representative difficult reliable and resistant to shortcutting or reward hacking.
Build systems for high-quality human data including expert task design annotation methodologies data-quality signals and data-attribution techniques.
Run fast rigorous iteration loops: prototype evaluate interpret results diagnose failure modes and turn learnings into the next benchmark or system.
Publicly contribute to the field through benchmarks open-source tools research and technical writing.
PhD in ML/AI computer science data science or related fields (or equivalent research experience in industry).
Publications at top AI/ML venues like NeurIPS ICML ICLR COLM.
Builders who enjoy tinkering with agents and shipping high-quality software benchmarks or datasets (e.g. a strong GitHub profile / OSS contributions or product portfolio).
Strong Python skills experience building scalable software working with agents.
Strong knowledge of frontier AI: benchmarks eval techniques agent harnesses post-training recipes data shapes.
Comfort operating in an ambiguous fast-moving environment with substantial ownership.
Work at the very frontier of AI with most major AI labs researching some of the most important problems in Data and Evaluations.
Publish results and work in public through open-source benchmarks and software papers and blogs.
Join a rapidly growing company whose data business grew from zero to nearly $1B run rate in a year.
Help build an early technical organization where your work shapes the roadmap standards and culture.
Attend (and publish at) conferences like NeurIPS ICML ICLR COLM.
Handshake delivers benefits that help you feel supportedand thrive at work and in life.
The below benefits are for full-time US employees.
Ownership: Equity in a fast-growing company
Financial Wellness: 401(k) match competitive compensation financial coaching
Family Support: Paid parental leave fertility benefits parental coaching
Wellbeing: Medical dental and vision mental health support $500 wellness stipend
Growth: $2000 learning stipend ongoing development
Office: Commuting support free lunch and gym in our SF office
Time Off: Flexible PTO 15 holidays 2 flex days
Connection: Team outings & referral bonuses
About Company
The better career platform for Gen Z changing how, where, and why the next generation of talent builds their career.