Staff Experience Designer, AI Evaluation Platform
New York City, NY - USA
Job Summary
Teams building AI features need to know if what they shipped works. Today that means navigating unfamiliar territory: scoring something non-deterministic telling a real regression from noise trusting a judgment a model made instead of a person. Most existing tools here were built for the researchers who invented these methods not for the people who now need to use them set the direction here and build the design system everyone else builds on. Wed rather get a rough version in front of real users than a polished one later. And because teams use this platform to decide what to ship trust matters more than usual: provenance uncertainty and edge cases determine whether someone acts on a result or quietly stops believing makes this team unusual is its interdisciplinary core. Alongside the platform we run a research group working on evaluation methodology itself including how to tell whether an evaluator is calibrated biased or measuring what it claims to. Their methods ship into this platform and you decide how anyone first encounters them. Youll be close to that work while its still forming instead of picking it up once its finished.
Design the path from a vague question about a systems quality or safety to a running evaluation without requiring evaluation dense multi-dimensional results legible enough to act on: scores comparisons distributions and the uncertainty attached to how someone inspects multi-step agent behavior traces a failure to its cause and audits a judgment a model with our research scientists as new evaluation techniques are being developed so you shape how they surface in the product instead of designing around them after theyre set. Youll translate these techniques into interfaces non-specialists can use and and maintain the platforms design system as the first designer on the against real data to validate ideas before theyre research directly with the engineers scientists and practitioners who use the platform.
8 years of experience designing digital products or experiences including end-to-end ownership of complex data-dense products for technical users. Youve set direction in spaces with no existing precedent and you get to clarity by running research yourself with the people who use what you portfolio you can share including at least one case study that walks through your process from problem framing to shipped ability to design complex information: dashboards comparison views large result sets and results that carry statistical interaction and visual design skills with a high bar for craft in dense information-heavy fluency in AI/ML concepts (benchmarks metrics model-based judging agentic systems) to work directly with a research scientist or with AI tools in your own practice. You use Claude Code or equivalents to build working prototypes and extend what you can make on your communication skills with the ability to build buy-in across engineering and research without formal authority.
Experience as the first or only designer on a platform and/or experience translating research output into shipped building and maintaining a design system including its components patterns and naming AI-first instinct for interface design: shipped interfaces organized around stated intent or surfaces meant to be operated by agents as well as designing coherent workflows across multiple surfaces (UI CLI SDK) that need to stay consistent with each with modern evaluation observability or visualization tooling (e.g. LangSmith Braintrust D3 Vega-Lite).nComfort with SQL and notebooks to explore your own data.
Required Experience:
Staff IC
About Company
Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more