Lead Machine Learning Engineer
New York City, NY - USA
Job Summary
The AI Engineering team is responsible for working closely with the research and modeling teams to create state-of-the-art NLP models for specifictasksand deploy them in a production setting designed to serve our customers at scale. We are looking for a Machine Learning Engineer to help build and evaluate the core intelligence behind our agentic AI systems. This role will play a key part in designing and owning evaluation frameworks that ensure quality safety and performance across complex agentic systems.
Were looking for a Lead Machine Learning Engineer to own and grow the evaluation platform that measures quality safety and performance across ASAPPs agentic AI systems- the infrastructure that tells us with confidence whether a model or agent change is actually an improvement before it reaches customers.
Thisa hybrid role with 10-12 days of in-office presence per month to balance flexibility with collaboration.
Help develop the technical roadmap and architecture for the evaluation platform from offline benchmarking to online/production monitoring of agentic and LLM-based systems.
Design eval methodologies appropriate to different stages of the pipeline: golden/regression test sets human-in-the-loop review workflows LLM-as-judge approaches and automated metrics for task success safety and hallucinations.
Build the data infrastructure evaluation depends on: annotation and labeling pipelines dataset versioning data quality checks and tooling that lets researchers and product teams run and interpret experiments without needing platform team help.
Partner closely with Research Product and Platform teams to productize experiments into robust AI solutions
Represent the eval platform to stakeholders outside the immediate team- set expectations on what good looks like for a model/agent release and report on platform health and coverage.
Stay current with advancements in ML NLP voice and LLM systems and contribute actively to technical discussions across teams.
Mentor and support other engineers through design reviews feedback and knowledge sharing.
Deep hands-on experience building and operating evaluation systems for modern ML/LLM/agentic systems- not just consuming existing eval tools.
Demonstrated experience leading the technical direction of a project or small team: setting architecture driving design reviews and being accountable for a systems long-term health (not just shipping features).
Strong architectural skills with proven experience designing complex data-intensive software systems and production experience with Python AWS Kubernetes and/or Docker.
Experience designing data pipelines for ML evaluation- labeling/annotation workflows dataset versioning and quality control and reproducible benchmarking.
A Bachelors Degree in CS or other related fields
Demonstrated technical mentorship of junior and mid-level engineers driving adoption of best practices and architectural alignment for scalability and extensibility.
Desire to learn teach and collaborate closely with cross-functional peers.
Experience building and evaluating agentic systems at scale.
Experience with voice/audio quality evaluations.
Production experience with LLM-centric services (e.g. inference orchestration evaluation monitoring)
Familiarity with large-scale ML experimentation benchmarking or simulation frameworks.
Experience with conversational/customer-support AI domains (e.g. containment rate conversation quality goal completion).
Knowledge of techniques for optimizing model architectures for faster inference.
Experience with AWS CI/CD Kafka Athena
Required Experience:
IC
About Company
Improve customer experience and radically increase CX performance at the same time. This AI-NativeĀ® software platform provides AI-driven predictions on what agents should and do throughout each interaction and increasingly automates routine tasks before, during, and after the conversa ... View more