AI Evaluations Engineer, US Decision Intelligence
Cupertino, CA - USA
Job Summary
Were seeking a visionary AI Evaluations Engineer to own the end-to-end evaluation pipeline for our AI products and agentic workflows. This role will focus on implementing and maintaining evaluation frameworks instrumentation and workflows that help us understand how well our AI systems perform where they fail and how they improve over time. You own the evaluation gate and the role will operate in both capacities to augment existing AI roadmap as well as innovate and trailblaze new frontier-technology projects crafting AI experiences that reduce time to insight and catalyze decision making.
Architect a comprehensive Evals framework to trace at every layer from agent responses breakdown to skill level and tool calling with the main objective to optimize for accuracy and performance improving latency and running AB tests to find the best recommendation across the and operate AI evaluation workflows that measure the quality of LLM outputs across chat summarization recommendations and agentic rubric-based evals to score outputs for correctness relevance grounding and beyond LLM-as-a-judge to agent-as-a-judge including harness-as-a-judge patterns applied against real traces in a LLM and agent workflows to capture traces prompts and responses metadata and user the platform-wide eval gate: define the pass/fail contract every track ships against and hold a release when it isnt define agent-specific evaluations (task completion tool correctness error recovery).nPartner with AI engineers and AI platform teams to translate product requirements into evaluation eval disputes with track leads and own the standard that geo eval engineers implement rerun policy and variance with the Evals team in India to maximize global to system design for observability retries and logging.
5 years of experience in data and AI-related fields such as AI engineering software development ML engineering data science or QA and ability to learn new skills and solve dynamic problems in an encouraging and expansive Python -on experience with AI evaluation techniques such as Golden datasets LLM-as-a-Judge or rubric-based with different LLM ecosystems (OpenAI Anthropic Gemini etc.) RAG pipelines vector databases (e.g. Pinecone FAISS Milvus PostgreSQL).nProficiency in SQL and experience with at least one major data analytics platform such as Hadoop Spark or with CI/CD or release validation working with data science teams on insights generation leveraging time management skills with the ability to collaborate across multiple to balance competing priorities long-term projects and ad hoc to work in a fast-paced dynamic constantly evolving business -on experience with Langfuse or similar tools for LLM working with product/domain experts to translate fuzzy correctness criteria into measurable rubrics or .S. degree in Computer Science/Engineering or equivalent work experience
Sound communication skills - expert at messaging domain and technical content at a level appropriate for the audience. Strong ability to gain trust with stakeholders and senior with embeddings retrieval algorithms agents and data modeling for vector and graph complementary technologies for distributed systems architecture and asynchronous messaging agent communication and caching like RabbitMQ Redis and Valkey are working across global teams to ensure alignment of product knowledge of GenAI and RAG strategies microservices recommendation systems and context knowledge of agent evaluation concepts like trajectory vs. end-to-end vs. component-level evaluation tool-call degree (MS or Ph.D.) in Economics Electrical Engineering Statistics Data Science or a similar quantitative field is preferred.
Required Experience:
IC
About Company
Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more