Prompt & Evaluation Engineer
Job Summary
Convo is looking for a Prompt & Evaluation Engineer with 5 years of experience in applied NLP/LLM engineering to design and optimize prompt chains agent behaviors and evaluation frameworks. The ideal candidate should have strong expertise in prompt engineering LLM evaluation Python structured outputs regression testing and failure analysis with the ability to translate business requirements and commercial use cases into reliable measurable AI solutions.
Engineer prompt chains agent constitutions and evaluation suites that convert governed CPG semantics into reliable commercial-agent behavior.
Design modular prompt chains and structured-output contracts for commercial analysis and advisory workflows.
Encode agent constitutions business constraints and escalation rules defined by the Guild.
Create representative evaluation sets scoring rubrics and regression benchmarks.
Implement Python-based evaluation harnesses and automated failure analysis.
Version prompts and evaluation assets track changes and prevent regressions across model updates.
Analyze errors with domain experts and convert findings into prompt data or workflow improvements.
Minimum experience: 3 years in applied NLP or LLM engineering including at least 1 year building systematic evaluation or regression workflows.
Hands-on prompt engineering for multi-step tool-using or structured-output LLM applications.
LLM evaluation design including rubric-based deterministic and model-graded approaches.
Python scripting for evaluation pipelines and data analysis.
Commercial analytics literacy and ability to reason about KPIs constraints and business decisions.
Experience building regression datasets and diagnosing model/prompt failures.
Version control and reproducible experiment practices.
Agent frameworks tool calls and retrieval-augmented generation.
CPG pricing trade category or RGM use cases.
Evaluation telemetry prompt registries and experiment-tracking tools.
Versioned prompt-chain and agent-constitution library.
Evaluation datasets rubrics and automated regression suite.
Benchmark results and failure taxonomy by commercial use case.
Release criteria and change documentation for approved prompt assets.
Works with the CPG Domain Expert Agent Runtime Agentic Governance application teams and AI Test & Dataset Engineering.
What We Have For You
- Great compensation package medical benefit for you and your family free lunch annual performance-tied increments & performance recognition awards and a great lean and agile work culture!
- Convo endorses a culture of diversity in all aspects and aims to build a diverse team of amazing individuals!
Required Experience:
IC
About Company
Our team of experienced consultants work closely with clients to understand their specific needs and goals, and provide customized solutions to help them leverage the power of Artificial Intelligence to optimize processes, gain insights, and make data-driven decisions to help achieve ... View more