Benchmarking Project Lead, Siri Evaluation
Cambridge, MA - USA
Job Summary
As part of the work on next generation Siri we are developing novel measurements of its quality. To ensure that the evaluation systems we are building are reliable we plan benchmarking their accuracy on a wide range of features locales and platforms using humans in the loop.
Design and execute efficient data collection processes using humans in the loop nLead annotation efforts for various languages and devicesnPlan and manage the budget of annotator resourcing with Finance Annotation Ops and International team and improve the quality of the human judgements through revision of annotator training materials clear annotation questions annotator training and efficient reviewing mechanismsnCollaborate with other engineering teams to design and build a tooling ecosystem for managing and browsing rich datasets n
Agentic Coding proficiency to achieve data-science data collection and visualisation tasksnGood understanding of metrics crowd science data collection annotation analysis statisticsnAbility to work independently and cross-functionally to integrate in partner team reporting systems and pipelinesnExcellent communication skills and the ability to thrive in a highly collaborative work environment
Attunement to computational linguistics language quality human in the loop evaluationnGood engineering practices to create sustainable and easy to use data management pipelinesnPython experience and other tools for data collection and visualisation
Required Experience:
Senior IC
About Company
Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more