Enter a job title or keyword

Principal Security Research Manager

Microsoft


Job Location:

Redmond, WA - USA

Yearly Salary: USD 142800 - 274800
Posted: 5 October 2026 (21 hours ago)
Application Deadline: 2 January 2027
Vacancies: 1 Vacancy

Job Summary

Overview

We are building the evaluation backbone for safe reliable and efficient agentic engineering in Microsoft Security. This team will create systems that determine when an AI agent model prompt tool memory strategy or orchestration pattern is ready to be used in production security and engineering workflows. The role is ideal for engineers who can operate end to end: understand the workflow design the benchmark build the harness implement validators and graders run experiments analyze quality and cost tradeoffs connect results to production feedback and help teams make evidence-based release decisions.

Why this role matters:Microsoft Securityis moving toward agentic engineering systems for security triage remediation repo readiness and scan-to-verified-closure workflows. Evals are the trust system for that shift. They help decide whether autonomy can safely expand whether a release should stop and which configuration achieves the required quality safety reliability latency and cost bar with the lowest practical human-review burden.

Role mission

As a Principal Security Research Manager on the AI Evaluation Systems team you will build the common evaluation platform and methodology used by MSec agent programs. You will work across evaluation design platform implementation test infrastructure telemetry measurement security workflow understanding and production learning. Your work will make agentic systems measurable reproducible governable and continuously improving.

Microsofts mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset innovate to empower others and collaborate to realize our shared goals. Each day we build on our values of respect integrity and accountability to create a culture of inclusion where everyone can thrive at work and beyond.



Responsibilities
  • Design and build end-to-end evaluation harnesses for agentic security and engineering workflows including triage remediation repo readiness escalation tool use and scan-to-verified-closure paths.

  • Create representative benchmark suites and golden datasets that include normal edge adversarial failure-recovery regression and production-derived cases.

  • Implement deterministic validators automated graders trace analyzers result stores comparison views and workflow adapters that make evaluations repeatable and actionable.

  • Measure task success correctness safety and policy compliance failure recovery latency tool-call behavior token usage total cost per successful outcome and human-review effort.

  • Compare models prompts tools memory strategies policies and orchestration patterns under consistent conditions and help teams understand quality-versus-efficiency tradeoffs.

  • Integrate evaluations into engineering workflows CI/CD release gates and decision processes so material agent changes are supported by reproducible evidence before production rollout.

  • Connect offline evaluation results with production feedback including accepted and rejected outputs human overrides incidents rollbacks escaped defects and customer or service-health signals.

  • Detect regressions benchmark drift evaluator miscalibration and cases where eval scores improve while real-world outcomes do not.

  • Partner with agent builders product teams security engineers data scientists program managers and leadership to turn evaluation results into release recommendations and autonomy-boundary decisions.

  • Convert learnings into reusable paved paths: playbooks templates onboarding guides dashboards scorecards and reference implementations that can scale across MSec.

What you will build

  • Versioned benchmark suites for security triage remediation repo readiness tool use escalation and end-to-end agentic workflows.

  • Evaluation runners and harnesses that can replay tasks capture traces evaluate outputs and compare multiple agent configurations.

  • Deterministic validation checks for code policy security provenance ownership deployment constraints and workflow-specific correctness.

  • Automated and human-in-the-loop grading pipelines with calibration sampling and confidence thresholds.

  • Pareto-style scorecards that show tradeoffs across quality risk latency tokens cost and human-review burden.

  • Telemetry and production feedback loops that continuously expand benchmark coverage and keep offline evaluation anchored to real-world outcomes.

  • Release-readiness gates and evidence packages that help leaders and product teams decide whether to scale stop or redesign an agent pattern.



Qualifications

Required Qualifications:

  • Doctorate in Statistics Mathematics Computer Science Computer Security or related field AND 3 years experience in software development lifecycle large-scale computing threat analysis or modeling cybersecurity vulnerability research and/or anomaly detection
    • OR Masters Degree in Statistics Mathematics Computer Science Computer Security or related field AND 4 years experience in software development lifecycle large-scale computing threat analysis or modeling cybersecurity vulnerability research and/or anomaly detection
    • OR Bachelors Degree in Statistics Mathematics Computer Science Computer Security or related field AND 6 years experience in software development lifecycle large-scale computing threat analysis or modeling cybersecurity vulnerability research and/or anomaly detection
    • OR equivalent experience.
  • 1 year(s) people management experience.

Preferred Qualifications:

  • Doctorate in Statistics Mathematics Computer Science Computer Security or related field AND 5 years experience in software development lifecycle large-scale computing threat analysis or modeling cybersecurity vulnerability research and/or anomaly detection
    • OR Masters Degree in Statistics Mathematics Computer Science Computer Security or related field AND 8 years experience in software development lifecycle large-scale computing threat analysis or modeling cybersecurity vulnerability research and/or anomaly detection
    • OR Bachelors Degree in Statistics Mathematics Computer Science Computer Security or related field AND 12 years experience in software development lifecycle large-scale computing threat analysis or modeling cybersecurity vulnerability research and/or anomaly detection
    • OR equivalent experience.
  • Proven software engineering experience building production systems developer platforms test infrastructure automation frameworks data pipelines quality systems or reliability tooling.
  • Ability to design and implement evaluation systems end to end including task definition dataset creation harness implementation scoring analysis and operational integration.
  • Demonstrated coding debugging system design and operational excellence skills.
  • Experience working with structured data logs traces metrics APIs automation workflows and engineering telemetry.
  • Experience with LLMs AI agents model evaluation prompt/tool orchestration automated grading evaluation harnesses or benchmark design.
  • Experience withexperimentation statistical confidence evaluator calibration regression analysis human-review protocols or quality measurement systems.
  • Experience with security engineering vulnerability management SAST/SCA SARIF remediation workflows secure development lifecycle or compliance-sensitive systems.


Security Research M5 - The typical base pay range for this role across the U.S. is USD $142800 - $274800 per year. There is a different range applicable to specific work locations within the San Francisco Bay area and New York City metropolitan area and the base pay range for this role in those locations is USD $188000 - $304200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
position will be open for a minimum of 5 days with applications accepted on an ongoing basis until the position is filled.



Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age ancestry citizenship color family or medical care leave gender identity or expression genetic information immigration status marital status medical condition national origin physical or mental disability political affiliation protected veteran or military status race ethnicity religion sex (including pregnancy) sexual orientation or any other characteristic protected by applicable local laws regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process read more about requesting accommodations.


Required Experience:

Manager