Enter a job title or keyword

Principal Machine Learning Engineer

NextDeavor


Job Location:

New York City, NY - USA

Yearly Salary: USD 200000 - 250000
Posted: 7 August 2026 (27 days ago)
Application Deadline: 4 November 2026
Vacancies: 1 Vacancy

Job Summary

Principal Machine Learning Engineer
Full-time
New York City NY US
Exclusive confidential search details shared with qualified applicants.
Become a Key Player as a Principal Machine Learning Engineer

You will own the ML infrastructure that turns research into reliable real-time compliance enforcement systems driving model training evaluation and production serving. You will partner closely with research stakeholders and engineering peers to ship reproducible pipelines and low-latency serving; the role is Hybrid (3 days onsite) in the New York City Metro area.

Heres How Youll Make an Impact on the Team
  • Build and own training pipelines: data preparation reproducible fine-tuning runs experiment tracking and release automation
  • Build evaluation infrastructure: automated eval runs regression gates dashboards and dataset versioning
  • Own model serving in production: low-latency inference batching optimization autoscaling and cost management
  • Ship model updates safely with versioning canarying rollback and drift monitoring
  • Create repeatable workflows to adapt models to new domains and customer needs
  • Turn expert labels and reviewer feedback into clean training and evaluation data
  • Set the engineering bar for ML infrastructure as the team grows
Heres What Youll Need to Be Successful in This Role
  • 8 years of software engineering experience including 4 years building infrastructure for ML or LLM systems in production
  • Hands-on experience with the modern LLM stack: PyTorch distributed training fine-tuning at scale (e.g. LoRA SFT) and inference engines such as vLLM or TensorRT-LLM
  • Experience building eval harnesses regression gates or dataset pipelines; strong understanding of precision recall and calibration
  • Proven ownership of production model serving with real latency reliability and cost constraints
  • Strong fundamentals in Python containers CI/CD cloud infrastructure and observability
  • Ability to scope work ship frequently and make pragmatic build-vs-buy decisions
  • Experience collaborating tightly with research partners and defining clear interfaces
Heres What Else Might Help You Out
  • Experience productionizing small or specialized language models
  • Experience with structured-output serving or constrained decoding in production
  • Prior work in regulated or high-stakes domains (fintech healthcare legal trust and safety)
  • Experience deploying models into customer-controlled environments
Pay Range

$200K - $250K/year

Ready to Make Your Mark

This role may fill quickly. Submit your resume to be considered.

Apply with Pioneers here


Required Experience:

Staff IC


About Company

Hire trusted candidates who BELONG STAY ADVANCE NextDeavor is a recruiting agency helping companies make more strategic hiring decisions. FIND YOUR NEXT GREAT HIRE Using AI technology to make the recruiting process more human AI speeds up, refines, and expands our initial search. This ... View more

View Profile View Profile