Senior Machine Learning Engineer Intelligent Document Processing Production AI Systems
Reston, VA - USA
Job Summary
Company Overview
Pantheon Data (a Kenific Holding company) is a private small business based in the Washington DC area. Pantheon Data was founded in 2011 initially providing acquisition and supply chain management services to the US Coast Guard. Our service offerings have grown in the past ten years including infrastructure resiliency contact center operations information technology software engineering program management strategic communications engineering and cybersecurity. We have also grown our customer base to include commercial clients. The company has used this experience to expand our service offerings to other agencies within the Department of Homeland Security (DHS) the Department of Defense (DoD) and other Federal Civilian Agencies.
Position Overview
We are seeking a Senior Machine Learning Engineer to help build production-oriented AI and Intelligent Document Processing (IDP) systems. This role is for a hands-on engineer who can move beyond experiments and build working software that ingests processes analyzes retrieves and explains information from complex unstructured and semi-structured sources.
The ideal candidate has real depth in machine learning NLP OCR computer vision LLMs and retrieval-based systems but also has the broader engineering judgment to understand the system around the model: data pipelines APIs databases cloud infrastructure containers testing evaluation observability and production failure modes.
This is not a notebook-only prompt-only or research-only role. A successful candidate should be prepared to discuss specific systems they have built including the data flow model or inference architecture deployment approach evaluation strategy failure modes and what they personally implemented.
What This Role Will Work On
- Design and build AI/ML capabilities for Intelligent Document Processing including OCR post-processing document parsing NLP/LLM extraction semantic search retrieval evidence grounding and structured data generation.
- Develop production-quality Python services pipelines and tooling that turn messy source documents into reliable traceable usable information.
- Work across the full lifecycle of AI systems: data ingestion preprocessing model or LLM integration evaluation deployment monitoring and iterative improvement.
- Build and improve systems that process PDFs scanned documents tables forms drawings images technical manuals and other complex document types.
- Collaborate with software engineers data engineers cloud engineers product leads customers and leadership to turn ambiguous technical problems into working solutions.
- Make practical engineering decisions about when to use deterministic logic classical NLP OCR embeddings LLMs fine-tuned models or human review workflows.
- Helpestablishengineering standards for evaluation reproducibility model behavior data quality traceability and responsible use of AI in customer-facingsystems.
Responsibilities
- Design implement andmaintainML/AI software components for IDP and Generative AI systems.
- Build data pipelines for unstructured and semi-structured data including document ingestion extraction cleaning enrichment validation and storage.
- Develop and evaluate NLP OCR computer vision embedding retrieval and LLM-based approaches for document understanding use cases.
- Create APIs internal tools review interfaces dashboards orvalidationworkflows that allow engineers and users to inspect correct and trust system output.
- Contribute production-quality code with clear structure tests logging error handling and documentation.
- Deploy and support ML/AI services in cloud or containerized environments including model serving batch processing and workflow automation.
- Design evaluation approaches for extraction quality retrieval quality model behavior hallucination risk and end-to-end system performance.
- Troubleshoot system behavior across model output data quality retrieval schema design infrastructure latency cost and user workflow issues.
- Mentor other engineers and help raise the technical quality of the team.
- Communicate clearly with both technical and non-technical stakeholders including project managers customers and executive leadership.
Required Skills and Experience
- Bachelors degree in Computer ScienceEngineering or a related technical fieldfrom an ABET accredited university.
- 5 years of professional hands-on experience in machine learning engineering AI engineering data science engineering or a closely related software engineering role.Plusanadditional5 years of experience in technology and software engineering.
- Demonstrated experience building AI/ML systems beyond notebooks prototypes or demos. Candidates should have shipped or supported pipelines services APIs inference endpoints evaluation harnesses or production-facing tools.
- Strong Python engineering experience including readable maintainable code; debugging; testing; packaging; and integration with other systems.
- Hands-on experience with NLP OCR computer vision LLMs embeddings semantic search RAG or other document-understanding techniques.
- Experience working with unstructured or semi-structured data such as PDFs scanned documents forms tables images logs contracts technical manuals or engineering documentation.
- Ability to design and reasonaboutend-to-end data flow: source data preprocessing transformation model/inference step persistence API/service boundary evaluation and user-facing output.
- Familiarity with common ML frameworks and tooling such asPyTorch TensorFlow scikit-learn Hugging FaceMLflow or similar technologies.
- Experience with databases and data stores including SQL and at least one relational or non-relational data platform.
- Experience using Git-based development workflows code review issue tracking and team-based software delivery practices.
- Clear written and verbal communication skills including the ability to explain technical tradeoffs limitations and failure modes.
- Ability to work effectively remotely in cross-functional teams.
- Ability to meet deadlines and produce quality work.
- Proficient in Microsoft Suite software including Outlook Word Excel SharePoint and PowerPoint.
Preferred Skills and Experience
- Direct experience with Intelligent Document Processing document AI OCR pipelines table extraction form extraction layout-aware processing or evidence-grounded retrieval.
- Experience building deploying or operating LLM-backed systems including inference serving prompt/version management model evaluation retrieval observability or cost/latency management.
- Experience with cloud platforms such as AWS or Azure including storagecompute serverless networking basics IAM monitoring or managed ML/AI services.
- Experience with containers and deployment workflows including Docker Kubernetes CI/CD pipelines automated tests and environment promotion.
- Experience building user-facing or internal tools such as validation interfaces review workflows dashboards admin tools or lightweight full-stack applications.
- Experience with data engineering tools such as pandas NumPy Spark/PySpark Databricks Airflow or similar workflow/data platforms.
- Experience with observability performance profiling or debugging tools such as Grafana CloudWatchTensorBoard tracing tools GPU profiling tools or application logs.
- Experience with evaluation design benchmarking reproducibility statistical analysis error analysis or human-in-the-loop validation.
- Bachelors or advanced degree in Computer Science Engineering Mathematics Physics Statistics Data Science oranothertechnical discipline. Equivalent professional experience will also be considered.
- Demonstrated ability to mentor junior developers or contribute to team technical direction.
Clearance Requirements
Applicants selected will be subject to a security investigation and may need to meet eligibility requirements. Secret Clearance is required for continued employment.
Work Location: Reston VA - Remote
- Our company prioritizes the benefits of flexibility and collaboration whether that happens in person or remotely.
- If the position is remote or hybrid you may periodically work from a Pantheon Data office location or client site.
- If this position is assigned to a Pantheon Data office location or client site youll work with colleagues and clients in person as needed for specific client requirements.
Interview Requirement: Candidates who are local to the area should be prepared to participate in an in-person interview as part of the selection process. Candidates outside the local area may be considered for a virtual interview.
Compensation
The salary range for this position is $140000 - $200000. This is not however a guarantee of compensation or salary. Rather salary will be set based on experience geographic location and possibly contractual requirements and could fall outside of this range.
Benefits Overview
We are always looking for good people! Pantheon Data is committed to providing its employees with competitive salaries and benefits in order to increase employee satisfaction and addition to our benefits we also offer SmartBenefits through the Washington Metro Area Transportation Authority where you specify an amount of your pre-tax wages be paid directly to your SmarTrip some cases tuition assistance may be available for continuing education expenses and certifications related to their position. Additional details may be found at Data Important Information
All qualified applicants will be considered for employment without regard to disability status as a protected veteran or any other status protected by applicable federal state local or international law.
As part of the application process you are expected to be on camera during interviews and assessments. We reserve the right to take your picture to verify your identity and prevent fraud.
If you require reasonable accommodation in completing this application interviewing completing any pre-employment testing or otherwise participating in the employee selection process please direct your inquiries to our Talent Team at or by phone .
This company uses E-Verify to confirm each employees work authorization. For more information click here E-Verify Participation Poster
Required Experience:
Senior IC