Sr. Staff AI Architect
Job Summary
Work Schedule
Standard (Mon-Fri)Environmental Conditions
OfficeJob Description
About Company:
ThermoFisher Scientific Inc. is the world leader in serving science with revenues of more than $25 billion and approximately 130000 employees globally. We help our customers accelerate life sciences research solve elite analytical challenges improve patient diagnostics deliver medicines to market and increase laboratory efficiency.
Through our world-class brandsThermo Scientific Applied Biosystems Invitrogen Fisher Scientific and Unity Lab Serviceswe offer an unmatched combination of innovative technologies purchasing convenience and comprehensive services.
About Team:
We are the Digital Foundation Platform team - the software center of excellence (CoE) for Thermo Fisher Scientific. We are responsible for developing and delivering foundational components and SaaS-based applications and digital lab (Cloud-based) tools foundational AI and AI agentic solutions to help scientists do their work more efficiently and with precision enabling them to make our world healthier cleaner and safer. Our elite software products and solutions accelerate scientific discovery and lab productivity. These solutions
- Provide rich content selection tools teamwork tools and scientific apps that allow our customers to focus on innovation and the complexities of their science.
- Build a connected world for our customers where discoveries happen in a thoughtful way where every device/product is connected self-aware and self-healingthereby enabling efficient workflows and collaborative science.
- Enable our customers to efficiently handle their labs by providing them with insight into workflow processes asset uptime and product availability.
We give them the flexibility to access what they need when they need it allowing them to select and receive products and services across multiple channels. We apply industry-standard methodologies to the design development and deployment of best-in-class software products built to demonstrate the power and scalability of the cloud.
Purpose:
The Senior Staff Architect provides enterprise-wide architectural design and technical leadership for Generative AI and agentic AI solutions across multiple Scrum teams. As a strategic technical leader and hands-on architect you will define the architecture standards and roadmap for an enterprise agentic harness that enables agents to reason plan use tools access knowledge maintain state collaborate and operate safely at scale.
The role requires deep expertise in designing production-grade agentic systems cloud-native platforms RAG architectures LLM integrations and AI engineering practices. You will influence platform strategy mentor architects and engineers and ensure AI-driven systems are scalable secure resilient observable governable and production-ready.
- Provide enterprise architecture leadership for Generative AI and agentic AI platforms across multiple teams.
- Define the architecture and roadmap for a reusable agentic harness and runtime platform.
- Own high-level and low-level system design including component architecture data flows integration patterns deployment topologies and runtime interactions.
- Design and evolve cloud-native event-driven API-first and distributed architectures for AI-enabled products.
- Define reference architectures design standards and best practices for:
- Agentic systems
- Multi-agent orchestration
- RAG and knowledge-grounded agents
- Tool and function execution
- Agent memory and state management
- Human-in-the-loop workflows
- Agent-to-Agent communication
- Define architecture patterns for agent planning task decomposition reasoning execution retries timeouts compensation and failure recovery.
- Establish standards for agent lifecycle management including agent creation versioning testing evaluation deployment monitoring and retirement.
- Act as the go-to authority for architecture design trade-offs scalability decisions and complex implementation challenges.
- Ensure architectural decisions address Non-Functional Requirements (NFRs) including performance scalability security reliability privacy governance cost and observability.
- Lead architecture reviews design reviews technical investigations and architecture decision records.
- Design and implement agentic systems using LangGraph LangChain and comparable orchestration frameworks.
- Architect the core capabilities of an agentic harness including:
- Agent runtime and execution management
- Planning and task orchestration
- Tool registry and tool discovery
- Function calling and API integration
- Agent memory and context management
- Workflow state persistence
- Multi-agent collaboration
- Human approvals and intervention
- Guardrails and policy enforcement
- Execution tracing auditability and replay
- Design model-agnostic architectures supporting Azure OpenAI Anthropic Claude OpenAI-compatible APIs and local models using Ollama.
- Define model routing fallback provider abstraction latency optimization token management and cost-control strategies.
- Architect and implement Retrieval-Augmented Generation (RAG) solutions including document ingestion chunking strategies embeddings retrieval reranking grounding and response synthesis.
- Integrate agents with APIs enterprise systems scientific applications structured data knowledge graphs and domain ontologies.
- Apply advanced prompt engineering structured outputs tool-calling memory patterns and context optimization techniques.
- Define safe and reliable patterns for autonomous and semi-autonomous agent execution.
- Define automated testing and evaluation strategies for agentic and GenAI systems.
- Establish evaluation frameworks for:
- Task completion
- Tool-use accuracy
- Response correctness
- Grounding and retrieval quality
- Safety and policy compliance
- Robustness and recovery
- Latency and cost
- Design regression testing for prompts models tools workflows and agent behaviors.
- Define observability standards using execution traces agent steps tool calls model interactions latency token usage cost and failure metrics.
- Establish mechanisms for agent debugging replay inspection and root-cause analysis.
- Define secure deployment monitoring governance and operational practices for AI systems.
- Ensure production systems support reliability scalability disaster recovery data protection and compliance requirements.
- Actively contribute to hands-on development using Python and modern backend frameworks such as FastAPI.
- Build reference implementations and reusable platform components for the agentic harness.
- Design and build well-structured maintainable and extensible APIs supporting AI agent and data-driven workloads.
- Define service boundaries integration contracts data models workflow interfaces and event schemas.
- Implement performance scalability security and observability patterns for distributed AI services.
- Guide the use of PostgreSQL pgvector Qdrant event streaming caching and persistent workflow storage.
- Promote clean code automated testing code reviews CI/CD and engineering quality standards.
- Mentor and guide architects staff engineers and senior engineers on:
- Agentic architecture
- Distributed systems
- System design
- GenAI engineering
- Evaluation and observability
- Influence technical direction across multiple Scrum teams and product areas.
- Lead Communities of Practice focused on agentic AI GenAI architecture and AI engineering excellence.
- Communicate effectively with technical and non-technical stakeholders through clear documentation architecture diagrams design reviews and technical presentations.
- Partner with product security platform data and engineering leadership to align agentic AI capabilities with business and scientific objectives.
- Anticipate architectural risks and opportunities and guide the organization through technology evolution.
Bachelors degree in Engineering or Masters degree in Computer Science with 15 years of proven industry experience including significant experience in software architecture platform engineering and technical leadership.
- Enterprise Architecture: Proven experience defining architecture and technical strategy across multiple products or engineering teams.
- Python Backend Development: 8 years of experience building scalable backend systems and RESTful APIs using Python and FastAPI.
- Distributed Systems: Strong knowledge of microservices event-driven architecture asynchronous processing messaging caching and workflow orchestration.
- API Design & Integration: Strong experience with API lifecycle management authentication authorization service-to-service communication and integration patterns.
- Cloud-Native Architecture: Experience designing scalable and resilient solutions on Azure AWS or GCP.
- Version Control & CI/CD: Proficiency with Git automated delivery pipelines infrastructure automation and release management.
- Testing & Automation: Experience with pytest unittest integration testing contract testing and automated quality strategies.
- Technical Leadership: Demonstrated ability to influence architecture and engineering decisions across multiple teams without relying solely on organizational authority.
- Deep hands-on experience designing and implementing agentic AI systems and agentic harnesses.
- Strong experience with LangGraph LangChain or comparable agent orchestration frameworks.
- Practical experience with:
- Agent planning and task decomposition
- Tool and function calling
- Multi-agent coordination
- Agent memory and state management
- Human-in-the-loop workflows
- Guardrails and policy enforcement
- Retry recovery and failure handling
- Agent execution tracing and replay
- Experience integrating LLMs using Azure OpenAI Anthropic Claude OpenAI-compatible APIs and local inference platforms such as Ollama.
- Strong prompt engineering skills including structured prompting few-shot learning prompt versioning context management and output validation.
- Experience implementing RAG architectures embeddings semantic search hybrid search reranking grounding and citation patterns.
- Hands-on experience with vector databases such as PostgreSQL with pgvector and Qdrant.
- Experience with model routing fallback strategies token optimization latency management and LLM cost control.
- Experience evaluating nondeterministic AI and agentic systems in production-like environments.
- Knowledge of evaluation-driven development golden datasets regression testing and quality metrics for agents.
- Experience implementing observability for prompts model calls tool calls agent steps workflow state latency token usage and cost.
- Understanding of AI security data privacy access control responsible AI auditability and governance.
- Experience with LLMOps model gateways prompt management model evaluation monitoring and deployment governance.
- Design and manage data stores and vector indexes supporting GenAI and RAG workloads using PostgreSQL pgvector and Qdrant.
- Strong data engineering skills including ETL pipelines data ingestion preprocessing metadata management and large-scale data processing.
- Experience integrating structured data knowledge graphs and domain ontologies with agentic systems.
- Expertise in Pandas NumPy and Python data-processing libraries.
- 5 years of experience working in Agile/Scrum or comparable product development environments.
- Excellent written and verbal communication skills.
- Ability to explain complex architecture and AI concepts clearly to technical and non-technical stakeholders.
- Demonstrated experience mentoring senior engineers architects and technical leads.
- Cloud Platforms: Azure preferred including AKS managed identity Azure AI services eventing security and monitoring.
- Agent Protocols: Experience with MCP Agent-to-Agent communication protocols or comparable tool and agent interoperability standards.
- Workflow Platforms: Experience with durable workflow engines or distributed orchestration platforms.
- LLMOps / MLOps: Familiarity with MLflow Kubeflow model gateways prompt registries evaluation platforms and AI governance.
- Machine Learning Frameworks: Exposure to scikit-learn PyTorch or TensorFlow.
- Observability: Experience with OpenTelemetry distributed tracing logging monitoring and AI-specific telemetry.
- Code Quality: Experience with SonarQube static analysis secure coding and enterprise engineering standards.
- Scientific or Regulated Domains: Experience supporting scientific workflows healthcare laboratory or regulated environments.
- A reusable secure and scalable enterprise agentic harness adopted by multiple product teams.
- Production-grade agents capable of reliable planning tool use knowledge retrieval collaboration and recovery.
- Standardized agent development evaluation deployment monitoring and governance practices.
- Measurable improvements in workflow automation developer productivity scientific outcomes product capability and customer experience.
- Engineering teams enabled through strong architectural guidance reference implementations mentorship and technical standards.
- AI systems that are observable explainable cost-aware resilient and maintainable in production.
Thermo Fisher Scientific is an equal opportunity employer and is committed to building a diverse and inclusive workforce.
Required Experience:
Staff IC
About Company
Electron microscopes reveal hidden wonders that are smaller than the human eye can see. They fire electrons and create images, magnifying micrometer and nanometer structures by up to ten million times, providing a spectacular level of detail, even allowing researchers to view single a ... View more