AWS DevOps Agentic SRE Engineer (Active TSSCI Clearance)
Chantilly, VA - USA
Job Summary
Clearance: Active TS/SCI
Certification: Active Security or equivalent required
Experience: 3 years of SRE DevOps Platform Engineering or Infrastructure Engineering experience
We are seeking an SRE / DevOps & Release Engineer to own the deployment reliability observability and operation of secure AWS environments supporting an advanced Agentic AI platform.
This is a hands-on engineering position combining DevOps Site Reliability Engineering platform engineering and release automation. The engineer will own the path from local development through production deployment into AWS Kubernetes environments as well as the operational health and reliability of those environments.
The platform is currently operating within IATT environments and progressing toward scale and ATO. At the current team size build/release and production operations are intentionally combined the engineer responsible for deploying the platform also has responsibility for ensuring that it operates reliably.
The environment includes AWS EKS CDK Kubernetes Helm Flux GitOps container registries Envoy Gateway distributed tracing and modern observability technologies.
- Design build maintain and operate highly available AWS infrastructure supporting an Agentic AI platform.
- Own the software delivery lifecycle from local development through build test packaging promotion and deployment into AWS environments.
- Deploy and operate Kubernetes workloads in production AWS EKS environments.
- Author maintain and troubleshoot Helm charts supporting platform applications and services.
- Build and maintain GitOps-based continuous delivery utilizing Flux Helm and container registries.
- Develop and maintain Infrastructure as Code using AWS CDK and TypeScript.
- Build and manage AWS infrastructure utilizing EKS RDS S3 IAM/IRSA ECR and related AWS services.
- Build and maintain container-image and Helm-chart promotion processes.
- Configure and support Envoy Gateway routing certificates and TLS.
- Implement and operate observability and distributed-tracing capabilities utilizing OpenTelemetry Grafana Tempo or equivalent technologies.
- Establish cross-service tracing to provide end-to-end visibility into distributed application and agent workflows.
- Monitor cluster application and service health and proactively identify reliability issues.
- Diagnose failed or stuck Helm and Flux reconciliations configuration drift pod failures and deployment issues.
- Lead production incident investigation and root-cause analysis using Kubernetes state logs metrics and distributed traces.
- Develop preflight validation diagnostic QA and operational tooling.
- Automate infrastructure and operational processes using Python and other scripting technologies.
- Support credential certificate and secrets rotation.
- Partner with software engineers and AI/ML teams to deploy and operate agentic applications model-serving infrastructure and supporting services.
- Support platform security compliance IATT and ATO activities.
- Use AI-assisted development tools to accelerate engineering while independently validating generated code and configuration before deployment.
- Drive an engineering approach based on measurable evidence automated testing telemetry and verification.
- Current and active TS/SCI security clearance
- Current Security certification or equivalent certification supporting privileged-user access.
- 3 years of professional experience in Site Reliability Engineering DevOps Platform Engineering Cloud Engineering or Infrastructure Engineering.
- Hands-on experience operating Kubernetes in production environments.
- Experience authoring and maintaining Helm charts.
- Experience implementing GitOps-based continuous delivery using Flux Argo CD or equivalent technologies.
- Hands-on AWS experience with services such as EKS RDS S3 IAM/IRSA and ECR.
- Experience developing and maintaining Infrastructure as Code.
- Experience working with TypeScript/AWS CDK or demonstrated ability to work within a TypeScript-based IaC environment.
- Experience implementing and utilizing production observability and distributed-tracing technologies such as OpenTelemetry Grafana and Tempo.
- Demonstrated experience diagnosing infrastructure and application failures using telemetry logs metrics and traces.
- Experience leading incident investigation root-cause analysis remediation and validation.
- Proficiency with Python or another scripting language for infrastructure automation and operational tooling.
- Strong understanding of CI/CD containers networking security and modern cloud architecture.
- TS/SCI with Poly
- Experience implementing registry-based GitOps architectures utilizing ECR or similar container registries.
- Experience troubleshooting Flux and Helm reconciliation issues including HelmRelease failures and configuration drift.
- Experience deploying and operating ML/LLM-serving infrastructure such as KServe MLflow or model-inference endpoints.
- Experience supporting AI or Agentic AI platforms.
- Familiarity with Amazon Bedrock LLM services MCP tools and agent-based architectures.
- Ability to analyze model and agent performance characteristics such as latency token utilization and cost using distributed traces.
- Experience with Envoy API gateways routing certificate management and TLS.
- Experience supporting secure Federal or Intelligence Community AWS environments.
- Experience supporting systems through IATT and ATO processes.
- Experience using AI coding assistants while independently validating generated code and infrastructure changes before deployment.
- Bachelors degree in Computer Science Engineering Information Technology or a related discipline or equivalent practical experience.
SBS is a trusted AWS partner supporting both federal and commercial customers. Our teams focus on deliverybuilding and operating secure scalable solutions that matter.
Youll be part of a team that:
- Works directly with AWS on meaningful mission-driven programs
- Operates in high-impact national security environments
- Builds and operates modern AWS and Kubernetes platforms
- Works across cloud infrastructure DevOps SRE GitOps and AI
- Values engineers who can build automate troubleshoot and deliver
COMPENSATION & BENEFITS
SBS offers a comprehensive total-rewards package including a market competitive salary along with:
Comprehensive medical dental and vision coverage; HSA-eligible plan options available
401(k) retirement plan with company match (vesting schedule per Plan Document)
Paid Time Off federal holidays and floating holiday for personal observance
Annual professional development support for AWS certifications training and conferences
Employee referral program where applicable and documented by program policy
Life AD&D and short- / long-term disability insurance
Telework and flexible-schedule support where mission and contract permit
Mission-focused federal contractor supporting national-security customers
Target Salary Range: $00.This range reflects the anticipated compensation for this role. Actual salary will be based on a combination of factors including the positions scope and level of responsibility the candidates relevant experience education technical expertise skills and qualifications geographic location and applicable business or contractual requirements.
About SBS
Strategic Business Systems Inc. (SBS)is a national Information Technology services company headquartered in the Washington D.C. metropolitan area. SBS provides IT infrastructure design integration and operational services. Our expertise spans the full spectrum of infrastructure technologies including networking servers data storage disaster recovery cybersecurity and internet technologies.
Equal Employment Opportunity
SBS is an equal opportunity employer; all qualified applicants will receive consideration for employment without regard to age gender gender identity sex sexual orientation color race creed national origin religion marital status parental status citizenship status ancestry physical or mental disability genetic information veteran status military status or any other classification protected by federal state or local laws.
Accommodations
If you need an accommodation while seeking employment with SBS please email. Accommodations are made on a case-by-case basis.
No Unsolicited Agency Referrals
SBS does not accept unsolicited resumes from staffing agencies. Any resumes submitted without a prior agreement will be considered the property of SBS.
Required Experience:
Unclear Seniority
About Company
SBS is a Dubai-based software house that provides the most innovative healthcare solutions & medical software for all healthcare providers with different sized across the MENA region with operations in the USA, Saudi Arabia, Kuwait, Bahrain, Qatar, and Egypt. Our team has 30 years ... View more