Enter a job title or keyword

SRE Operations Lead

Diverse Lynx LLC


Job Location:

Charlotte, VT - USA

Yearly Salary: USD 90000 - 110000
Posted: 10 October 2026 (Yesterday)
Application Deadline: 7 January 2027
Vacancies: 1 Vacancy

Job Summary

Job Role: SRE Operations Lead
Location: Charlotte NC (onsite)
Job Type: Full-Time
Salary: $90000 - $110000 a year

Job Description
We are looking for an SRE Operations Lead to drive reliable scalable and automated production operations across AWS and GitLab environments. The role will focus on SRE operations incident management observability DevSecOps cloud automation AIOps and AI-assisted troubleshooting.
The ideal candidate will have strong experience with AWS infrastructure CI/CD GitLab Python/Shell scripting observability platforms incident response and reliability engineering along with the ability to leverage AI tools such as Claude and AWS Bedrock for troubleshooting and automation.
Must Have Technical/Functional Skills
Strong experience in Site Reliability Engineering (SRE) Application Reliability Engineering (ARE).
Strong knowledge of CI/CD DevSecOps production operations incident management and reliability engineering.
Hands-on experience with AWS services including:
CloudWatch
Route 53
S3
CloudFront
ECR
EC2
Lambda
AWS Bedrock
Experience with GitLab repositories CI/CD pipelines branching strategies and source-code management.
Experience using Claude/Claude Code or similar AI-assisted development tools for troubleshooting analysis and automation.
Strong Python and Shell/Bash scripting skills for operational automation and integration development.
Experience with Grafana Prometheus CloudWatch Splunk Dynatrace Datadog OpenTelemetry APM and distributed observability.
Strong understanding of SLIs SLOs SLAs MTTR reliability KPIs service health and operational resilience.
Experience with incident management and P1/P2 major incident response including root-cause analysis and service restoration.
Experience implementing automation AIOps event correlation self-healing and AI-driven operations.
Strong understanding of AWS containerized environments Docker Amazon ECR and scalable cloud architectures.
Experience with Terraform or similar Infrastructure as Code technologies.
Experience developing and consuming REST APIs and integrating cloud DevOps and source-control platforms.
Knowledge of GitLab APIs Personal Access Tokens (PATs) custom workflows and CI/CD automation.
Experience with security scanning vulnerability management remediation and DevSecOps controls.
Knowledge of AWS Bedrock and AI-driven operational automation.
Experience supporting resilient and multi-region AWS architectures.
Strong troubleshooting analytical communication and problem-solving skills.
Roles & Responsibilities
SRE operations and coordinate P1/P2 major incident response through service restoration root-cause analysis and follow-up actions.
Define and monitor SLIs SLOs SLAs MTTR availability reliability and operational performance KPIs.
Drive end-to-end observability across applications and infrastructure using Grafana Prometheus CloudWatch Splunk Dynatrace Datadog and OpenTelemetry.
Lead initiatives around AIOps event correlation self-healing automated remediation and AI-driven operations.
Manage and Eligible to workimize AWS platform operations including compute storage networking containers serverless workloads and deployment environments.
Support AWS Lambda ECR EC2 S3 CloudWatch CloudFront Route 53 and Bedrock environments.
Design and implement automation to improve operational efficiency reliability and incident response.
Integrate Claude/Claude Code with AWS Bedrock and GitLab using APIs PATs and custom automation workflows.
Develop AI-assisted solutions for analyzing GitLab projects vulnerabilities pipelines security findings and operational data.
Design and maintain GitLab API integrations custom workflows DevSecOps controls and CI/CD improvements.
Build and maintain Python-based GitLab integrations and REST API automation.
Support vulnerability remediation and the onboarding configuration and maintenance of security scanning tools.
Eligible to workimize AWS Lambda configurations including concurrency scaling performance and deployment automation.
Design and support resilient multi-region AWS architectures and containerized deployments using Docker and Amazon ECR.
Support batch-processing and scheduling environments including Control-M where applicable.
Partner with development security infrastructure and application teams to improve system reliability and operational maturity.
Identify recurring incidents and operational bottlenecks and implement permanent automation-based solutions.
Establish operational standards documentation runbooks monitoring strategies and continuous improvement practices.
Leverage AI and automation to reduce manual operational effort and improve troubleshooting incident resolution and service reliability.










Disclaimer: Diverse Lynx LLC is an Equal Opportunity Employer. All applicants and employees are evaluated without discrimination based solely on their qualifications ability competence and performance. This email and its attachments may contain confidential or proprietary information and is intended only for the recipient(s). If you received this message in error please disregard it and notify the sender. If you no longer wish to receive our communications you may unsubscribe at any time.
Security Notice: Our official website is We do not operate or authorize any other websites representing Diverse Lynx LLC.

Required Experience:

Manager