Senior SRE Engineer
Job Location:
Atlanta, GA - USA
Monthly Salary:
Not provided by the employer
Posted:
29 September 2026 (18 hours ago)
Application Deadline:
27 December 2026
Vacancies:
1 Vacancy
Job Summary
At Staffworxs we dont just connect talent we power transformation. Headquartered in Frisco TX with teams in Bengaluru and Hyderabad we combine global reach with deep expertise. Our Digital & Data Analytics practice drives growth and innovation for some of the worlds top brands who continue to retain us as their trusted partner. If youre ready to make an impact youre in the right place.
Job Details:
Job Title: Senior SRE Engineer
Work Location: Atlanta GA (Hybrid)
Duration: 12 months Contract
Work Mode for delta roles:
50% Hybrid Work From Office Schedule
The role follows a 50% hybrid work model. Employees are required to work from the office for 5 consecutive business days from Wednesday through the following Tuesday on alternate weeks.
Job Description:
Enterprise Site Reliability Engineer
Position Summary
We are seeking an experienced Enterprise Site Reliability Engineer to advance reliability resilience observability automation and operational excellence across the organization.
This is a highly visible hands-on engineering role that will work across SRE application architecture cloud platform infrastructure security data and other IT domains. The successful candidate will combine deep technical expertise with curiosity thorough analysis and strong engineering judgment to identify risks others may overlook connect technical findings to business impact and develop scalable enterprise solutions.
The role requires someone who can move effectively between detailed technical analysis software development application architecture reviews complex incident leadership enterprise standards and executive-level communication. The candidate must be able to influence SRE and engineering practices across organizational boundaries without relying on direct authority.
Key Responsibilities
Staffworxs is an Equal Opportunity Employer. We celebrate diversity and are committed to creating an inclusive workplace for all employees regardless of race color religion gender sexual orientation national origin age disability or veteran status.
Job Details:
Job Title: Senior SRE Engineer
Work Location: Atlanta GA (Hybrid)
Duration: 12 months Contract
Work Mode for delta roles:
50% Hybrid Work From Office Schedule
The role follows a 50% hybrid work model. Employees are required to work from the office for 5 consecutive business days from Wednesday through the following Tuesday on alternate weeks.
Job Description:
Enterprise Site Reliability Engineer
Position Summary
We are seeking an experienced Enterprise Site Reliability Engineer to advance reliability resilience observability automation and operational excellence across the organization.
This is a highly visible hands-on engineering role that will work across SRE application architecture cloud platform infrastructure security data and other IT domains. The successful candidate will combine deep technical expertise with curiosity thorough analysis and strong engineering judgment to identify risks others may overlook connect technical findings to business impact and develop scalable enterprise solutions.
The role requires someone who can move effectively between detailed technical analysis software development application architecture reviews complex incident leadership enterprise standards and executive-level communication. The candidate must be able to influence SRE and engineering practices across organizational boundaries without relying on direct authority.
Key Responsibilities
- Define and advance enterprise SRE standards for SLIs SLOs error budgets production readiness incident management and operational excellence.
- Establish practical observability standards for OpenTelemetry traces logs metrics events telemetry correlation tagging data quality and service ownership.
- Conduct thorough application maturity assessments across observability reliability resilience operability automation and incident readiness.
- Analyze application architecture and production operations to uncover dependencies failure modes capacity constraints performance bottlenecks and risks that may not be immediately visible.
- Connect technical and operational findings to customer experience business impact service risk and investment priorities.
- Translate assessments and operational data into clear prioritized and measurable improvement roadmaps.
- Influence architecture and engineering decisions by recommending reliability patterns such as fault isolation graceful degradation circuit breakers retries rate limiting load shedding high availability and disaster recovery.
- Design end-to-end observability across distributed services APIs business transactions customer journeys cloud platforms and cross-domain dependencies.
- Develop production-grade automation applications APIs and integrations that reduce toil improve consistency and scale reliability practices across the enterprise.
- Automate service onboarding telemetry validation SLO reporting production readiness checks incident enrichment and remediation workflows.
- Integrate observability cloud CI/CD IT service management incident management and configuration platforms using APIs SDKs webhooks and event-driven patterns.
- Lead complex incident investigations and post-incident reviews challenge assumptions identify contributing factors and drive corrective actions through completion.
- Analyze operational and observability data to detect patterns quantify risks identify systemic gaps and generate actionable insights for engineers and leaders.
- Proactively identify opportunities to improve application operability resilience telemetry support readiness and operational processes before issues affect customers.
- Define observability for AI agents and AI-enabled applications including workflows model and tool interactions dependencies latency failures quality token consumption cost and reliability.
- Use AI-assisted engineering tools such as Kiro GitHub Copilot and AI agents to accelerate development incident analysis correlation pattern detection and operational insights.
- Develop reusable reference architectures engineering patterns assessment frameworks maturity models scorecards and implementation guidance.
- Work with SREs and engineering teams across the organization to resolve cross-domain reliability issues and promote consistent practices.
- Mentor engineers facilitate difficult technical decisions constructively challenge existing approaches and build alignment across teams.
- Communicate complex technical risks business impact recommendations and progress clearly to engineering teams and senior leadership.
- 8 years of experience in Site Reliability Engineering Production Engineering Platform Engineering or software engineering for large-scale production systems.
- Deep knowledge of SRE principles operational excellence distributed systems microservices APIs and cloud-native architecture.
- Deep hands-on experience architecting operating and troubleshooting large-scale AWS environments across compute containers serverless networking databases storage identity and cloud observability including EC2 EKS Lambda VPC Elastic Load Balancing Route 53 RDS/Aurora DynamoDB S3 IAM and CloudWatch.
- Strong hands-on experience with Kubernetes and Red Hat OpenShift Service on AWS (ROSA) architecture operations performance and troubleshooting.
- Strong experience with OpenTelemetry distributed tracing logs metrics events and Dynatrace or a comparable enterprise observability platform.
- Experience defining and operationalizing SLIs SLOs error budgets production readiness criteria and reliability scorecards.
- Demonstrated ability to assess application architecture and operational maturity from observability reliability resilience and operability perspectives.
- Strong software development skills in Python Java Go JavaScript/TypeScript or similar languages.
- Experience developing production-grade APIs integrations automation services and internal engineering tools.
- Proficiency with REST APIs SDKs Git automated testing CI/CD Infrastructure as Code secure development and software lifecycle practices.
- Experience leading complex incidents technical investigations root cause analysis post-incident reviews and corrective-action programs.
- Strong analytical and investigative skills with the curiosity and technical judgment to look beyond immediate symptoms and uncover systemic problems.
- Demonstrated ability to convert technical findings and operational data into clear actionable insights and enterprise recommendations.
- Ability to connect reliability risks and technical decisions to customer experience business impact and organizational priorities.
- Excellent written verbal technical and executive communication skills.
- Proven ability to mentor experienced engineers challenge assumptions constructively facilitate decisions and influence outcomes without direct authority.
- Ability to collaborate effectively with SREs architects application teams and specialists across multiple IT domains.
- AWS certification preferably AWS Certified Solutions Architect Professional AWS Certified DevOps Engineer Professional or a relevant specialty certification.
- Kubernetes or OpenShift certification such as CKA CKS or Red Hat Certified OpenShift Administrator.
- Experience integrating enterprise platforms such as Dynatrace AWS ServiceNow GitLab Kubernetes and ROSA.
- Experience developing self-service reliability capabilities internal developer platforms or enterprise engineering products.
- Experience with performance engineering capacity planning resilience testing and chaos engineering.
- Experience establishing observability and reliability controls for AI agents LLM-enabled applications or AI-driven workflows.
- Experience working with large-scale highly available business-critical enterprise systems.
Staffworxs is an Equal Opportunity Employer. We celebrate diversity and are committed to creating an inclusive workplace for all employees regardless of race color religion gender sexual orientation national origin age disability or veteran status.
Required Experience:
Senior IC