Lead Infrastructure & DevOps
Job Summary
1. Infrastructure & OS Management
- Manage and maintain Linux-based systems (Rocky Linux CentOS RHEL) in production and staging environments.
- Perform kernel patching security updates and system hardening following industry best practices.
- Troubleshoot and debug complex OS-level issues (performance memory I/O kernel panics).
- Manage and support Windows Server environments where applicable.
2. Security & Compliance
- Perform regular security scanning and vulnerability assessments using tools like Nessus Qualys or OpenVAS.
- Conduct code security scans (SAST/DAST) and integrate them into CI/CD pipelines.
- Configure and maintain LDAP authentication for centralized access management.
- Ensure all systems comply with internal security policies and external regulatory standards.
- Implement and manage WAF (Web Application Firewall) configurations to protect against OWASP Top 10 threats.
3. Automation & Configuration Management
- Write and maintain Ansible playbooks for automated provisioning patching and configuration management.
- Manage GitHub Enterprise repositories branching strategies and CI/CD workflows (GitHub Actions preferred).
- Develop automation scripts using Python Shell or Go to reduce toil and improve operational efficiency.
4. Application & Web Server Management
- Configure and maintain Apache and Nginx web servers including performance tuning and reverse proxy setup.
- Troubleshoot web server logs SSL/TLS issues and load balancing configurations.
5. Container Orchestration & Cloud-Native Technologies
- Deploy manage and scale applications on Kubernetes clusters (EKS AKS or on-prem).
- Build and maintain Docker images with secure base layers and minimal attack surfaces.
- Monitor cluster health resource usage and implement auto-scaling policies.
6. Monitoring & Incident Management
- Deploy and maintain monitoring tools (Prometheus Grafana Datadog or ELK stack).
- Define SLIs SLOs and error budgets; drive blameless post-mortems.
- Participate in on-call rotations and lead incident response efforts.
- 8 years in SRE DevOps or Systems Engineering roles.
- Experience with Terraform or other IaC tools.
- Familiarity with service mesh (Istio/Linkerd) and CNI plugins (Calico/Cilium) or similar.
- Understanding of zero-trust networking and SSO integrations.
- Certifications: RHCE CKA or CISSP are a plus.
OS: Expert in Rocky Linux CentOS RHEL (6 years); Windows Server experience
Security: WAF vulnerability scanning LDAP OS hardening kernel patching
Automation: Ansible (playbooks roles AWX/Tower) Python/Shell scripting
CI/CD/SCM: GitHub Enterprise administration GitHub Actions branch protection rules
Containers: Docker (build/security scanning) Kubernetes (deployment networking storage)
Web Servers: Apache Nginx (configuration SSL reverse proxy)
Monitoring: Prometheus Grafana ELK or commercial tools (Datadog/New Relic)
Code Security: SAST/DAST tools (SonarQube Snyk Checkmarx)
- Strong communication and documentation skills.
- Ability to mentor junior engineers and conduct technical interviews.
- Collaborative mindset with DevOps/DevSecOps philosophy.
- Calm under pressure during incidents; drive for continuous improvement.
XTIUM is an equal opportunity employer.