Enter a job title or keyword

OpenShift & GitLab L2


Job Location:

Mumbai - India

Monthly Salary: Not provided by the employer
Posted: 11 September 2026 (14 hours ago)
Application Deadline: 9 December 2026
Vacancies: 1 Vacancy

Job Summary

Monitor OpenShift cluster health including control plane components API responsiveness kubelet node status cluster operators etcd health CPU memory disk network pod restarts and overall platform availability using enterprise monitoring tools.
Perform daily/shift health checks for nodes pods cluster operators disk pressure certificate validity scheduled jobs and overall platform status.
Monitor platform alerts acknowledge incidents within defined SLAs execute approved L1 runbooks and coordinate timely incident response.
Troubleshoot and resolve control plane issues including API instability etcd quorum/latency node NotReady conditions operator failures CNI/OVN-Kubernetes networking issues Ingress/Route failures admission webhook problems and image registry outages.
Collect analyze and provide diagnostics using oc/kubectl pod descriptions logs events manifests must-gather oc adm inspect audit logs API server metrics etcd metrics node logs and journalctl.
Perform first-line remediation by restarting pods/services deleting failed pods scaling workloads clearing PVC locks rotating authorized tokens and executing documented recovery procedures.
Escalate unresolved or complex issues to L2/L3 with complete diagnostics perform Root Cause Analysis (RCA) workload impact assessment and recommend corrective and preventive actions.
Monitor and administer Machine API Autoscaler node lifecycle operations including drain cordon uncordon reimage node replacement and cluster resource optimization.
Plan coordinate and execute OpenShift cluster upgrades z-stream patching Operator (OLM) upgrades cluster add-on maintenance compatibility validation canary deployments rollback planning and post-upgrade verification.
Support scheduled maintenance by performing pre/post patch validation node drain/uncordon controlled reboots cluster health validation and maintenance documentation.
Administer and troubleshoot OpenShift networking including OVN-Kubernetes/CNI NetworkPolicies Egress/IPs ExternalIPs Routes Ingress Controllers HAProxy/load balancing DNS Services EndpointSlices and network connectivity.
Administer storage services including CSI drivers StorageClasses PVC/PV lifecycle reclaim policies storage performance registry storage ODF/ODF-LVM/Ceph platforms and persistent storage troubleshooting.
Monitor and manage the internal image registry registry storage image pull policies registry trust image cleanup garbage collection artifact repositories replication retention and geo-distribution.
Validate cluster backups etcd snapshots backup reports restore procedures Velero integration disaster recovery readiness DR runbooks and recovery testing.
Execute approved routine platform administration tasks including ConfigMap and Secret updates Route/Ingress management namespace/project creation service account support and simple configuration changes.
Administer RBAC enforce least-privilege access manage quotas LimitRanges Pod Disruption Budgets (PDBs) priority classes eviction policies and multi-tenancy best practices.
Manage TLS certificates including API Ingress internal trust chains certificate rotation renewal expiry validation and certificate lifecycle management.
Monitor administer and optimize CI/CD platforms including Jenkins GitLab CI GitLab Runners Tekton Argo Workflows controllers agents runners job queues artifact retention caching and execution capacity.
Troubleshoot CI/CD pipeline failures involving credentials registry access network connectivity runner capacity image pull failures flaky tests deployment failures pipeline parameters and build logs.
Develop maintain and optimize reusable CI/CD pipeline templates shared libraries GitOps repository structures Infrastructure-as-Code automation using Terraform/Ansible approval workflows and policy validation.
Support developers by troubleshooting application deployments image pull issues liveness/readiness probe failures resource limits deployment errors namespaces service account tokens and cluster access issues.
Operate and optimize the monitoring and logging platforms including Prometheus Alertmanager Grafana Loki/EFK log forwarding indexing retention scrape targets dashboards recording rules SLO/SLA reporting and alert optimization.
Enforce platform security by administering SCC/Pod Security Admission (PSA) RBAC reviews image security registry trust service mesh security Vault/KMS integration secrets management dynamic credentials key rotation authentication monitoring and compliance controls.
Implement release strategies including blue-green deployments canary deployments feature flags rollout policies automated rollback and deployment validation.
Optimize platform performance through resource requests/limits tuning HPA/VPA optimization node right-sizing workload bin-packing capacity forecasting and compute storage and network cost optimization.
Maintain operational documentation update patch backup maintenance and incident records ensure compliance with organizational standards and support change management processes.
Coordinate with L2/L3 teams during incidents maintenance windows platform upgrades disaster recovery exercises and production support while ensuring adherence to operational procedures SLAs and security policies.
GitLab Administration & Monitoring: Administer configure monitor secure troubleshoot upgrade patch and maintain the complete GitLab platform including GitLab Server GitLab CI/CD GitLab Runners Gitaly PostgreSQL Redis Praefect (where applicable) Container Registry Package Registry LDAP/Active Directory integration SSL/TLS certificates backups and disaster recovery repository administration user and group management RBAC authentication performance tuning high availability monitoring capacity planning log analysis security hardening integrations with external systems and overall platform health availability and lifecycle management.

Required Experience:

IC


About Company

3i Infotech Limited is a global Information Technology company committed to Empowering Business Transformation.A comprehensive set of IP based software solutions (20+), coupled with a wide range of IT services, uniquely positions the company to address the dynamic requirements of a va ... View more

View Profile View Profile