Lead Senior Production Support Operations Engineer
Job Summary
Job Title: Lead Senior Production Support / Operations Engineer
Education: Any Graduate
Experience: 8years
Job Location: Mumbai
Role Summary
We are seeking a highly Senior Production Support / Operations Engineer to represent and operationally support our in-house monitoring and observability platform across enterprise database environments.
This role acts as the critical operational bridge between:
- The in-house monitoring product development and QE team(s)
- Internal ITSM/ticketing teams using ServiceNow
- Multiple enterprise DBA organizations across:
- Oracle Database
- Open-source database platforms (MySQL PostgreSQL MongoDB Cassandra etc.)
- Microsoft SQL Server
The ideal candidate combines strong operational troubleshooting skills production support leadership monitoring expertise incident coordination capabilities and excellent cross-team collaboration skills in large-scale enterprise environments.
This is a senior leadership position: the candidate must bring substantial experience leading Production Support Engineers combined with exceptional communication skills for direct confident engagement with customers and senior stakeholders of in-house management.
Key Responsibilities
Production Operations & Monitoring Support
- Provide operational ownership and production support for the in-house enterprise monitoring platform.
- Monitor health performance alerting quality and operational stability of monitoring services.
- Analyze monitoring gaps false positives missed alerts and operational inefficiencies.
- Ensure monitoring coverage across Oracle Open-source and SQL Server database environments.
Incident & Escalation Management
- Act as the operational point-of-contact during production incidents involving monitoring failures alerting gaps or infrastructure issues.
- Coordinate incident triage across:
- DBA teams
- Monitoring development teams
- Infrastructure teams
- Service management teams.
- Drive bridge calls and ensure effective stakeholder communication during critical outages.
- Perform root cause analysis (RCA) and post-incident operational reviews.
ServiceNow & Ticket Workflow Coordination
- Work with ServiceNow for:
- Operational escalations
- Service requests.
- Review ticket quality and ensure operational accuracy of issue classification and routing.
- Improve ticket workflows between DBA teams and monitoring platform support teams.
- Collaborate with internal support organizations to streamline escalation processes.
Cross-Functional DBA Collaboration
- Collaborate closely with enterprise DBA teams supporting:
- Oracle Database
- MySQL
- PostgreSQL
- MongoDB
- Apache Cassandra
- Microsoft SQL Server
- Cloud services (AWS AZURE GCP)
- Understand operational monitoring requirements specific to each database technology.
- Work with DBAs to validate alert thresholds event correlation and monitoring accuracy.
- Serve as the operational liaison between DBAs and monitoring team developers and QE.
Operational Excellence & Reliability Engineering
- Identify recurring operational pain points and recommend automation opportunities.
- Improve alert quality event correlation and monitoring reliability.
- Participate in operational readiness reviews for new monitoring features.
- Help define operational standards playbooks and escalation procedures.
Monitoring & Observability Engineering
- Support enterprise observability initiatives involving:
- Metrics
- Events
- Alerting
- Dashboards
- Health monitoring
- Incident correlation.
- Work with both commercial and in-house monitoring systems.
- Analyse operational telemetry to identify systemic reliability concerns.
DevOps & CI/CD Enablement
- Collaborate with engineering teams to improve CI/CD pipelines.
- Implement deployment strategies (blue-green canary rolling updates).
- Advocate for reliability-focused design patterns.
Security & Compliance
- Ensure infrastructure adheres to security standards and compliance requirements.
- Participate in vulnerability assessments and remediation.
Required Technical Skills
- Strong production support and operations experience in enterprise environments.
- Substantial experience leading Production Support Engineers in a senior/lead or team-management capacity.
- Strong experience with cloud platforms (AWS Azure and GCP).
- Substantial practical hands-on experience with both Windows and Linux operating systems and sound working familiarity with Databricks and Microsoft Fabric for analytics workloads.
- Expertise in monitoring & observability tools (e.g. Prometheus Grafana Datadog or in-house tools).
- Working knowledge of:
- ServiceNow
- Incident workflows
- Escalation management
- Operational support models.
- Exposure to database technologies including:
- Oracle Database
- Microsoft SQL Server
- MySQL
- PostgreSQL
- NoSQL ecosystems preferred.
- Strong understanding of:
- Windows and Linux systems
- Infrastructure monitoring
- Alerting concepts
- Production operations.
- Experience supporting 24x7 enterprise production environments.
Preferred Qualifications
- Experience working with in-house monitoring or observability product teams.
- Familiarity with SRE/DevOps operational practices.
- Exposure to enterprise event management systems.
- Knowledge of automation/scripting (Python Shell PowerShell).
- Experience handling high-severity production incidents.
Critical Non-Technical Skills
An ideal candidate must demonstrate:
Operational Intuition
- Ability to detect operational anomalies early.
- Strong troubleshooting instinct and pattern recognition.
Fearless Communication
- Ability to speak confidently during incidents and escalations.
- Comfortable engaging customers and senior stakeholders of in-house management along with multiple technical teams.
- Exceptional written and verbal communication skills able to present operational status and risk directly to customers and senior leadership with clarity and confidence.
Cross-Team Collaboration
- Ability to coordinate effectively across DBA teams support organizations and development groups.
Calmness Under Pressure
- Structured decision-making during high-severity incidents.
Ownership Mindset
- Drives issues to closure rather than relying solely on assigned ownership boundaries.
Investigative Curiosity
- Continuously analyses why operational failures occur and how they can be prevented.
- Substantial track record leading mentoring and developing Production Support Engineers.
- Sets the standard for operational excellence and coaches junior/mid-level engineers toward it.
- Acts as an escalation point and mentor for less experienced engineers during high-severity incidents.
- Team Leadership & Mentorship
Preferred Qualifications
- Relevant Certifications in cloud platforms (AWS/Azure/GCP).
- Familiarity with SRE/DevOps operational practices.
- Familiarity with Databricks and Fabric domain.
- Exposure to enterprise event management systems.
- Knowledge of automation/scripting (Python Shell PowerShell).
Required Experience:
Senior IC
About Company
Datavail is a leading provider of data management, application development, analytics, and cloud services, with more than 1,000 professionals helping clients build and manage applications and data via a world-class tech-enabled delivery platform and software solutions across all leadi ... View more