JobHabor

Lead Senior Production Support / Operations Engineer

Datavail
Location
Mumbai, Maharashtra, India
Workplace
Hybrid
Employment
Full Time
Salary
Apply on the employer’s site

Posted 2mo ago

Job Title

Lead Senior Production Support / Operations Engineer

Education: Any Graduate

Experience: 8+years

Job Location: Mumbai

Role Summary

We are seeking a highly Senior Production Support / Operations Engineer to represent and operationally support our in-house monitoring and observability platform across enterprise database environments.

This role acts as the critical operational bridge between:

  • The in-house monitoring product development and QE team(s),
  • Internal ITSM/ticketing teams using ServiceNow,
  • Multiple enterprise DBA organizations across:
  • Oracle Database
  • Open-source database platforms (MySQL, PostgreSQL, MongoDB, Cassandra, etc.)
  • Microsoft SQL Server

The ideal candidate combines strong operational troubleshooting skills, production support leadership, monitoring expertise, incident coordination capabilities, and excellent cross-team collaboration skills in large-scale enterprise environments.

This is a senior leadership position

the candidate must bring substantial experience leading Production Support Engineers, combined with exceptional communication skills for direct, confident engagement with customers and senior stakeholders of in-house management.

Key Responsibilities

Production Operations & Monitoring Support

  • Provide operational ownership and production support for the in-house enterprise monitoring platform.
  • Monitor health, performance, alerting quality, and operational stability of monitoring services.
  • Analyze monitoring gaps, false positives, missed alerts, and operational inefficiencies.
  • Ensure monitoring coverage across Oracle, Open-source, and SQL Server database environments.

Incident & Escalation Management

  • Act as the operational point-of-contact during production incidents involving monitoring failures, alerting gaps, or infrastructure issues.
  • Coordinate incident triage across:
  • DBA teams,
  • Monitoring development teams,
  • Infrastructure teams,
  • Service management teams.
  • Drive bridge calls and ensure effective stakeholder communication during critical outages.
  • Perform root cause analysis (RCA) and post-incident operational reviews.

ServiceNow & Ticket Workflow Coordination

  • Work with ServiceNow for:
  • Operational escalations,
  • Service requests.
  • Review ticket quality and ensure operational accuracy of issue classification and routing.
  • Improve ticket workflows between DBA teams and monitoring platform support teams.
  • Collaborate with internal support organizations to streamline escalation processes.

Cross-Functional DBA Collaboration

  • Collaborate closely with enterprise DBA teams supporting:
  • Oracle Database
  • MySQL
  • PostgreSQL
  • MongoDB
  • Apache Cassandra
  • Microsoft SQL Server
  • Cloud services (AWS, AZURE, GCP)
  • Understand operational monitoring requirements specific to each database technology.
  • Work with DBAs to validate alert thresholds, event correlation, and monitoring accuracy.
  • Serve as the operational liaison between DBAs and monitoring team developers and QE.

Operational Excellence & Reliability Engineering

  • Identify recurring operational pain points and recommend automation opportunities.
  • Improve alert quality, event correlation, and monitoring reliability.
  • Participate in operational readiness reviews for new monitoring features.
  • Help define operational standards, playbooks, and escalation procedures.

Monitoring & Observability Engineering

  • Support enterprise observability initiatives involving:
  • Metrics,
  • Events,
  • Alerting,
  • Dashboards,
  • Health monitoring,
  • Incident correlation.
  • Work with both commercial and in-house monitoring systems.
  • Analyse operational telemetry to identify systemic reliability concerns.

DevOps & CI/CD Enablement

  • Collaborate with engineering teams to improve CI/CD pipelines.
  • Implement deployment strategies (blue-green, canary, rolling updates).
  • Advocate for reliability-focused design patterns.

Security & Compliance

  • Ensure infrastructure adheres to security standards and compliance requirements.
  • Participate in vulnerability assessments and remediation.

Required Technical Skills

  • Strong production support and operations experience in enterprise environments.
  • Substantial experience leading Production Support Engineers, in a senior/lead or team-management capacity.
  • Strong experience with cloud platforms (AWS, Azure, and GCP).
  • Substantial, practical hands-on experience with both Windows and Linux operating systems, and sound working familiarity with Databricks and Microsoft Fabric for analytics workloads.
  • Expertise in monitoring & observability tools (e.g., Prometheus, Grafana, Datadog, or in-house tools).
  • Working knowledge of:
  • ServiceNow
  • Incident workflows,
  • Escalation management,
  • Operational support models.
  • Exposure to database technologies including:
  • Oracle Database
  • Microsoft SQL Server
  • MySQL
  • PostgreSQL
  • NoSQL ecosystems preferred.
  • Strong understanding of:
  • Windows and Linux systems,
  • Infrastructure monitoring,
  • Alerting concepts,
  • Production operations.
  • Experience supporting 24x7 enterprise production environments.

Preferred Qualifications

  • Experience working with in-house monitoring or observability product teams.
  • Familiarity with SRE/DevOps operational practices.
  • Exposure to enterprise event management systems.
  • Knowledge of automation/scripting (Python, Shell, PowerShell).
  • Experience handling high-severity production incidents.

Critical Non-Technical Skills

An ideal candidate must demonstrate

Operational Intuition

  • Ability to detect operational anomalies early.
  • Strong troubleshooting instinct and pattern recognition.

Fearless Communication

  • Ability to speak confidently during incidents and escalations.
  • Comfortable engaging customers and senior stakeholders of in-house management, along with multiple technical teams.
  • Exceptional written and verbal communication skills, able to present operational status and risk directly to customers and senior leadership with clarity and confidence.

Cross-Team Collaboration

  • Ability to coordinate effectively across DBA teams, support organizations, and development groups.

Calmness Under Pressure

  • Structured decision-making during high-severity incidents.

Ownership Mindset

  • Drives issues to closure rather than relying solely on assigned ownership boundaries.

Investigative Curiosity

  • Continuously analyses why operational failures occur and how they can be prevented.
  • Substantial track record leading, mentoring, and developing Production Support Engineers.
  • Sets the standard for operational excellence and coaches junior/mid-level engineers toward it.
  • Acts as an escalation point and mentor for less experienced engineers during high-severity incidents.
  • Team Leadership & Mentorship

Preferred Qualifications

  • Relevant Certifications in cloud platforms (AWS/Azure/GCP).
  • Familiarity with SRE/DevOps operational practices.
  • Familiarity with Databricks and Fabric domain.
  • Exposure to enterprise event management systems.
  • Knowledge of automation/scripting (Python, Shell, PowerShell).

Skills

  • ServiceNow
  • Oracle Database
  • MySQL
  • PostgreSQL
  • MongoDB
  • Cassandra
  • SQL Server
  • Apache
  • AWS
  • Azure
  • GCP
  • Windows
  • Linux
  • Databricks
  • Microsoft Fabric
  • Prometheus
  • Grafana
  • Datadog
  • Python
  • Shell
  • PowerShell

More jobs at Datavail

All 16

Similar roles