About the role
Database Site Reliability Engineer (Database Operations)
Position Summary
We are seeking an experienced Database Site Reliability Engineer (SRE) to support and operate mission-critical database platforms within a fast-paced enterprise environment. This role is focused on operational excellence, reliability, resiliency, automation, and continuous improvement across multiple database technologies.
This is not a development role. We are looking for hands-on database professionals who thrive in production operations, take ownership of issues, understand urgency, and are passionate about improving systems, processes, and themselves.
The ideal candidate views reliability as a product, proactively identifies risks before they become incidents, and continuously seeks opportunities to automate repetitive tasks and improve platform stability.
Key Responsibilities
Database Operations & Reliability
Install, configure, upgrade, patch, and maintain enterprise database platforms.
Ensure availability, performance, recoverability, and security of production database environments.
Monitor database platforms and associated infrastructure, responding rapidly to incidents and service degradations.
Lead troubleshooting efforts for database, operating system, storage, replication, and application connectivity issues.
Execute failovers, disaster recovery testing, and recovery procedures.
Partner with application teams to provide database guidance and operational support.
Platform Engineering
Deploy, maintain, and optimize database infrastructure across physical, virtual, and cloud environments.
Implement scalable, resilient database solutions.
Evaluate and recommend improvements to architecture, monitoring, automation, and operational processes.
Support capacity planning, performance tuning, and platform lifecycle management.
Automation & Continuous Improvement
Develop and maintain automation solutions using Python, Shell, Ansible, or similar technologies.
Help eliminate manual operational activities through engineering and automation.
Improve monitoring, alerting, reporting, and operational workflows.
Drive incremental improvements that reduce risk, improve reliability, and increase operational efficiency.
Performance & Incident Management
Analyze and resolve database performance issues.
Troubleshoot replication, backup/recovery, storage, network, and infrastructure-related incidents.
Participate in root cause analysis and drive permanent corrective actions.
Review operational metrics and trends to identify opportunities for improvement.
Operational Excellence
Maintain accurate operational documentation, standards, and procedures.
Generate and present operational metrics, service health indicators, and reliability reporting.
Participate in incident response activities.
Demonstrate strong ownership from issue identification through resolution.
Required Qualifications
Strong experience administering enterprise database platforms, including:
o Sybase ASE
Oracle RAC
Additional database technologies such as MongoDB, Cassandra, Redis, PostgreSQL, MySQL, or similar platforms are a plus.
Experience performing:
o Installation
Configuration
Upgrades
Patching
Performance tuning
Backup and recovery
High availability and disaster recovery
Experience with database replication technologies including:
o SAP Replication Server
Data Guard
HVR (preferred)
Strong Linux administration skills.
Experience with automation and scripting:
o Python
Ansible
Shell scripting
Understanding of storage, networking, operating systems, and infrastructure services.
Experience with Veritas Cluster Server, ASM, LVM, and SAN technologies.
Familiarity with enterprise operational tooling such as Jira, Service Now and Confluence.
Strong analytical, troubleshooting, and problem-solving skills.
What Success Looks Like
The successful candidate:
Takes ownership and drives issues to closure.
Understands the urgency required to support critical production environments.
Continuously improves systems, processes, and operational effectiveness.
Learns quickly and adapts to new technologies.
Balances operational stability with engineering innovation.
Communicates clearly and effectively during incidents and high-pressure situations.
Demonstrates a strong sense of accountability and professionalism.
Leaves the platform better than they found it every day.
Preferred Mindset
We hire for attitude as much as technical skill.
We're looking for individuals who are:
Customer obsessed
Accountable and dependable
Urgent without being reckless
Continuously learning
Driven to automate repetitive work
Detail-oriented
Collaborative but willing to lead
Focused on long-term platform reliability rather than short-term fixes