CockroachDB Database Administrator
Sunnyvale, CA - USA
Job Summary
8-10 Years (with strong experience in distributed databases and production database administration)
We are seeking an experienced CockroachDB Database Administrator (DBA) to design deploy administer and optimize CockroachDB clusters in large-scale production environments. The ideal candidate should have hands-on experience managing distributed SQL databases ensuring high availability disaster recovery performance optimization monitoring automation and production support.
The candidate will be responsible for maintaining highly available fault-tolerant multi-region CockroachDB clusters while collaborating with infrastructure cloud and application teams to ensure database reliability scalability and operational excellence.
- Design deploy configure and manage CockroachDB clusters in production environments.
- Build and maintain multi-region distributed database clusters.
- Ensure high availability fault tolerance and data consistency across geographically distributed environments.
- Monitor cluster health node status replication latency and resource utilization.
- Perform database capacity planning and cluster scaling.
- Monitor and optimize SQL query performance.
- Analyze execution plans and identify bottlenecks.
- Optimize database schema design and indexing strategies.
- Resolve high-latency and throughput issues.
- Address hotspotting leaseholder imbalance and replication lag.
- Troubleshoot complex database and infrastructure issues including:
- Node failures
- Network partitions
- Replication issues
- Leaseholder imbalance
- Range imbalance
- High CPU or memory utilization
- Storage issues
- Performance bottlenecks
- Participate in incident management and on-call support.
- Design and implement disaster recovery strategies.
- Configure backup and restore processes.
- Implement Point-in-Time Recovery (PITR).
- Validate backup integrity through regular recovery testing.
- Manage failover and failback procedures.
- Automate provisioning deployment upgrades scaling and maintenance of CockroachDB clusters.
- Perform rolling upgrades with minimal or zero downtime.
- Develop automation scripts using Bash Python or similar scripting languages.
- Improve operational efficiency through Infrastructure as Code (IaC).
- Monitor database health using observability tools.
- Configure alerts and dashboards.
- Track:
- Cluster health
- Replication status
- Latency
- Resource utilization
- Storage growth
- Query performance
- Work with monitoring tools such as Grafana Prometheus and Cloud monitoring platforms.
- Create operational runbooks.
- Develop Standard Operating Procedures (SOPs).
- Maintain architecture diagrams and recovery procedures.
- Document production support processes and troubleshooting guides.
- CockroachDB
- Distributed SQL Databases
- SQL Performance Tuning
- Database Replication
- High Availability (HA)
- Multi-region Database Deployment
- Database Clustering
- AWS / Azure / GCP
- Linux Administration
- Kubernetes (Preferred)
- Docker
- Networking Fundamentals
- Prometheus
- Grafana
- Cloud Monitoring Tools
- Performance Monitoring
- Alert Management
- Bash
- Python
- Shell Scripting
- Terraform (Preferred)
- Ansible (Preferred)
- Backup & Restore
- Point-in-Time Recovery (PITR)
- Disaster Recovery
- Failover / Failback
- Business Continuity Planning
- Production Support
- Incident Management
- Root Cause Analysis (RCA)
- Capacity Planning
- Performance Tuning
- Bachelors degree in Computer Science Information Technology Engineering or a related field.
- 8-10 years of database administration experience.
- Strong hands-on experience with CockroachDB or other distributed SQL databases.
- Experience managing production database environments.
- Expertise in SQL performance tuning and database optimization.
- Experience with Linux system administration.
- Knowledge of cloud platforms (AWS Azure or GCP).
- Strong troubleshooting and analytical skills.
- Excellent communication and documentation abilities.
- Experience with Kubernetes-based database deployments.
- Knowledge of Infrastructure as Code (Terraform Ansible).
- Experience with PostgreSQL or other distributed databases.
- Experience with CI/CD pipelines and DevOps practices.
- CockroachDB certification (if available).
- Experience in mission-critical production support environments.