Senior DevOps & Infrastructure
Job description
- 2+ years of experience in a DevOps, SRE, or cloud infrastructure role.
- Hands-on experience operating production workloads on DigitalOcean (or a comparable cloud
provider with willingness to work primarily in DigitalOcean).
- Strong grasp of infrastructure security: firewalls, VPCs, secrets management, certificate
management, and hardening Linux servers.
- Experience setting up centralized logging and making logs useful for debugging.
- Experience building monitoring and alerting with tools like Prometheus, Grafana, Datadog,
or similar.
- Practical knowledge of database replication, backups, and high-availability configurations
(PostgreSQL, MySQL, or similar).
- Proficiency with Linux system administration.
- Experience with Infrastructure as Code (Terraform, Ansible, or similar) and scripting (Bash,
Python).
- Familiarity with CI/CD tooling (GitHub Actions, GitLab CI, Jenkins, or similar).
- Solid understanding of networking fundamentals (DNS, TCP/IP, HTTP/HTTPS, load
balancing).
Nice to Have
- Experience with containerization (Docker).
- Familiarity with security and compliance frameworks (SOC 2, ISO 27001, or similar).
- Experience with secrets management tools (Vault, Doppler, or cloud-native equivalents).
- Exposure to other cloud providers (AWS, GCP) for multi-cloud or migration scenarios.
- A track record of cost optimization on cloud infrastructure.
- Provision, configure, and maintain cloud infrastructure on DigitalOcean (Droplets, Managed
Databases, Spaces, Load Balancers, VPCs, and related services).
- Design and enforce security best practices across the stack: network segmentation, firewalls
and security groups, SSH key and secrets management, least-privilege IAM, TLS/SSL
certificates, and regular patching.
- Stand up and maintain centralized logging pipelines so engineers can quickly trace issues
across services (e.g., ELK/Elastic, Loki, or a managed logging stack).
- Build and tune monitoring and alerting for infrastructure and application health — uptime,
resource utilization, latency, and error rates — with sensible, actionable alerts (e.g.,
Prometheus + Grafana, or equivalent).
- Configure and manage database replication and high availability (primary/replica setups,
failover, backups, and restore testing) to protect data and minimize downtime.
- Automate provisioning and configuration using Infrastructure as Code and scripting (e.g.,
Terraform, Ansible, Bash).
- Build and maintain CI/CD pipelines that let developers deploy reliably and roll back safely.
- Manage backups, disaster recovery procedures, and routine restore drills.
- Respond to incidents, perform root-cause analysis, and document runbooks to prevent
recurrence.
- Continuously optimize for cost, performance, and reliability across the infrastructure footprint.