About the Role
Heirs Insurance is a general insurance company challenging traditional insurance by providing simple and accessible protection for vehicles, homes, business and more. We are looking for a dedicated Site Reliability Engineer to keep our high-stakes platform running smoothly, securely, and efficiently.
Key Responsibilities- Work within the 24/7 Network and Security Operations Centre actively monitoring platform health, network layers, and cloud environments.
- Serve as a first responder for incident response, executing runbooks and conducting blameless post-mortems.
- Implement and maintain observability stacks including centralized logging, metrics collection, and distributed tracing.
- Build automation scripts in Python or Bash to eliminate operational toil and enhance platform reliability.
- Support reliability engineering through chaos testing, load tests, and proactive capacity monitoring on AWS and Azure.
- Track SLO attainment, error budgets, and DORA metrics to maintain high standards of uptime and performance.
- Minimum of 2 years of SRE, DevOps, or platform operations experience in a high-availability production environment.
- Hands-on expertise with observability tooling such as Prometheus, Datadog, ELK, or Loki.
- Strong scripting skills in Python and Bash for automation and toil reduction.
- Solid operational understanding of cloud infrastructure on AWS and/or Azure.
- Demonstrated experience in incident response, triage, and root-cause analysis under pressure.
- Familiarity with financial services or other regulated environments is an added advantage.
- Competitive salary and performance-based bonuses.
- Comprehensive health insurance and wellness coverage.
- Continuous learning opportunities with professional certification support.
- A collaborative and dynamic work culture at a leading Nigerian insurer.