Dear Job Seeker, to receive Job Digest alerts, please confirm your email address. Check your inbox or spam folder for the confirmation email once you subscribe.

Team Lead, Site Reliability Engineering

Moniepoint Inc. | Lagos, Nigeria

Full-Time Posted 3 weeks ago 5 views

About the Role

Moniepoint is a leading financial technology company digitizing Africa's real economy by building a comprehensive financial ecosystem for businesses. We are currently seeking a dynamic and visionary Team Lead, Site Reliability Engineering to guide a squad of talented site reliability engineers. In this pivotal role, you will design high-level reliability architecture, mentor engineers, define our technical roadmap, and drive a culture of engineering excellence across our rapidly scaling financial platform in Nigeria.

Key Responsibilities
  • Set the technical direction for the SRE team, architecting self-healing systems, defining Production Readiness Reviews, and driving automation best practices.
  • Enforce end-to-end system visibility by guiding teams to instrument code effectively and minimizing alert fatigue.
  • Lead, mentor, and grow a team of Senior and Associate SREs through code reviews and technical workshops.
  • Act as the ultimate escalation point for major incidents, refining incident management processes and ensuring thorough root cause analyses.
  • Partner with Engineering Managers and Product Leads to define Service Level Objectives that align with business goals.
Qualifications
  • Minimum of 6 years of experience in SRE or Backend Engineering, with at least 2 years in a leadership or senior role.
  • Expert-level proficiency in Java, Go, Rust, or Python.
  • Mastery of distributed systems patterns and microservices architecture.
  • Deep expertise with cloud platforms like GCP or AWS, alongside extensive experience running Kubernetes at scale.
  • Strong communication skills with the ability to manage high-pressure situations with calm authority.
Benefits
  • Competitive salary with pension and comprehensive health insurance plans.
  • A people-first culture that prioritizes team member well-being and professional growth.
  • Dedicated learning and development environment featuring knowledge sharing and technical talks.

Skills

Site Reliability Engineering Distributed Systems Kubernetes GCP AWS Python Go Java Observability Incident Management
Report this job listing