Team Lead, Site Reliability Engineering

Moniepoint Inc. | Lagos, Nigeria

Full-Time Posted today 0 views

About the Role

Moniepoint is a leading financial technology company digitizing Africa's real economy by building a comprehensive financial ecosystem for businesses. We are currently seeking a dynamic and visionary Team Lead, Site Reliability Engineering to guide a squad of talented site reliability engineers. In this pivotal role, you will design high-level reliability architecture, mentor engineers, define our technical roadmap, and drive a culture of engineering excellence across our rapidly scaling financial platform in Nigeria.

Key Responsibilities
  • Set the technical direction for the SRE team, architecting self-healing systems, defining Production Readiness Reviews, and driving automation best practices.
  • Enforce end-to-end system visibility by guiding teams to instrument code effectively and minimizing alert fatigue.
  • Lead, mentor, and grow a team of Senior and Associate SREs through code reviews and technical workshops.
  • Act as the ultimate escalation point for major incidents, refining incident management processes and ensuring thorough root cause analyses.
  • Partner with Engineering Managers and Product Leads to define Service Level Objectives that align with business goals.
Qualifications
  • Minimum of 6 years of experience in SRE or Backend Engineering, with at least 2 years in a leadership or senior role.
  • Expert-level proficiency in Java, Go, Rust, or Python.
  • Mastery of distributed systems patterns and microservices architecture.
  • Deep expertise with cloud platforms like GCP or AWS, alongside extensive experience running Kubernetes at scale.
  • Strong communication skills with the ability to manage high-pressure situations with calm authority.
Benefits
  • Competitive salary with pension and comprehensive health insurance plans.
  • A people-first culture that prioritizes team member well-being and professional growth.
  • Dedicated learning and development environment featuring knowledge sharing and technical talks.

Skills

Site Reliability Engineering Distributed Systems Kubernetes GCP AWS Python Go Java Observability Incident Management
Report this job listing