About the Role
Duplo is building the platform to power the next generation of financial services. We're on a mission to expand financial access for all through our simple yet powerful banking-as-a-service API. We're seeking an experienced Senior Site Reliability Engineer to lead our infrastructure evolution and ensure the reliability, scalability, and resilience of our backend systems serving the fintech ecosystem.
Key Responsibilities- Lead technical design and implementation for large-scale, high-complexity initiatives that improve system availability, scalability, and resilience
- Design robust service boundaries, define maintainable infrastructure architecture, and establish reliability standards across the engineering team
- Own observability end-to-end—metrics, logging, tracing, and alerting—to catch issues before customers experience them
- Build and refine incident response processes, lead production incident diagnosis and resolution, and drive thorough postmortems
- Collaborate seamlessly with product, compliance, finance operations, support, and backend teams while communicating across technical and non-technical stakeholders
- Make sound architectural decisions on trade-offs, deployment strategies, and rollout safety with broader business impact awareness
- Drive capacity planning, performance tuning, and cost-efficiency while balancing reliability against operational spend
- Define and track SLOs/SLIs to prioritize reliability work against feature delivery
- Mentor engineers and strengthen both systems and team standards
- Minimum 7 years of experience in SRE, DevOps, or backend infrastructure roles with proven expertise in Kubernetes, Docker, and AWS production environments
- Strong scripting and programming ability in TypeScript and/or similar languages for automation and tooling
- Practical expertise with PostgreSQL operations at scale including replication, backup/recovery, and financial system protection patterns
- Production expertise with Redis for caching, coordination, rate-limiting, and distributed locking
- Strong command of observability tools such as Prometheus, Grafana, Datadog, or ELK
- Demonstrated experience designing CI/CD pipelines, infrastructure-as-code tools like Terraform, and secrets management
- Proven track record in fintech or high-stakes systems with provider integration and compliance expertise
- Strong communication skills across engineering, product, compliance, and external partners
- Track record of mentoring and raising operational standards within teams