1. Leadership & Strategy
- Lead, mentor, and grow the SRE team; set clear goals, on-call structure, and career paths.
- Define and own SRE strategy, roadmap, and best practices aligned with business and compliance requirements.
- Drive a culture of reliability, automation, and blameless postmortems.
2. Reliability & Availability
- Own SLAs, SLOs, and SLIs for all production platforms (core banking, APIs, payments).
- Ensure 99.9%+ availability of critical services and lead efforts to eliminate single points of failure.
- Manage capacity planning, scalability, and disaster recovery (DR/BCP) strategies.
3. Infrastructure & Automation
- Own and evolve our cloud and on-prem infrastructure (AWS/Azure, Kubernetes, Docker, Terraform).
- Drive Infrastructure as Code (IaC), CI/CD, and GitOps maturity to enable safe, frequent releases.
- Lead automation of operational toil, provisioning, and configuration management.
4. Incident & Problem Management
- Own the incident response lifecycle - detection, escalation, resolution, and post-incident review.
- Build and improve monitoring, alerting, logging, and observability stacks (Prometheus, Grafana, ELK/Datadog, PagerDuty).
- Act as final escalation for P1/P2 incidents.
5. Security & Compliance
- Partner with Security and Compliance to ensure infrastructure meets PCI-DSS, NDPA, CBN, and ISO 27001 requirements.
- Embed security, secrets management, and vulnerability remediation into SRE practices.
- Own change management and audit readiness for infrastructure changes.
6. Collaboration
- Collaborate closely with Software Engineering, Product, Security, and Client Success to ensure reliability is built-in.
- Provide technical guidance to engineering teams on resilient architecture patterns.