Directly manage and develop 3 SRE Managers/Leads and own the overall health, growth, and performance of an ~20-person SRE organization supporting the Paylo product suite.
Set the vision, priorities, and operating cadence for the SRE function; translate business and product priorities into a reliability roadmap your managers can execute against.
Build a strong bench by hiring, coaching, and developing managers and senior engineers while creating clear career paths and succession plans.
Foster a blameless, learning-oriented culture around incidents, on-call, and operational excellence.
Partner closely with engineering directors, product managers, and business stakeholders across the Paylo organization to align reliability investments with business risk and customer impact.
Stay technically engaged day to day by participating in architecture and design reviews, troubleshooting complex production issues, and directly contributing to infrastructure-as-code, Kubernetes manifests/Helm charts, and CI/CD pipelines when needed.
Set and enforce engineering standards for multi-cloud infrastructure across AWS and Azure and for container orchestration on Kubernetes at scale.
Own adoption and standards for GitOps-based continuous delivery using Argo CD/Argo Workflows, including deployment strategy, rollout policy, and multi-cluster promotion.
Own the Infrastructure-as-Code strategy across teams (Terraform, OpenTofu), including module standards, state management, drift detection, and remediation.
Own CI/CD pipeline architecture and standards built on Jenkins, driving build/deploy automation, pipeline reliability, and progressive delivery practices such as blue-green/canary deployments and automated rollback.
Evaluate and guide adoption of new infrastructure tooling and patterns as the platform evolves across AWS and Azure.
Own the observability strategy across all supported products, with deep, hands-on expertise in Datadog (APM, infrastructure monitoring, log management, dashboards, and alerting) as the standard platform for metrics, tracing, and alerting.
Define and drive adoption of SLIs/SLOs, error budgets, and reliability KPIs across the organization, holding managers and teams accountable to them.
Own the incident management program end to end, including on-call structure, escalation paths, severity definitions, postmortems, and follow-through on remediation actions.
Drive root-cause analysis and long-term reliability investments that reduce Sev1/Sev2 frequency and recurrence.
Ensure appropriate resilience, disaster recovery, and capacity planning practices are in place given the sensitivity of payment- and transaction-related systems.
Partner with Security and Compliance to maintain awareness of PCI DSS and related compliance requirements and ensure the SRE organization supports audit and compliance readiness.
Track and report cost, capacity, and operational KPIs to senior leadership.