Peraton is currently seeking skilled and qualified candidates for our Cloud Operations & Engineering Team Lead role to provide technical leadership for a team of five cloud engineers responsible for provisioning, operating, and maintaining the full breadth of infrastructure services across that estate. Peraton staff support a Cloud Business Office operating a large enterprise AWS environment — 500+ AWS accounts, along with a smaller Azure footprint — supporting a broad portfolio of enterprise workloads.
This is a team lead position, and it is a hands-on one. The person in this role sets team priorities, provides technical direction, and represents the team to stakeholders — while remaining personally in the environment doing the work. Expect roughly 50–60% hands-on engineering and operations, and 40–50% team leadership and coordination. This is not a role that directs from a distance.The lead does not carry formal supervisory authority, but does have input into hiring, interviews, and performance feedback.
This position is 100% remote, working East Coast hours. This is an operations role; support needs may arise outside standard business hours.
Day to Day Responsibilities:
- Lead the day-to-day technical direction of the cloud engineering and operations team, including workload balancing across the account estate
- Scope and prioritize ServiceNow requests and changes, and Jira tasks, for provisioning, modifying, troubleshooting, and decommissioning cloud services
- Plan operational support activities — managed service patching, maintenance windows, cost optimization support — with stakeholder engagement
- Manage the team's on-call and escalation model, including severity definitions and escalation paths
- Identify recurring operational pain points and toil, and drive improvements to monitoring, automation, and operational process
- Advise on and improve team processes using cloud best practices and established team standards
- Coordinate with counterparts across the Cloud Business Office and other technical teams
- Communicate status, issues, risks, and concerns effectively to leadership
- Contribute to hiring, interviews, and performance feedback for the team
- Performs incident management and troubleshooting in an operations environment including stakeholder engagement for issue resolution
- Manages operations in cloud environments and ensure the high availability and performance of those environment
- Serves in the on-call and escalation rotation alongside the team, and personally leads technical bridges for incidents
- Performs root cause analysis and writes post-incident reports
- Executes deployments using the provided Terraform modules and pipelines — running plans, applying changes, and troubleshooting failed runs and state issues across the account estate
- Reads, troubleshoots, and maintains existing CloudFormation stacks in legacy environments
- Provisions, configures, and operates infrastructure across a broad AWS service catalog, and grants and governs the permissions that allow application developers to work within their environments
- Manages full EC2 instance lifecycle: provisioning, resizing, AMI selection and refresh, snapshots, and replacement
- Provisions RDS instances and performs managed service patching and maintenance window planning
- Owns backup and recovery for provisioned resources using enterprise AWS Backup policies, including restore requests
- Administers AWS access management: authors IAM Identity Center permission sets and creates the group, account, and permission set assignments; manages the Entra ID groups that SCIM into AWS; maintains roles, policies, and permission boundaries at scale
- Builds operational automation using Python/boto3 and Lambda to eliminate recurring manual toil — for example, resource inventory and reporting, and monitoring automation
- Evaluates and operationalizes new cloud services, tools, and technologies