AWS Infrastructure Services owns the design, planning, delivery, and operation of all AWS global infrastructure. In other words, we’re the people who keep the cloud running. We support all AWS data centers and all of the servers, storage, networking, power, and cooling equipment that ensure our customers have continual access to the innovation they rely on. We work on the most challenging problems, with thousands of variables impacting the supply chain — and we’re looking for talented people who want to help.
The Data Center Global Operations: Central Operations Install Team within Amazon Web Services (AWS) is looking for a talented systems engineer to join our team. We are a team that provides technical support to the global operations business within AWS data centers. Our vision is to ensure that work driven into AWS Data Centers is safe, secure, serviceable, supportable and highly available for customers. As a Subject Matter Expert (SME) for the operations team, you will primarily be working within the organization to improve our install workflows and fix systemic issues to allow the location operations teams to hand off capacity more smoothly and efficiently.
Aside from your direct team, your daily work engagements will be with the network engineering, service teams, and local operations teams. You’ll join a diverse team of software, hardware, and network engineers, supply chain specialists, security experts, operations managers, and other vital roles. You’ll collaborate with people across AWS to help us deliver the highest standards for safety and security while providing seemingly infinite capacity at the lowest possible cost for our customers. And you’ll experience an inclusive culture that welcomes bold ideas and empowers you to own them to completion.
Key job responsibilities
- You will develop solutions to recover capacity for customers.
- You will work with support and systems engineers to push upstream fixes to enable capacity recovery.
- You will own and make improvements to global directed workflows that enable local teams to hand off capacity.
- You will work to deep dive defects throughout the capacity installation process, and work to develop solutions to mitigate defects.
- You will contribute to data analytics to support your role.
- You will evaluate data and trending patterns to uncover defects across product life cycles.
A day in the life
A day in the life of this role will involve a mix of working with our internal DCC Ops, Support/System Engineering, TPM's, Hardware Engineering, and Software teams to address systemic issues and resolve capacity recovery campaigns.
- 4+ years of site reliability engineering (SRE), systems engineering, systems administration, DevOps, security administration, or network administration experience
- 5+ years of systems engineering experience
- Bachelor's degree in Systems Engineering, Computer Science, or related field or relevant work experience
- Experience in site reliability engineering (SRE), systems engineering, systems administration, DevOps, security administration, or network administration
- Experience working with Linux
- Experience in any of the following: Python, Java, Perl, PHP, Ruby, Bash, Shell or equivalent
- Knowledge of TCP/IP and networking protocols such as HTTP and DNS
- Experience designing and developing scripts to automate operational burdens and reviewing scripting changes to ensure they meet the standards for maintainability, scalability and security
- Experience working in 24/7 production environment
- Experience with service-oriented architecture and web services
Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.
Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.
The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.
USA, OH, Dublin - 104,500.00 - 160,000.00 USD annually
USA, VA, Herndon - 104,500.00 - 160,000.00 USD annually
USA, WA, Seattle - 104,500.00 - 160,000.00 USD annually