Role : Cloud Infrastructure Architect
Location : Hybrid-4 days a week in Redwood City, CA
Full time
We need a senior or staff-level engineer with 4+ years of experience in cloud infrastructure who is equally strong at hands-on coding and systems design. You should be comfortable working in a small, fast-moving team with high ownership and minimal red tape, and have a track record of building and operating scalable, reliable cloud infrastructure. You must have a CS or engineering degree and be proficient in Python — this is a Python-heavy environment.
Role Requirements
-
Cloud Infrastructure Engineer needed to join a growing team.
-
Must handle both coding and systems design; Python, Terraform, and Kubernetes expertise required.
Candidate Requirements
-
4+ years of experience preferred; strong coding skills necessary.
-
Must have a computer science or relevant engineering degree.
Team Structure and Reporting
-
Reports to VP of Engineering; flat structure beyond that.
-
Small engineering team of about 20, offering broad exposure across the stack.
Work experience
-
Cloud infrastructure engineering at production scale (not just backend/microservices)
-
Designing and operating highly available, scalable infrastructure systems
-
Managing Kubernetes-based systems for compute workloads
-
Implementing infrastructure-as-code across environments
-
Education
-
BS in Computer Science, Electrical Engineering, or related engineering field
-
Hard skills
-
Strong Python coding ability (not just scripting) + Bash or Go
-
Strong Kubernetes, Docker, and containerized architecture expertise
-
AWS required; multi-cloud a plus (GCP, Azure, OCI, on-prem)
-
Infrastructure-as-code (Terraform or Pulumi)
-
CI/CD pipelines, monitoring/observability & systems-level debugging
What you'll do:
-
Architect and maintain multi-cloud infrastructure (AWS primary, plus GCP, Azure, OCI, on-prem) to support customer deployments across diverse environments
-
Define and implement infrastructure-as-code using Terraform or Pulumi, driving best practices across the team
-
Design and manage Kubernetes-based systems for model training, inference, and data processing workloads
-
Write production-quality Python code daily — this is not a scripting-only role; you'll pass rigorous coding interviews and build real systems
-
Optimize CI/CD pipelines and streamline deployment of services across customer environments
-
Build monitoring, alerting, and logging systems to ensure high availability and observability
-
Collaborate closely with research and engineering teams to provide infrastructure support for training large-scale ML models
-
Drive cost-efficiency strategies across compute and storage resources
-
Respond to and resolve infrastructure incidents with ownership and urgency
Tech stack
Python, AWS, GCP, Azure, OCI, Kubernetes, Terraform, Pulumi, Docker, Helm, CI/CD, Bash, Go, CloudFormation
-
Cloud infrastructure and developer tools startups (fast-moving companies with strong cloud/Kubernetes expertise)
-
AI research labs and foundation model companies (teams with large-scale cloud infrastructure for training)
-
Cloud-native infrastructure teams at major cloud providers and big tech