We are building the next generation of AI infrastructure by making GPU computing more flexible, efficient, and accessible. Our platform enables customers to consume GPU capacity through real-time marketplaces, spot bidding, and flexible reservations—bringing cloud-scale infrastructure together with market-driven resource allocation.
We're growing rapidly and are looking for a talented Backend Software Engineer to help build the distributed systems that power our platform.
-
Palo Alto / San Francisco, CA
-
Hybrid/Onsite (4 days per week)
-
Team is onsite in Palo Alto every Monday
-
$170,000 – $230,000 base salary
-
Exceptional candidates may be considered up to $250,000
-
Competitive equity package
-
Comprehensive benefits
-
Design, build, and maintain highly scalable backend systems supporting real-time GPU allocation.
-
Develop and evolve distributed services responsible for resource scheduling, reservation management, and capacity allocation.
-
Build reliable backend infrastructure capable of handling concurrency, asynchronous workflows, and race conditions.
-
Design and optimize pricing, allocation, and orchestration logic for large-scale distributed systems.
-
Partner closely with cross-functional engineering teams to ensure efficient infrastructure utilization.
-
Own projects from technical design through implementation, deployment, monitoring, and long-term maintenance.
-
Improve platform reliability, resiliency, observability, and operational excellence.
-
Build and maintain CI/CD pipelines, cloud infrastructure, and deployment automation.
-
Participate in production on-call rotations and continuously improve operational tooling and runbooks.
-
3+ years of professional backend software engineering experience (5–8 years preferred).
-
Strong expertise building and maintaining production distributed systems at scale.
-
Proven experience handling:
-
Concurrency
-
Race conditions
-
Asynchronous request processing
-
Distributed system reliability
-
Strong proficiency in Python.
-
Experience with PostgreSQL or other relational databases.
-
Hands-on experience with Kubernetes and Docker.
-
Experience deploying cloud-native applications on AWS and/or GCP.
-
Familiarity with Infrastructure as Code tools such as Terraform.
-
Strong understanding of software architecture, scalability, and production operations.
-
Experience participating in system design and architecture discussions.
-
Experience at Big Tech or VC-backed startups.
-
Experience building marketplace platforms, auction engines, exchanges, reservation systems, billing platforms, or pricing engines.
-
Experience with observability, monitoring, and production incident response.
-
Strong understanding of distributed scheduling or orchestration systems.
-
Comfortable contributing across multiple areas of the technology stack.
We're looking for someone who:
-
Enjoys solving complex distributed systems problems.
-
Has built and maintained production systems—not just prototypes or MVPs.
-
Takes ownership from design through long-term operational support.
-
Is comfortable working in a fast-paced startup environment.
-
Can influence architecture while remaining highly hands-on.
-
Thrives in collaborative, cross-functional engineering teams.
-
Python
-
PostgreSQL
-
Kubernetes
-
Docker
-
AWS
-
GCP
-
Terraform
-
CI/CD
-
Distributed Systems
-
Cloud Infrastructure