Yobitel builds and operates the AI-native compute layer that our customers' training and inference workloads run on. We own the hardware rather than reselling someone else's: H100 fleets in production today, B200 capacity landing next. When a rack degrades at 3am, it is our pager and our problem.
We are hiring a senior systems engineer to keep that fleet fast, healthy and full. This is a deep systems role that sits close to the metal, spanning firmware and driver stacks, NUMA and topology tuning, RDMA fabric behaviour, and the performance regressions that only ever show up under real multi-node load. You will work alongside the platform and ML infrastructure teams, and your work sets the ceiling on what every workload above you can achieve.
- In your first 90 days, you have commissioned a rack end to end and written the runbook that lets someone else repeat it.
- Node-level performance regressions get caught by automated checks before customers feel them.
- Fleet utilisation and time-to-recovery both improve, and you can show the numbers.