About Our Client
Our client is a neocloud building purpose-built GPU infrastructure for AI workloads. They operate large-scale clusters powering training and inference for some of the most demanding AI customers in the market and are rapidly expanding. Their infrastructure runs on NVIDIA and AMD accelerators, InfiniBand and high-speed Ethernet fabrics, and a production stack spanning bare-metal provisioning, Kubernetes, and OpenStack.
The team is small, senior, and moving fast. The people who do well in this environment own problems end-to-end and make decisions with incomplete information.
The Role
Our client is hiring a Senior SRE / DevOps Engineer to own major pieces of its GPU cloud platform end-to-end, including the reliability, automation, and operability of the systems that turn racks of GPUs into customer-facing infrastructure.
This is a high-autonomy individual contributor role. You’ll own subsystems outright, lead incidents on the bridge, build automation that removes toil, and mentor mid- and junior-level engineers around you. You’ll debug in production, write the runbook that didn’t exist yesterday, and leave systems better instrumented than you found them. Expect to participate in a shared on-call rotation.
What You’ll Own
Subsystem ownership. Own the reliability, performance, and operability of major platform subsystems end-to-end.
Incident response. Lead live incidents alongside platform and network engineers; author RCAs and drive action items through completion.
Automation & IaC. Build and maintain Ansible, Terraform, and equivalent tooling to reduce toil and codify operational knowledge.
Observability. Define what good monitoring looks like for the systems you own using tools such as Prometheus, Grafana, Checkmk, and Loki, while eliminating low-signal alerts.
Operational readiness. Help bring up new clusters and data center sites in partnership with Data Center Build and Networking teams.
Mentorship. Level up mid- and junior-level engineers through reviews, pairing, and on-call coaching.
What We’re Looking ForRequired
6+ years of experience in SRE, production engineering, DevOps, or infrastructure operations.
Strong hands-on Linux experience at scale, including debugging real networking, storage, and performance issues in production.
Strong Kubernetes operational experience—you understand what breaks at scale, not just how to write a manifest.
Solid networking fundamentals, including L2/L3 and BGP basics.
Fluency with modern automation and Infrastructure as Code tools such as Ansible, Terraform, or equivalent, with a strong bias toward codifying operational knowledge.
A track record of participating in on-call rotations, leading incident response, and producing post-mortems that engineers trust.
Clear, direct written communication. Much of the work is asynchronous.
Strongly Preferred
GPU or HPC infrastructure experience, including technologies such as NCCL, InfiniBand, or DCGM—or a strong appetite to learn them quickly.
OpenStack or bare-metal provisioning experience.
Familiarity with observability stacks such as Prometheus, Grafana, and Checkmk, along with a strong point of view on effective monitoring.
Prior experience at a neocloud, cloud service provider, or infrastructure vendor where uptime had a direct impact on revenue.
How the Team Works
Our client is remote-first across US, LATAM, and EU time zones, with a strong operational culture.
The team writes things down—decisions, architecture, runbooks, and post-mortems—and values those artifacts over unnecessary meetings.
AI tooling is used heavily in day-to-day operations, and engineers are expected to be comfortable working with it and finding ways to gain additional leverage from it.
The team ships, debugs, and iterates. They avoid adding unnecessary process around problems that simply need to be solved.
They value calm over performative urgency and intentionally protect both focus time and sleep.
AI as a Force Multiplier


