← Усі вакансії

Senior Observability Engineer, Cloud Platform

Рівень:
senior
Джерело:
djinni.co
Відгукнутись на вакансію →

About Our Client

Our client builds and operates GPU cloud infrastructure at the scale modern AI demands. They run dense accelerated-compute clusters on bare metal and Kubernetes, stitched together with high-performance Ethernet and InfiniBand fabrics, and deliver them to customers as reliable, secure, high-throughput platforms. They are building Cloud v2, their next-generation, clean-slate, Kubernetes-native GPU cloud, and scaling it across new data center sites. The infrastructure team sits at the center of that effort, building the systems that turn racks of hardware into a product customers depend on.

The role

We are hiring a Senior Observability Engineer to build the observability layer for Cloud v2. This role exists to close the single most damaging gap in our client’s current product: today their observability is reactive, so customers are their failure detector. Some of their worst incidents went undetected internally because they occurred below the layer they instrument (a corrupted forwarding entry, a flapping firewall sync process), and the customer reported the outage before their dashboards did. Cloud v2 is architected from day one against the industry’s Gold-tier bar for neocloud platforms (SemiAnalysis ClusterMAX), and proactive detection, where they see problems before the customer does, is the trust gap they are building this platform to fix.

You own the metrics, logs, alerting, dashboards, and detection stack that make that true. That means operator-facing observability that catches failures first and drives automated remediation, and tenant-facing observability that gives customers the usage, health, and status visibility a hyperscaler-grade cloud implies. This is the observability workstream the platform validates earliest, because detection is where the current product loses trust, and it maps directly to the ClusterMAX Monitoring and Reliability criteria.

This is a build-focused, hands-on senior IC role. You will spend your time engineering the telemetry pipelines, detection logic, dashboards, and alerting that turn raw signals from multi-vendor GPUs (AMD/NVIDIA), hosts, fabrics, and switches into problems we catch before customers feel them, and you will set the technical direction and standards for how we instrument the platform. You will partner closely with our client’s SRE, platform, network, and security engineers, who own the control-plane architecture, the fabric, and day-to-day operations; you own the observability that makes the whole platform legible and its failures visible.

What you’ll own

The OpenTelemetry foundation: own OpenTelemetry as the instrumentation standard the whole platform plugs into: the collector topology and pipelines, semantic conventions for metrics, logs, and traces, and consistent resource and tenant attribution across GPU, host, fabric, switch, and control-plane signals. Everything we instrument, internal and tenant-facing, lands on this common backbone so signals are portable, correlated, and not locked to any one backend.

Proactive detection: build the detection that fires before a customer reports, GPU failure-rate tracking, Xid detection, and the anomaly and threshold signals that feed the automated detect-drain-remediate loop, so mean time to detect stops being inverted and our client sees the problem first.

The metrics pipeline: own the metrics stack across GPU, CPU, network planes, thermal, and power (Prometheus and OpenTelemetry-based), including the DCGM profiling-metric suite (DCGM_FI_PROF_* for SM activity and tensor TFLOPs, plus PCIe AER, ECC, power and thermal, and NVLink and fabric throughput), with custom exporters and OpenTelemetry receivers where DCGM falls short.

Instrumenting below the layer we watch today: push instrumentation down into the sub-layers where our most damaging silent failures live (forwarding state, firewall and HA sync, fabric health), so the failure classes that used to reach us through the customer surface on our o

Схожі вакансії

З блогу Trackr

Усі статті →

Знайдено через trackr.help/jobs · Канал: @trackrhelp · Бот для персональних сповіщень: @trackrhelpBot