jobs
Senior AI Infrastructure Engineer, Observability
Firmus Technologies Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific. Founded in Australia in 2019, our mission is to create the most efficient AI infrastruct
- Company
- Firmus
- Location
- Singapore
- Status
- Open
- Posted
- 2026-09-04T05:23:56+00:00
Firmus is hiring a senior engineer to own observability for its GPU fleets. The role sets GPU and host health criteria and service-readiness gates, then turns them into dashboards, alerts, PromQL/LogQL queries, DCGM and NCCL diagnostic checks and runbooks. It covers fault isolation — separating a bad GPU from cooling, host, network or power-limit problems using host and BMC telemetry — and joining incidents. It asks for 7+ years in GPU, HPC or AI infrastructure, production GPU fault diagnosis, Linux and Python, and Grafana-class tooling. On-call and occasional overseas travel are expected.
Original job posting