jobs
Senior AI Infrastructure Engineer, Observability
Firmus Technologies Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific. Founded in Australia in 2019, our mission is to create the most efficient AI infrastruct
In inglese
- Company
- Firmus
- Location
- Singapore
- Status
- Open
- Posted
- 2026-09-04T05:23:56+00:00
Firmus is hiring a senior engineer to own observability for its GPU fleets. The role sets GPU and host health criteria and service-readiness gates, then turns them into dashboards, alerts, PromQL/LogQL queries, DCGM and NCCL diagnostic checks and runbooks. It covers fault isolation — separating a bad GPU from cooling, host, network or power-limit problems using host and BMC telemetry — and joining incidents. It asks for 7+ years in GPU, HPC or AI infrastructure, production GPU fault diagnosis, Linux and Python, and Grafana-class tooling. On-call and occasional overseas travel are expected.
Original job posting