jobs
Operations Engineer, Fleet Reliability
CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and
- Company
- CoreWeave
- Location
- Dublin, Ireland
- Status
- Open
- Posted
- 2026-07-22T14:35:18+00:00
CoreWeave's Fleet Reliability Operations team keeps its server fleet provisioned and online, and this role sits on the front line of that effort in Dublin. The engineer configures and maintains large-scale GPU supercomputing clusters, troubleshoots hardware and software faults, monitors system performance, writes documentation, and joins on-call rotations with nights and weekends. The posting asks for Linux system administration and scripting in bash or Python, and prefers data centre troubleshooting experience, observability tools such as Grafana and Prometheus, Kubernetes administration, and HPC work with GPUs.
Original job posting