jobs
Member of Technical Staff, Performance Optimization
About Us:Fireworks is the platform for specialized intelligence, enabling companies to build, train, and serve AI models tailored to their own data, workflows, and products. Founded by the team behind PyTorch and backed by AMD, Atreides, Be
- Company
- Fireworks AI
- Location
- San Mateo
- Status
- Open
- Posted
- 2025-05-06T20:04:42.424+00:00
Fireworks AI builds and serves open AI models and is hiring an engineer for performance work across its model stack. The job covers GPU kernel and system optimization for training and inference: profiling latency, throughput and memory, writing CUDA or Triton kernels, tuning mixed precision, quantization and graph optimization, and scaling workloads across multiple GPUs and nodes. Candidates need five or more years in performance optimization or high-performance computing, CUDA or ROCm with profiling tools, PyTorch, and distributed debugging experience. The role is hybrid in San Mateo.
Original job posting