jobs
Member of Technical Staff (AI Inference Engineer)
We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets. Our stack is Rust, Python, CUDA, and CuTe DSL - and we need another engineer to join
- Company
- Perplexity
- Location
- San Francisco; Palo Alto; New York City
- Status
- Open
- Posted
- 2026-04-13T19:39:49.104+00:00
Perplexity is hiring an inference engineer for the team running the engine behind every query, deploying dozens of model architectures under latency and cost budgets. Work includes bringing up transformer retrieval, text-generation and multimodal models, porting in-house CUDA kernels to CuTe DSL for GB200 and future racks, developing a Rust serving runtime, profiling bottlenecks, and reliability tooling. The posting asks for 3+ years of software engineering experience, GPU programming depth, and familiarity with LLM architectures, inference optimization, Rust, Python and CUDA. On-site across three US cities.
Original job posting