Skip to content
AI.info

jobs

Member of Technical Staff (AI Inference Engineer)

We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets. Our stack is Rust, Python, CUDA, and CuTe DSL - and we need another engineer to join

Company
Perplexity
Location
San Francisco; Palo Alto; New York City
Status
Open
Posted
2026-04-13T19:39:49.104+00:00

Perplexity is hiring an inference engineer for the team running the engine behind every query, deploying dozens of model architectures under latency and cost budgets. Work includes bringing up transformer retrieval, text-generation and multimodal models, porting in-house CUDA kernels to CuTe DSL for GB200 and future racks, developing a Rust serving runtime, profiling bottlenecks, and reliability tooling. The posting asks for 3+ years of software engineering experience, GPU programming depth, and familiarity with LLM architectures, inference optimization, Rust, Python and CUDA. On-site across three US cities.

Original job posting