jobs
Member of Technical Staff, Model Efficiency
Who are we?Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.We’re training and deploying frontier models f
- Company
- Cohere
- Location
- New York; San Francisco; Toronto; Montreal
- Status
- Open
- Posted
- 2025-11-07T23:19:41.712+00:00
Cohere's Model Efficiency team is hiring an engineer to improve LLM inference in production. The role involves working across the inference stack, diagnosing bottlenecks, measuring experiments, and shipping optimizations with modeling and systems teams. The posting asks for 5+ years of high-performance production code, strong C++ or Python, LLM inference ecosystem experience such as vLLM or SGLang, and bottleneck diagnosis. GPU/CUDA, low-level systems, MoE, speculative decoding, KV-cache, and distributed-systems experience are pluses. Work is remote-friendly, with the team concentrated in EST/PST.
Original job posting