Skip to content
AI.info

jobs

Member of Technical Staff, Model Efficiency

Who are we?Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.We’re training and deploying frontier models f

Company
Cohere
Location
New York; San Francisco; Toronto; Montreal
Status
Open
Posted
2025-11-07T23:19:41.712+00:00

Cohere's Model Efficiency team is hiring an engineer to improve LLM inference in production. The role involves working across the inference stack, diagnosing bottlenecks, measuring experiments, and shipping optimizations with modeling and systems teams. The posting asks for 5+ years of high-performance production code, strong C++ or Python, LLM inference ecosystem experience such as vLLM or SGLang, and bottleneck diagnosis. GPU/CUDA, low-level systems, MoE, speculative decoding, KV-cache, and distributed-systems experience are pluses. Work is remote-friendly, with the team concentrated in EST/PST.

Original job posting