Skip to content
AI.info

jobs

Staff Engineer, Distributed Storage and HPC & AI Infrastructure

About the Role In this role, you will operate, scale, and optimize multi-petabyte storage systems purpose-built for the world’s largest AI training and inference workloads. You’ll manage and scale high-performance parallel filesystems and o

Company
Together AI
Location
Bangalore India
Status
Open
Posted
2026-09-11T07:09:12+00:00

Together AI is hiring a Staff Engineer to run and scale multi-petabyte storage for AI training and inference. The work covers parallel filesystems and object stores (Vast, Weka, Ceph, Lustre), Kubernetes storage operators and controllers, tiered caching, and data paths reaching 10+ GB/s per GPU node. The posting asks for 8+ years in distributed storage, production GPU/HPC cluster experience, Go and Python, and depth in one or more parallel filesystems. Distinctive: open-source contributions and network-layer multi-tenancy tuning.

Original job posting