The Pulse
Gimlet Labs Raises $300 Million for Multi-Silicon AI Inference
Gimlet Labs has raised a $300 million Series B led by Andreessen Horowitz for its multi-silicon inference cloud. The San Francisco startup says it has billions of dollars in contracted revenue and is scaling toward hundreds of megawatts of

AI.info Team ·
Gimlet Labs announced a $300 million Series B on September 4, putting fresh capital behind its effort to run AI inference across different types of processors instead of building data centers around GPUs alone.
Andreessen Horowitz led the round, which includes Sapphire Ventures, Menlo Ventures, 645 Ventures, Arm, Eclipse, Emergence, Factory, Hudson River Trading, M12, OnePrime Capital, Prosperity7, QuantumLight, Samsung Ventures, Tiger Global Management, Triatomic, Wing Ventures and XTX Markets. A company-distributed financing release puts Gimlet’s valuation at $3 billion and its total funding at $392 million.
Gimlet says it has added billions of dollars in contracted revenue since its $80 million Series A in March and is expanding toward hundreds of megawatts of managed capacity. The company says the new funding will support its inference cloud, infrastructure operations and hiring.
Gimlet’s $300 Million Bet on Mixed Hardware
Gimlet’s central argument is that inference does not place the same demands on hardware from one stage to the next. The prefill phase of a model call is generally compute-intensive, while decoding relies more heavily on memory bandwidth. Agentic systems add further variation by chaining model calls with retrieval, code execution and external tool use.
The company’s software traces and decomposes workloads, then schedules their parts across GPUs, near-memory compute, dataflow architectures and CPUs. Gimlet says its system can split a model across accelerators, separate prefill from decode, and rebalance work when one class of processor reaches capacity.
“We’ve reached a turning point where inference is the dominant AI workload and the demand for tokens is explosive,” said Zain Asgar, co-founder and chief executive of Gimlet Labs. “With Gimlet, our customers are able to serve massive volumes of tokens at very low latency, even as their AI workloads continue to grow across all dimensions.”
The company claims 5x to 10x speedups at the same power footprint, or comparable throughput improvements at the same latency. Its financing announcement describes the product as a multi-silicon cloud built specifically for inference rather than a repurposed training cluster.
Power, Latency and the Limits of GPU-Only Scaling
Gimlet’s funding arrives as AI companies face constraints beyond chip availability. Data-center power, grid connections, cooling and memory supply all affect how quickly providers can add inference capacity. Gimlet cites approximately 18 gigawatts of AI data-center capacity consumed in 2025 and says that figure could triple by 2030.
The company also says monthly token generation has increased sixfold over the past year. Its argument is that operators can serve more requests from a fixed power budget by assigning each phase of a workload to hardware suited to that task, rather than forcing every phase through the same accelerator.
Andreessen Horowitz managing partner Raghu Raghuram, who joins Gimlet’s board, described the investment in the company’s financing release as a response to the physical limits facing AI infrastructure.
“AI demand is growing exponentially, while data centers and silicon can’t keep pace. The answer isn’t just more infrastructure - it’s a better architecture.”
Raghu Raghuram, managing partner, Andreessen Horowitz
Andreessen Horowitz’s investment note says Gimlet’s system can route models and tools onto different processors, separate prefill from decode, and divide a model at the layer or operation level. The firm says the company has already deployed its approach with a frontier laboratory and a hyperscaler, though neither customer is named.
From Series A to Hundreds of Megawatts
Gimlet raised its Series A in March, led by Menlo Ventures, and said at the time that its customer base had tripled to include a top frontier lab and a hyperscaler. The company emerged from stealth in October 2025 with a stated goal of making AI workloads more efficient by widening the pool of usable compute.
The startup now says its architecture supports chips from NVIDIA, AMD, Intel, Arm, Cerebras and d-Matrix. Sapphire Ventures, which participated in the Series B, says Gimlet is scaling heterogeneous infrastructure to hundreds of megawatts and has secured billions of dollars in contracted revenue.
That claim reflects a large gap between a software demonstration and an inference service that operates at industrial scale. Supporting multiple chip architectures requires compilers, runtimes, networking, scheduling, power systems and cooling arrangements that can account for different device requirements. Andreessen Horowitz notes that even water-inlet temperatures can differ between accelerator systems.
The Test Is Utilization, Not the Funding Total
Gimlet is selling an efficiency proposition at a moment when new capacity is expensive and difficult to connect to the grid. Its system has to show that the gains from moving work between processors outweigh the added complexity of operating a mixed fleet.
The company’s own announcement says its workloads can be scheduled according to service-level requirements and available hardware, with unused capacity reassigned when demand shifts. That approach could matter most for agentic workloads, where a single user task may generate dozens of sequential model calls and where latency accumulates across each step.
Gimlet says it is working with frontier labs and other large inference consumers and plans to expand access beyond its current customer base. The immediate expansion target is not a new model or chip, but more managed capacity for the inference cloud it has been building since its launch.