Skip to content
AI.info

The Pulse

Cirrascale launches production platform for multi-accelerator inference

Cirrascale Cloud Services has released a production inference platform that routes models across NVIDIA, AMD, Qualcomm, and Tenstorrent accelerators. The platform adds private model hosting, fine-tuning, governance controls, spend managemen

Cirrascale launches production platform for multi-accelerator inference

AI.info Team ·

“Customers get working applications, spend controls, and agent guardrails on day one, running privately on the accelerator that makes the most sense for their workload.”

Alex Nataros, CTO, Cirrascale Cloud Services

Cirrascale Cloud Services has released the production version of its Cirrascale Inference Platform, a private enterprise AI stack designed to route workloads across accelerators from NVIDIA, AMD, Qualcomm, and Tenstorrent. The company announced the release on September 15, 2026, at the AI Infra Summit in Santa Clara, California.

The platform combines model deployment, routing, accelerator selection, security controls, and application tooling in a single web console. Cirrascale says customers can run open-source models, private models, and closed model ecosystems, including Google Gemini delivered on-premises through Google Distributed Cloud and operated by Cirrascale.

One console for models, applications, and controls

Cirrascale is targeting search systems, chatbots, workplace copilots, agentic services, coding assistants, document intelligence, and video generation. Its platform connects those pipelines to existing enterprise data and services across on-premises environments and hyperscaler infrastructure.

The company packages a private chat application connected to an organization’s knowledge base, controls for managing AI consumption by team, and policy guardrails for agentic workloads. Cirrascale says the deployment model supports requirements aligned with HIPAA, SOC 2, and FedRAMP where applicable.

“Enterprises do not struggle to stand up a model endpoint anymore,” Nataros said. “They struggle with everything around it: governance, cost control, and getting a secure application in front of employees.”

Routing across four accelerator vendors

The platform’s model and hardware selection layer is intended to separate model choice from a fixed infrastructure commitment. Cirrascale says it can automatically route each request to the appropriate model and select an available accelerator without requiring application code changes when hardware changes.

The supported accelerator vendors named in the release are NVIDIA, AMD, Qualcomm, and Tenstorrent. The company says the optimization layer is designed to increase throughput, keep latency stable as demand rises, and run larger models on existing systems while producing more predictable cost per token.

Cirrascale does not publish benchmark results, pricing, or a detailed list of accelerator models in the announcement. The release instead positions hardware choice as an operational feature for organizations that want to compare model and accelerator combinations within one service.

Private fine-tuning and enterprise deployment

Customers can fine-tune models on private data without moving that data outside their environment, according to Cirrascale. The resulting models can be deployed through the same platform and endpoints used for other workloads.

The company also says the production service includes multi-region operation, allowing enterprise pipelines to connect to data and services in the locations where they already run. Cirrascale presents that arrangement as an alternative to building separate inference, governance, and application layers around raw accelerator capacity.

Cirrascale’s bet on infrastructure choice

Dave Driggers, Cirrascale’s CEO and co-founder, framed the product as a response to the way cloud AI services typically bind models to a provider’s own hardware.

“They pick the model and we put it on the best hardware for the job, in a private environment, at a price their CFO can plan around,” Driggers said.

The Cirrascale Inference Platform is available across the company’s U.S. and international regions. Its release marks a move from the platform’s earlier preview positioning to a production service centered on multi-vendor accelerator selection, private deployment, and the software controls needed to operate enterprise inference.

Source

Cirrascale Cloud Services

Explore

More articles