tools
Not Diamond Model Router
Not Diamond routes AI requests to suitable models to balance quality, cost, latency, and coding-agent performance.

Not Diamond analyzes AI requests and recommends which model to use from a selected pool. It supports pre-trained and custom routers, cost or latency tradeoffs, prompt optimization, and integrations through Python, TypeScript, and REST APIs.
The product is aimed at developers, engineering teams, and coding-agent operators. It is not an AI gateway and does not execute model requests itself; teams continue using their existing gateway or provider. Current pricing is usage-based, with enterprise features and discounts available separately.
Features
- Routes each request to a suitable model from a configured model pool
- Supports quality, cost, latency, and cost-quality routing tradeoffs
- Trains custom routers on evaluation data and response scores
- Provides prompt optimization for supported language models
- Returns a recommended model and session ID through its API
- Supports Python, TypeScript, and REST API integrations
- Supports custom models and arbitrary inference endpoints
- Offers privacy-preserving routing for coding-agent workloads
Use cases
- Reduce coding-agent inference costs by selecting models per session step
- Route application queries across models from multiple providers
- Train a router for domain-specific evaluation criteria
- Balance response quality against model cost or latency
- Optimize prompts for different target models
- Track routing recommendations and session outcomes
Pros
Cons
Pricing
- Starting price
- $0.05 per million tokens routed
- Pricing checked
- 2026-09-19
Pay-as-you-go
$0.05 per million tokens routed
- Multi-harness support
- Multi-provider support
- Savings and usage dashboard
- Continuous learning from harness usage
Enterprise
Custom
- Volume-based discounts
- Organization-wide analytics
- Advanced security and admin controls
- SSO with SAML
- Privacy-preserving deployment