Deploy
Hardware that ships withthe models already loaded.
Five inference-ready configurations, from a desktop box to an eight-GPU rack. Each arrives with an open LLM, real-time voice, and retrieval pre-loaded and benchmarked, so it runs the moment you plug it in. You own the hardware and the stack.
Why on-prem
Some data cannot go to the cloud.
Data residency
Regulated data that legally cannot leave your walls.
Latency
Inference next to the data, no round trip to a cloud region.
Cost at scale
Steady high-volume inference is cheaper on hardware you own.
Control
Your keys, your kill switch, no third party in the path.
The configs
Five boxes. Pick your intelligence.
Every recommended stack clears 50 tokens a second with real-time voice on the box, and fits its VRAM with room for retrieval. No stack you can configure misses that bar. Prices are approximate and GST-inclusive; the premium buys a benchmarked performance guarantee.
Entry
Desktop · 32GBPerformance benchmarked and guaranteed.
A team piloting its first agent on hardware it owns.
Studio
Workstation · 48GBPerformance benchmarked and guaranteed.
A department running several agents, with the premium voice stack.
Pro
Single-GPU · 96GBPerformance benchmarked and guaranteed.
A frontier-class open model on a single accelerator.
Max
Dual-GPU · 192GBPerformance benchmarked and guaranteed.
Two cards, real headroom, and the sovereign 105B option.
Rack
8-GPU clusterPerformance benchmarked and guaranteed.
Hundreds of concurrent sessions, or the largest open models.
How it works
From configuration to running hardware.
You configure it
Pick the box, the models, and the customizations. We confirm the stack clears its performance target.
We build and benchmark it
We assemble the hardware, pre-load your models, and benchmark the stack so the guarantee holds.
It ships ready, you own it
It arrives inference-ready where your data lives. The hardware and the stack are yours.
Questions
What teams ask before they deploy.
What is Deploy?
Inference hardware that ships with the models already loaded. Five configurations, from a desktop box to an eight-GPU rack, each pre-loaded with an open LLM, real-time voice, and retrieval, and benchmarked so it runs the moment you plug it in.
Why run AI on-prem instead of in the cloud?
Four reasons: data residency for regulated data that cannot leave your walls, lower latency with inference next to the data, lower cost for steady high-volume inference, and full control with your keys and your kill switch.
What performance do the boxes guarantee?
Every recommended stack clears 50 tokens a second with real-time voice on the box, and fits its VRAM with room for retrieval. The premium buys a benchmarked performance guarantee.
Do we own the hardware?
Yes. The hardware and the stack are yours. It arrives inference-ready where your data lives.
How does it work end to end?
You configure the box, the models, and the customizations; we assemble and benchmark it so the guarantee holds; then it ships inference-ready and you own it.