Deploy

Hardware that ships withthe models already loaded.

Five inference-ready configurations, from a desktop box to an eight-GPU rack. Each arrives with an open LLM, real-time voice, and retrieval pre-loaded and benchmarked, so it runs the moment you plug it in. You own the hardware and the stack.

Why on-prem

Some data cannot go to the cloud.

01

Data residency

Regulated data that legally cannot leave your walls.

02

Latency

Inference next to the data, no round trip to a cloud region.

03

Cost at scale

Steady high-volume inference is cheaper on hardware you own.

04

Control

Your keys, your kill switch, no third party in the path.

The configs

Five boxes. Pick your intelligence.

Every recommended stack clears 50 tokens a second with real-time voice on the box, and fits its VRAM with room for retrieval. No stack you can configure misses that bar. Prices are approximate and GST-inclusive; the premium buys a benchmarked performance guarantee.

Entry

Desktop · 32GB
~₹6.5–8.5 Lincl. GST

Performance benchmarked and guaranteed.

A team piloting its first agent on hardware it owns.

Runsgpt-oss-20b · 519 tok/s
VoiceParakeet + Kokoro
RetrievalArctic-Embed L v2
Hardware1× RTX 5090 · 32GB

Studio

Workstation · 48GB
~₹9.5–12.5 Lincl. GST

Performance benchmarked and guaranteed.

A department running several agents, with the premium voice stack.

Runsgpt-oss-20b · 278 tok/s
VoiceCanary + XTTS-v2
RetrievalArctic-Embed L v2
Hardware1× RTX 6000 Ada · 48GB

Pro

Single-GPU · 96GB
~₹20–24 Lincl. GST

Performance benchmarked and guaranteed.

A frontier-class open model on a single accelerator.

Runsgpt-oss-120b · 366 tok/s
VoiceCanary + XTTS-v2
RetrievalArctic-Embed L v2
Hardware1× RTX PRO 6000 Blackwell · 96GB

Max

Dual-GPU · 192GB
~₹40–48 Lincl. GST

Performance benchmarked and guaranteed.

Two cards, real headroom, and the sovereign 105B option.

Runsgpt-oss-120b · 622 tok/s
VoiceCanary + XTTS-v2
RetrievalArctic-Embed L v2
Hardware2× RTX PRO 6000 Blackwell · 192GB

Rack

8-GPU cluster
~₹4.3–5.8 Crincl. GST

Performance benchmarked and guaranteed.

Hundreds of concurrent sessions, or the largest open models.

Runsgpt-oss-120b · 6275 tok/s
VoiceCanary + XTTS-v2
RetrievalArctic-Embed L v2
Hardware8× H200 · 1128GB

How it works

From configuration to running hardware.

01

You configure it

Pick the box, the models, and the customizations. We confirm the stack clears its performance target.

02

We build and benchmark it

We assemble the hardware, pre-load your models, and benchmark the stack so the guarantee holds.

03

It ships ready, you own it

It arrives inference-ready where your data lives. The hardware and the stack are yours.

Questions

What teams ask before they deploy.

What is Deploy?

Inference hardware that ships with the models already loaded. Five configurations, from a desktop box to an eight-GPU rack, each pre-loaded with an open LLM, real-time voice, and retrieval, and benchmarked so it runs the moment you plug it in.

Why run AI on-prem instead of in the cloud?

Four reasons: data residency for regulated data that cannot leave your walls, lower latency with inference next to the data, lower cost for steady high-volume inference, and full control with your keys and your kill switch.

What performance do the boxes guarantee?

Every recommended stack clears 50 tokens a second with real-time voice on the box, and fits its VRAM with room for retrieval. The premium buys a benchmarked performance guarantee.

Do we own the hardware?

Yes. The hardware and the stack are yours. It arrives inference-ready where your data lives.

How does it work end to end?

You configure the box, the models, and the customizations; we assemble and benchmark it so the guarantee holds; then it ships inference-ready and you own it.

Tell us where your data has to live.

Book a 30-minute callNo deck. Straight to your problem.
Or write to us with the use case →