Run AI models inside your four walls.
A cluster of NVIDIA-class inference nodes, supplied and managed by us, that joins Performance Hub like any other device. The Model Router detects it, your AI Agents route to it, and the latest open-weight models arrive on it automatically. Your prompts never leave the building.

Prompts never leave the building
Inference runs on hardware inside your own walls. What your agents see, ask, and answer stays on your premises, on machines only you occupy.
Zero infrastructure work
We supply the nodes, set them up before they ship, and manage them for their whole life. There are no servers to build, no drivers to fight, and no networking project.
Metered in tokens, not credits
You own the hardware, so on-premise requests don't draw down a balance - they are measured in tokens. Every request still lands in the same logs and analytics as the cloud catalog, with rate limits enforced the same way.
A private AI cloud, without the cloud.
The cluster behaves like any other part of Performance Hub: detected automatically, governed by the access controls you already use, and available to everything in your organisation that speaks to the Model Router.
The Model Router finds it
Pair the nodes like any other Performance Hub device and the Model Router detects them on its own. The models they serve appear in your catalog beside the cloud ones, with no endpoints to configure and no keys to wire in.
Your models, available anywhere
The cluster lives at one site, but every AI Agent and every API key in your organisation can use it as if it were local. Nothing has to be repointed - the same endpoint your software already speaks routes the request to your own hardware.
The catalog keeps itself current
As the labs release new open-weight models, they are provisioned onto your nodes automatically, sized to what the hardware can serve. Turn any model on or off for your facility, exactly as you do in the cloud catalog.
Scaling and throughput, handled
Requests spread across the nodes in the cluster, and adding capacity means adding a node. Load, batching, and placement are managed for you - there is no serving stack to tune.
Access control you already know
Per-agent and per-key model lists, token budgets, and rate limits apply to on-premise models the same way they apply to cloud ones. Decide exactly which agents, keys, and facilities can reach the cluster.
Pin what must stay local
Mark an agent, a key, or a workload so its requests can only ever run on your own hardware - no silent fallback to the cloud. The boundary is enforced by the router, not by convention.
No networking project
Requests are encrypted end to end - whether they arrive from another site through the Model Router or from an agent on the same local network. Nothing is exposed to the public internet, and there is no VPN to run or port-forwarding to set up.
Watch it work
Node health, throughput, and per-model usage sit in the same device management and analytics views as the rest of your fleet. You see what the cluster is doing without logging into a server.
The open-weight catalog, on your hardware.
The leading open-weight models arrive on your nodes automatically as they are released - no downloads, no conversion, no serving stack to configure. How much hardware they need is the only real decision: most of the catalog runs on a single Station-class node, and the very largest flagships run on rack-scale clusters.
Station class: frontier AI without the data centre.
A Station-class node is a single box that lives in your comms room, not a data hall. It serves most of the open-weight catalog at full speed - and for most mid-size and large businesses it is the entire deployment: one node, paired in minutes, at a fraction of the cost of a data-centre build.

GLM-5.2 Distilled
Z.ai
Our rack-scale flagship, distilled to a single-node footprint - the most capable model Station class can serve.
gpt-oss
OpenAI
OpenAI's open-weight models with configurable reasoning effort, trained for tool use.
DeepSeek V4 Flash
DeepSeek
Built for coding and agentic workflows, with 13B active parameters per request.
Qwen3.6
Qwen
A dependable general-purpose family tuned for stability and real-world coding.
Qwen3 Coder
Qwen
A coding specialist built for agents: long-horizon reasoning, tool use, and recovery from failures.
Gemma 4
Google's most capable open family, built from Gemini research, with vision input across sizes.
Llama 4 Scout
Meta
Meta's open multimodal workhorse, with very long context in a single-node footprint.
Magistral
Mistral AI
Mistral's open-weight reasoner, producing long chains of reasoning before it answers.
Devstral 2
Mistral AI
Built for software-engineering agents: explores codebases, edits across files, uses tools.
MiniMax M2
MiniMax
A coding and agent workhorse with 10B active parameters, fast for its size.
Phi-4
Microsoft
Small models with outsized reasoning, trained on carefully curated synthetic data.
Nemotron 3
NVIDIA
NVIDIA's open hybrid family, tuned for efficient multi-agent workloads on its own silicon.
LFM2
Liquid AI
Hybrid models designed for the edge: strong quality at very low latency and memory.
Starts with a single node, and grows by adding another - not by building a data centre.
Enterprise rack scale, for the largest frontier models.
The very largest flagships need a Rubin and Grace Blackwell class rack rather than a single node. It's an enterprise deployment we design and run together - our engineers and your team, from scoping through go-live and beyond.

GLM-5.2
Z.ai
Z.ai's flagship for long-horizon coding and agentic work, served at full quality across a rack.
Kimi K3
Moonshot AI
Moonshot's trillion-parameter-class flagship for long-horizon agentic work, with native vision.
Llama 4 Maverick
Meta
Meta's flagship open multimodal model, with the memory headroom for full context and concurrency.
For organisations that need the absolute frontier running on their own floor.
A sample of the catalog. Models are provisioned as the labs release them, sized to what your nodes can serve, and each one can be enabled or disabled per facility.
Runs at one site. Serves your whole organisation.
The cluster is physical hardware at one of your sites, but everything in your organisation that uses the Model Router can reach it - as if the models were local to each of them.
In the building
- A cluster of NVIDIA-class inference nodes, racked at your site and managed by us.
- Models execute entirely on your hardware. Prompts and responses are processed on machines you house.
- Capacity grows by adding nodes - the cluster absorbs them without reconfiguration.
- Node health and throughput report into device management beside your other hardware.
Across the organisation
- AI Agents at any facility route to the cluster through the Model Router, automatically.
- Your own software reaches it through the same OpenAI-compatible endpoint and the keys it already holds.
- One meter and one set of logs - local requests counted in tokens, side by side with cloud usage.
- Access control decides which agents, keys, and facilities can use the cluster at all.
Cloud, on-premise, or both.
Every model in the catalog is reached the same way - the only decision is where it runs. Route to the cloud, keep everything local, or mix the two per workload.
01
Cloud catalog
Hundreds of models from the leading providers, no hardware to own, one prepaid balance. The right default for most workloads, and it is already there the moment you enable the Model Router.
02
On-premise cluster
The same catalog experience on your own NVIDIA-class nodes. Prompts never leave the building, and because you own the inference, requests are metered in tokens rather than credits.
03
Local first, cloud fallback
A routing rule per agent, key, or workload: sensitive work pinned to your own hardware, overflow and the models your nodes don't carry served from the cloud - with one set of logs across both.
From a conversation to local inference.
01
Talk to sales
We start by scoping your workloads, sites, and throughput - and whether a Station-class node or a rack-scale cluster is the right fit.
02
We size and supply the cluster
NVIDIA-class nodes matched to your workload, set up before they ship. From compact units to rack-scale hardware, mixed as needed.
03
Pair it like any device
The nodes join your facility the way an Edge Processor does: claimed in minutes, connecting outbound only, detected by the Model Router on their own.
04
Your agents route locally
Requests start landing on your own hardware, visible in the same logs and analytics as everything else. Nothing about how you use your agents changes.
The fine print, up front
Who it's for
Station-class deployments suit mid-size and large businesses - one node, scoped and deployed with our team rather than self-serve. Rack-scale clusters are enterprise builds that we plan and run together with your team. Either way, it starts with a conversation with sales.
NVIDIA-class nodes
Deployments come in two classes: Station-class nodes that sit in an office and serve most of the catalog, and Rubin and Grace Blackwell class racks for the trillion-parameter flagships. A deployment can mix both, and capacity is added by adding nodes.
The security posture
Encrypted end to end on every path: remote requests routed through the Model Router and local requests on your own network alike. Outbound-only connections, nodes never exposed to the public internet - the same posture as every other device on the platform.
The data boundary
Inference happens on your premises. What travels to the platform is the meter - model, tokens, latency, and status for each request. Prompt and response content is never stored, locally or in the cloud.
Kept current
Models and node software are updated automatically as part of the managed service, on maintenance windows agreed with you. You are never the one applying patches.