Skip to main content

  1. Agentycs
  2. Deployment
  3. On-premises

On-premises

The whole platform in your own data centre, on hardware you own.


What it is #

On-premises means the platform runs on machines you own, in a room you control. Apex AI servers are built for exactly that: purpose-designed for dense inference rather than adapted from a general-purpose server, and vertically integrated with the platform so placement, numeric format and memory budgeting are decided for you rather than tuned by hand. If you would rather use your own hardware or a virtualised estate, that is supported too.

Everything above the metal is the same platform as everywhere else, reconciled by the same operator from the same declarations. That matters most when the site is quiet: an on-premises instance does not need a human to keep it converged, and it does not need the internet to keep serving what it already has.

Who it is for

  • Data cannot leave your premises, by policy or by law.
  • You have, or want, a data centre footprint and the people to run it.
  • You want inference capacity you own outright rather than rent by the hour.
  • You would rather trade some elasticity for complete physical control.

An on-premises instance does not need a human to keep it converged, and it does not need the internet to keep serving.

What you get #

The whole platform runs here. These are the parts this pattern changes the shape of.

01

Purpose-built accelerators #

Apex servers are built to run the largest open models in the fewest racks, and are integrated with the platform's scheduler rather than bolted underneath it. What a given model costs you in power, heat and space is a sizing conversation against your models, not a headline figure.

02

One control plane #

The same operator and the same custom resources as every other deployment, so what you validated elsewhere behaves the same here.

03

Sovereign model supply #

Import open models once into the private hub. Weights are content-addressed, signed and seeded to workers, so serving never reaches out to a model host at runtime.

04

Storage that grows in place #

S3-compatible object storage starts single-zone in the cluster and expands to multiple zones and sites without a migration.

The operating model #

Nobody should discover who owns a failure while it is happening. This split is explicit before anything is signed, and it is the part of a deployment decision that outlives the architecture.

You operate
The facility, power, cooling, physical security, the network and the hardware lifecycle.
The operator handles
Cluster and platform reconciliation, tenant composition, model deployment and rollout, backups and restore.
We can operate alongside
Platform on-call, upgrades and diagnostics through an access path you define, including one that requires an escort.
You keep
Physical custody, key custody, identity, tenant policy and every audit record.