The machines underneath, built for the models on top.
Apex is the entire Agentycs hardware range, from a 6 litre micro server you can carry in a backpack to racked datacentre systems, every one of them running the same platform and every one of them yours.
One range, one platform, from a backpack to a hall #
Frontier models are expensive to serve not because the mathematics is hard, but because most infrastructure is assembled from parts that were never designed for each other and then rented by the hour from somebody else. Apex is the opposite: servers designed for the workload, sold rather than metered, and running the platform that schedules them.
The range runs the whole way. At the micro end, Agentycs-in-a-Box is a 6 litre ultra-small-form-factor server, backpack-sized, battery-capable and drawing 400W, holding a full-fat 96GB NVIDIA GPU and full server capability, and running a proper platform instance with Atom and Anima on board. At the other end are racked clusters, single-site or geodistributed across datacentres. The instance model is the same at both ends, so the shape of a deployment is a configuration rather than a different product.
A box you can carry onto an aircraft and a hall full of racks run the same platform, with the same identity model and the same operational surface.
What Apex provides #
Four ideas carry the product. Everything else is detail.
Servers designed for the workload #
The servers and the inference plane are designed together, so accelerator architecture, numeric format and model placement line up without tuning.
- Purpose-built for inference and training
- Mixed accelerator fleets welcome
- Or bring your own hardware
A range, not a single box or a single rack #
The same platform instance model covers Agentycs-in-a-Box at the micro end, compact installations at remote sites, racked clusters in a datacentre, and facilities with no route to the internet at all.
- 6 litres, backpack-sized, battery-capable, 400W
- Racked, single-site or geodistributed
- Fully disconnected installations
Declared, not configured #
Instances, tenants, applications, databases, models, storage and routes are all reconciled to their declared state by one operator and its sub-operators.
- Kubernetes-native and operator-managed
- Version-controlled resource definitions
- A second site is configuration, not a project
Storage and secrets you own #
S3-compatible object storage starts in-cluster and expands in place to multi-zone capacity, and every credential passes through one audited doorway.
- S3-compatible, standard tooling unchanged
- Encrypted store of record with automatic unseal
- Or point at your own bucket and your own vault
In more detail #
The full capability surface, grouped by the job it does.
AI servers
Hardware designed for inference and training, integrated with the software that schedules it.
- Designed for the workload
- Purpose-built servers for inference and training rather than general-purpose machines adapted to it, which is where the efficiency comes from. What that is worth in your environment depends on the models and the load, so we size against those rather than publish a headline figure.
- Vertically integrated
- The servers and the inference plane are designed together, so accelerator architecture, numeric format and model placement all line up without tuning.
- Mixed fleets welcome
- Accelerators are discovered and classified automatically, so several generations and vendors can serve side by side.
- Your hardware, if you prefer
- Alternative hardware and virtualised cloud infrastructure are supported where that suits your procurement or your timeline better.
The range, end to end
From the smallest thing that can run the whole platform to a geodistributed estate, under one instance model.
- Agentycs-in-a-Box, the micro end
- A 6 litre ultra-small-form-factor server: backpack-sized, battery-capable, drawing 400W, with a full-fat 96GB NVIDIA GPU and full server capability, running a proper platform instance including Atom and Anima. It is a whole platform you can carry, not a cut-down appliance.
- Portable and edge
- Compact installations for remote sites, vehicles and vessels, serving locally and syncing back when a link is available.
- Datacentre
- Racked capacity in your own or a colocated facility, as one cluster or several geodistributed for scale and resilience.
- Disconnected, and combinations
- Fully air-gapped installations hold their own model weights, artefacts, registries and secrets, and any of these can be combined under one platform instance with central governance and local autonomy at each site.
Cluster and lifecycle
Kubernetes-native, operator-managed, declared rather than configured.
- Hardened Kubernetes
- A secure open-source Kubernetes distribution with enterprise support available, or the distribution you already standardise on.
- One platform operator
- A single operator manages every core resource and supervises the sub-operators that own each subsystem's full lifecycle.
- Declarative resources
- Platform instances, tenants, applications, databases, models, storage and routes are all custom resources reconciled to their declared state.
- Geodistribution
- A platform instance spans datacentres for scalability and resilience, presenting one surface across all of them.
Sovereign object storage
S3-compatible storage for files, datasets, model weights and search indexes, owned by your platform.
- S3-compatible
- Standard tooling and clients work unchanged against the platform's own object store.
- Grows in place
- Starts single-zone in-cluster and expands to multi-zone, geo-distributed storage across your sites without a migration.
- Least-privilege access
- Buckets, access keys and grants are reconciled from declarative resources, so every consumer gets only the access it needs.
- Bring your own
- A tenant may instead point at their own external S3-compatible storage, so where data rests stays their decision.
Secrets and credentials
One audited doorway to every credential the platform and its tenants use.
- Encrypted store of record
- An operator-managed vault holds platform and tenant secrets, with automatic unseal so recovery does not wait on a person with a key.
- Brokered access
- Every request routes through one gateway that enforces per-tenant access and rate limits, caches safely, and forwards writes to the secret's home region.
- Your own vault
- A tenant can keep secrets in their own external vault, reached through the same brokered path.
- Never in an image or a log
- Credentials are referenced rather than embedded, and redacted out of status, events, metrics and traces.
Durability and resilience
Assume hardware fails, sites go offline and links drop, then keep serving.
- Off-site backup
- Scheduled backups and multi-zone replication for durability and recovery.
- Spread serving capacity
- Platform-critical services run with multiple replicas spread across hosts, so losing a node cannot empty a service.
- Local autonomy
- A site keeps serving what it has loaded when the core is unreachable, then reconciles when the link returns.
- Observed health
- Cluster, accelerator, storage and network health all report into the platform's shared metrics, dashboards and alerts.
Where Apex goes #
Three situations the range is built for.
Bring inference in-house
Replace metered model APIs with owned capacity, sized to your actual load, at a cost per token you control.
Regulated estates
Keep data, weights and audit inside a boundary you can point at on a map, with no external service in the request path.
Remote and mobile operations
Carry Agentycs-in-a-Box to where the work is and run the whole platform on 400W of battery or a generator, operating on its own when connectivity is poor or absent.
Read moreHow you operate it #
Everything is a declared resource with an operational surface in front of it.
- Declarative resources
- S3-compatible storage API
- Fleet administration
- Metrics, dashboards and alerts
What each interface gives you
- Declarative resources
- Platform instances, clusters, tenants, databases, storage, models and routes, all reconciled from version-controlled definitions.
- S3-compatible storage API
- Standard object-storage clients and tooling against the platform's own sovereign store.
- Fleet administration
- A first-party console for platform-wide operations, tenant configuration and capability management.
- Metrics, dashboards and alerts
- Cluster, accelerator, storage and network health in the same observability fabric as every other product.
Size it against your workload
Tell us the models you want to serve, the load you expect and the constraints on where it can run. We will come back with a configuration.