Private AI GPU Cloud

Run your AI models on a private GPU cloud dedicated to your organisation.

AI in production needs GPU capacity you can plan against, and data that stays where you decide. We do not just supply GPUs: we design the GPU, storage and network architecture around your workload, install it, deploy your models on it, and operate it as one platform. It runs in our own cage space in Zurich, Frankfurt or Istanbul, or in your own data centre. Your data, your prompts and your model weights stay inside the boundary you choose.

It is built for production LLM, ML and image-processing workloads. That includes fintech, trading and life sciences, where data location, confidentiality and operational control matter. The cluster serves your organisation. Your teams share it under agreed quotas.

Tell us what you are training or serving, and on what data.

Boundary

What stays inside your boundary

Your data covers more than the training set. All of the following stay inside the boundary you choose, in the locations you choose.

This matters when your organisation has to state where processing happened, not estimate it. With a hosted service, retention, subprocessors and processing location follow that provider's terms. For many workloads those terms are acceptable. This platform is for the ones where they are not.

01

Training and fine-tuning data.

02

Model weights you have fine-tuned.

A fine-tuned model is derived from your data. Treat it with the same care as the source.

03

Prompts and inputs.

At inference time these can carry customer records, transaction details or clinical information.

04

Outputs.

What the model returns is derived from your data. Where it is stored or logged, that copy needs the same treatment.

05

Embeddings and vector stores.

An embedding is derived from the document it was computed from. Embeddings look like numbers, so they are easy to overlook.

06

Logs and traces.

Depending on how the application is configured, traces and logs can contain prompts and outputs.

Model independence

You are not tied to one model

The platform is built around a class of workload, not around one model. Nothing in the storage layout, the serving layer or the operational tooling assumes a specific model. That gives you four things.

Routing layerFig.
Requests
Router
Larger model
Smaller model
Same hardwareReplaceable
01

You can benchmark several candidate models on the same hardware before you choose.

02

You can run more than one at a time behind a routing layer. A larger model for the requests that need it, a smaller one for the rest.

03

You can replace a model later without rebuilding the platform underneath it.

04

The capacity is sized for a class of workload, so it stays useful when the model changes.

We do not sell a model. Our recommendation is not tied to our own revenue.

Benchmarking

We measure your workload, not ours

We publish no performance figures for this platform. A number measured on someone else's model, data and concurrency tells you nothing about yours.

Benchmarking is part of the platform. We measure candidate models on the hardware you would use, with your data, against your evaluation set. The results decide which model you deploy.

What we measure

01

Quality on your evaluation set.

02

Behaviour as concurrent load rises.

03

Memory use at realistic context lengths.

04

Throughput per unit of hardware.

05

How the system behaves at its limit.

06

Cost per unit of work, for each candidate.

What you get

A comparison across candidate models on identical hardware. It shows the capacity each one implies, with a recommendation and the reasoning behind it.

Commercial model

What it costs

Setup feeIn most cases none
Hardware investmentIn most cases none
BillingMonthly
AssessmentBefore any commitment
From inventory

In most cases there is no setup fee and no hardware investment. Where suitable capacity exists in our inventory, it is provided from there, and you pay monthly.

Dedicated hardware

Accelerators are expensive and demand is high, so the specification your workload needs may not be in our inventory. Where dedicated hardware is required, that is established during design, before you commit. The investment model is then agreed with you.

Assessment first

The workload assessment comes before any commitment. It can show that a serving platform is enough where a training cluster was assumed. That changes the cost.

Renting today

If you rent GPU capacity today, we can run a cloud cost analysis on those lines.

Workload fit

Which workloads belong here

Good fit

GPU demand that runs continuously rather than in short bursts.

Training or fine-tuning on data that has to stay inside your organisation.

Inference on prompts that carry customer, clinical or commercial records.

Several teams that need predictable, allocated capacity.

Production serving where latency and capacity have to be planned.

Retrieval systems built on internal documents, source code or pricing.

Regulated work with a stated processing location.

Not a fit

A large cluster for a few days a quarter, and nothing in between.

Work that is not yet defined, where training and serving volumes are unknown.

A product built on one provider's hosted model. A private platform cannot run it.

A requirement to be on the newest accelerator within weeks of its release.

A problem that is data quality or task definition. More GPUs produce the same results faster.

Next step

Start with the workload

Tell us what you are training or serving, on what data, and at what concurrency. Tell us what is in the way today.

You get an architecture conversation with an engineer and a view of what the platform should look like. If renting capacity suits you better, we say so.

Before any commitment
Request a GPU workload assessment

The assessment also answers whether you need a cluster at all. We reply within 2 business days.

How it is built and runBelow

Delivery scope

What we deliver

Design

Workload analysis.

What you run, in what proportion, on what data.

Architecture design.

GPU, CPU, memory, storage and network, designed together.

Hardware selection.

Specified against the workload. Where new hardware is needed, we source it.

Build

Installation.

Racking, power, cooling, cabling, firmware and drivers.

Kubernetes environment.

Cluster, GPU scheduling, namespaces, quotas, storage classes and ingress.

Model deployment.

Your models brought up, with serving, routing and autoscaling around them.

Run

Benchmarking.

Measured on your hardware, with your data.

Load balancing.

Request distribution designed around how inference behaves.

Monitoring and operations.

Utilisation, memory, queues, job outcomes and the infrastructure underneath.

Two designs

Training and inference need different designs

Training

Training takes all the capacity it is given, for as long as the job runs. It can wait in a queue.

Inference

Inference cannot. Work arrives when users send it. The number that matters is what happens to the slowest requests when the system is busy.

We design for both, and we keep them apart. Serving capacity is reserved. Batch training cannot take it. Where the workloads are large enough, they run in separate node pools.

Multi-tenancy

Sharing the cluster between your teams

Namespaces

Each team gets its own namespace, with quotas covering GPUs, CPU, memory and storage. A quota sets what a team may consume, so no team can take the whole cluster by accident.

Queues

Batch work queues. Priority decides the order. Fair-share behaviour stops one team's backlog from holding the cluster.

Reserved serving

Capacity for customer-facing inference is set aside for it. Batch work runs at lower priority and can be preempted. Where the workloads are large enough, serving and batch run in separate node pools.

Visibility

Per-team GPU hours, utilisation, queue times and job outcomes are visible to you. That makes idle capacity easy to find and release.

Storage and network

Fast GPUs need fast data

A GPU that waits for data is capacity you pay for and do not use. So storage and network are designed together with the GPUs, not added afterwards.

01

Fast storage for training reads.

02

A local cache on each GPU node.

03

Separate space for checkpoints.

04

A store for model versions.

GPU-to-GPU traffic and storage traffic use separate network paths, so one does not slow the other when both are busy.

Responsibilities

What we run, what you keep

We run

The hardware, the Kubernetes environment and the model deployment layer. Uptime, performance, security, capacity and network are our job. Monitoring and maintenance do not stop.

You keep

Your models, your data and your data science. They stay with you or your partners.

Location and access

Where your data lives, and who can see it

Where

Zurich, Frankfurt, Istanbul or your own data centre. You choose before we build, and it is written into the agreement.

Who sees it

Nobody on our team. Running the infrastructure does not require reading your data or your application content, so we do not.

Admin access

Named people only, with MFA, VPN and logging. Where access must be limited to one jurisdiction, we design for that and put it in the contract.

If you ask us to run an application as well, the access model is different and is set out in that agreement.

Questions

Questions we are asked

  • 01Can the platform run in our own data centre?+
    Yes. The design work, the Kubernetes environment, the model deployment and the operations are the same. What changes is where the hardware sits and who owns it. If it runs in your facility, we need an early assessment of rack power, cooling capacity and floor loading. We check those before the design is fixed, because they can change it.
  • 02Can a long training run survive a GPU failure?+
    Yes. The platform is designed for it. Where the job checkpoints, it resumes from the last checkpoint, and the work lost is whatever happened after that checkpoint. We set the checkpoint interval against the cost of lost work, and we size storage for the checkpoint write. The failed node is drained from the pool, the fault is diagnosed, and hardware is replaced under our maintenance. We plan capacity headroom, so losing a node reduces throughput instead of blocking the queue.
  • 03Can we keep our own models and data-science partners?+
    Yes. The models and the data science stay with you or your partners. We build and operate the platform underneath: we deploy your models, benchmark them, and configure the serving layer around them.
  • 04Can our data scientists run their own experiments?+
    Yes. Each team gets a Kubernetes environment with its own namespace, quota and storage, a way to submit and monitor jobs, and a place to deploy and serve models. They see their own queue position and consumption, so they do not need a ticket to run an experiment. Where your teams have existing tools and workflows, we build around them.
  • 05Can it connect to our existing systems and data?+
    Yes. Training data lives in the systems that produced it, so this connectivity is part of the design. We design the connectivity, the data movement path and the access controls as part of the architecture. That includes where a staging copy is needed and how long it stays. Where it runs alongside a private cloud we already operate for you, it uses the same network design and identity model.

We built our first private cloud infrastructure in 2020. Our team also builds and operates a live crypto exchange platform, open to you on request. It is a different product, built and run by our own engineers.

We do not publish client names. Ask, and we will share relevant references under NDA.

Tell us what you are training or serving. If renting capacity suits you better, we say so. We reply within 2 business days.