GPU SERVER

NVIDIA GPUs, rented or purpose-built.

GPU Cloud for hourly on-demand capacity, and Cluster Build for dedicated clusters we design, deploy, and operate. Both run on B300, B200, GB300, and H200 SXM.

Two ways to get GPUs running.

Need GPUs now? Rent them by the hour on GPU Cloud. Need dedicated infrastructure, a specific topology, or data residency guarantees? We'll build it for you. Both run the same GPU line-up.

GPU CLOUD

Rent capacity

Pick a GPU in the portal and start. On-demand bills hourly; reserved instances lower the rate for longer runs.

From $7.38/hr on B300
No commitment to start
PFS storage at $0.20/GiB per month
CLUSTER BUILD

Build dedicated

One team handles GPU selection, fabric design, power and cooling, installation, and operations for a dedicated cluster.

Dedicated to one customer
Interconnect topology designed to spec
Our facility or yours
GPU LINE-UP

Blackwell and Hopper — pick the generation you need.

BLACKWELL

NVIDIA B300

288GB
VRAM

Largest memory per node — fewer nodes for big models.

BLACKWELL

NVIDIA B200

192GB
VRAM

The default Blackwell choice, used for both training and inference.

BLACKWELL

NVIDIA GB300

288GB × N
VRAM

Delivered as NVL rack units, allocated by reservation only.

HOPPER

NVIDIA H200 SXM

141GB
VRAM

Stable supply and lower rates — a good fit for inference.

AROUND THE GPU

GPUs alone don't make a training run.

A narrow fabric leaves GPUs idle; slow storage stalls checkpoints. Both products include the layers below.

Interconnect
InfiniBand NDR / XDR

Node-to-node traffic runs over InfiniBand. Topology depends on scale and communication pattern.

Storage
PFS cluster storage + Object Storage

Checkpoints and training data on NVMe-backed PFS; archives on S3-compatible Object Storage.

Network
Own backbone

Data ingress and egress ride the PacketStream backbone and direct ISP interconnects.

Observability
Telemetry export

GPU, node, and fabric metrics exported to Prometheus, Grafana, Datadog, or your existing stack.

Delivery
Bare metal or Kubernetes

Handed over at the OS layer, or with Kubernetes and a scheduler configured.

Use Cases

Model training

Foundation model pre-training
Fine-tuning and alignment
Multi-node distributed training

Inference serving

LLM serving
Long-context workloads
Batch inference

Data processing

Pre-processing and embeddings
Vector indexing
RAG pipelines

Regulated environments

Data residency requirements
Isolated networks
Audit logging

Not sure which one fits?

Tell us the workload, how many GPUs you need, and for how long — we'll work out which option costs less.