Nebius
A GPU cloud that builds its own data centres, renting NVIDIA clusters with managed Slurm and Kubernetes.
Highlights
- Bare-metal and virtualized NVIDIA GPUs: H100, H200, B200, B300, GB200 NVL72 and GB300 NVL72
- High-bandwidth InfiniBand networking for multi-node distributed training
- Managed Slurm and managed Kubernetes orchestration for cluster scheduling
- Token Factory: a hosted LLM inference API for serving open-weight and custom models
- Managed data services including MLflow, PostgreSQL and Apache Spark
- Terraform provider, public API and CLI for infrastructure-as-code provisioning
- Owned data centers and racks rather than rented third-party capacity
- 24/7 support with solution architects for large training deployments
External link — opens nebius.com in a new tab. Nebius is a third-party product; we are not affiliated with it.
About Nebius
What it is
Nebius operates large NVIDIA GPU clusters for training and inference, from H100 through GB200, with managed Slurm and Kubernetes and an inference API on top. It designs its own racks and data centres rather than reselling capacity. It is listed on Nasdaq as NBIS, spun out of Yandex's non-Russian assets, and led by Yandex co-founder Arkady Volozh.
Why it's different
Owning the physical layer is the distinction. A lot of GPU cloud companies are resellers with a billing portal, which means they cannot control availability, density or price beyond what their supplier allows. Nebius building its own facilities is why it can offer large contiguous clusters, which is what actually matters for training — a hundred GPUs that cannot talk to each other quickly is not the same resource as a hundred that can. Weigh the corporate history honestly: the Yandex lineage means some organisations will have questions about ownership and jurisdiction that a listing cannot answer for them.
How people use it
It is used by teams training or fine-tuning models at a scale where the hyperscalers are either too expensive or unable to allocate what is needed. Managed Slurm is the tell for the intended user — that is a research and HPC workflow, not a web application one. Smaller teams use the inference API instead. As with any GPU cloud, the numbers that decide it are interconnect and real availability, not the headline hourly rate.
Written by the n3os team. We are not affiliated with Nebius.
This listing was written from public information, without Nebius’s involvement. If you own it and something here is wrong — or you would rather not be listed at all — email us and we will correct or remove it.
Get the ones worth knowing about
We write one of these for every tool worth the trouble. Get the new ones, plus what we have found genuinely useful lately.
Your address goes to Buttondown, who send the emails on our behalf. One click unsubscribes, and the list is never sold or shared.