Red Hat Launches llm-d, a Kubernetes-Based Platform for Scalable AI Inference

watch 1m, 1s
views 2

13:10, 22.05.2025

Article Content
arrow

  • Key Features of llm-d
  • Cooperation with Leading Players in the AI Industry
  • Technology and Architecture

Red Hat has introduced llm-d, a new open source project designed for high-performance distributed inference of large language models (LLMs). The platform is developed on Kubernetes and is focused on simplifying the scaling of generative AI. The source code is available on GitHub under the Apache 2.0 license.

Key Features of llm-d

The main features of the platform include

  • Optimized Inference Scheduler for vLLM;
  • Disaggregated service architecture;
  • Reuse of prefix caches;
  • Flexible scaling depending on traffic, tasks, and available resources.

Cooperation with Leading Players in the AI Industry

The development is carried out in partnership with such companies as Nvidia, AMD, Intel, IBM Research, Google Cloud, CoreWeave, Hugging Face, and others. Such cooperation emphasizes the seriousness of the approach to llm-d and the potential of the platform as an industry standard.

Technology and Architecture

The project uses the vLLM library for distributed inference, as well as components such as LMCache for KV cache offloading, AI-enabled intelligent traffic routing, highly efficient communication APIs, and automatic scaling to load and infrastructure.

All this allows you to adapt the system to different usage scenarios and performance requirements. And the launch of llm-d can be a significant step towards democratizing powerful AI systems and making them accessible to a wide audience of developers and researchers.

Share

Was this article helpful to you?

VPS popular offers

sale

-20%

CPU
CPU
3 Xeon Cores
RAM
RAM
1 GB
Space
Space
40 GB HDD
Bandwidth
Bandwidth
Unlimited
wKVM-HDD 1024 Windows

12.1 /mo

/mo

Billed monthly

sale

-20%

CPU
CPU
6 Xeon Cores
RAM
RAM
16 GB
Space
Space
150 GB SSD
Bandwidth
Bandwidth
Unlimited
KVM-SSD 16384 Linux

49.99 /mo

/mo

Billed monthly

sale

-20%

CPU
CPU
4 Epyc Cores
RAM
RAM
4 GB
Space
Space
50 GB NVMe
Bandwidth
Bandwidth
Unlimited
KVM-NVMe 4096 Linux

16.45 /mo

/mo

Billed semiannually

sale

-20%

CPU
CPU
4 Xeon Cores
RAM
RAM
2 GB
Space
Space
60 GB HDD
Bandwidth
Bandwidth
300 Gb
wKVM-HDD HK 2048 Windows

11.55 /mo

/mo

Billed monthly

sale

-20%

CPU
CPU
4 Xeon Cores
RAM
RAM
4 GB
Space
Space
50 GB SSD
Bandwidth
Bandwidth
Unlimited
KVM-SSD 4096 Linux

15.95 /mo

/mo

Billed monthly

-10%

CPU
CPU
8 Epyc Cores
RAM
RAM
32 GB
Space
Space
200 GB NVMe
Bandwidth
Bandwidth
Unlimited
Keitaro KVM 32768
OS
CentOS
Software
Software
Keitaro

77.54 /mo

/mo

Billed annually

sale

-20%

CPU
CPU
4 Xeon Cores
RAM
RAM
2 GB
Space
Space
75 GB SSD
Bandwidth
Bandwidth
2 TB
wKVM-SSD 2048 Metered Windows

24 /mo

/mo

Billed monthly

sale

-20%

CPU
CPU
6 Xeon Cores
RAM
RAM
8 GB
Space
Space
100 GB SSD
Bandwidth
Bandwidth
Unlimited
KVM-SSD 8192 Linux

25.85 /mo

/mo

Billed monthly

sale

-21.5%

CPU
CPU
2 Xeon Cores
RAM
RAM
4 GB
Space
Space
100 GB SSD
Bandwidth
Bandwidth
300 GB
wKVM-SSD 4096 HK Windows

40 /mo

/mo

Billed annually

sale

-20%

CPU
CPU
4 Xeon Cores
RAM
RAM
4 GB
Space
Space
100 GB HDD
Bandwidth
Bandwidth
Unlimited
KVM-HDD 4096 Linux

15 /mo

/mo

Billed monthly

Other articles on this topic

cookie

Accept cookies & privacy policy?

We use cookies to ensure that we give you the best experience on our website. If you continue without changing your settings, we'll assume that you are happy to receive all cookies on the HostZealot website.