Go to file

biondizzle 360b0dea58 Restore CUDA 13.0.1 + patch vLLM for cuMemcpyBatchAsync API change

CUDA 13 removed the fail_idx parameter from cuMemcpyBatchAsync.
Patch cache_kernels.cu to match new API signature instead of downgrading.

- Restore CUDA 13.0.1, PyTorch 2.9.0+cu130, flashinfer cu130
- Patch: remove fail_idx variable and parameter from cuMemcpyBatchAsync call
- Simplify error message to not reference fail_idx

2026-04-03 07:53:12 +00:00

lmcache

Updated for vllm v0.10.2

2025-09-24 05:52:11 +00:00

vllm

Restore CUDA 13.0.1 + patch vLLM for cuMemcpyBatchAsync API change

2026-04-03 07:53:12 +00:00

.gitignore

move things

2026-04-03 04:27:21 +00:00

README.md

Updated to vLLM v0.11.1rc3

2025-10-23 18:16:57 +00:00

README.md

Building containers for GH200

Currently, prebuilt wheels for vLLM and LMcache are not available for aarch64. This can make setup tedious when working on modern aarch64 platforms such as NVIDIA GH200.

Further, Nvidia at this time does not provide the Dockerfile associated with the NGC containers which makes replacing some of the components (like a newer version of vLLM) tedious.

This repository provides a Dockerfile to build a container with vLLM and all its dependencies pre-installed to try out various things such as KV offloading.

If you prefer not to build the image yourself, you can pull the ready-to-use image directly from Docker Hub:

docker run --rm -it --gpus all -v "$PWD":"$PWD" -w "$PWD" rajesh550/gh200-vllm:0.11.0 bash

# CUDA 13
docker run --rm -it --gpus all -v "$PWD":"$PWD" -w "$PWD" rajesh550/gh200-vllm:0.11.1rc2 bash

👉 Docker Hub

Version info:

CUDA: 13.0.1
Ubuntu: 24.04
Python: 3.12
PyTorch: 2.9.0+cu130
Triton: 3.5.x
xformers: 0.32.post2+
flashinfer: 0.4.1
flashattention: 3.0.0b1
LMCache: 0.3.7
vLLM: 0.11.1rc3