Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .agents/adding-backends.md
Original file line number Diff line number Diff line change
Expand Up @@ -92,7 +92,7 @@ If you have a `prepare.sh` doing the clone, delete it — the recipe belongs in
- CUDA 13 builds: Add after other CUDA 13 builds (e.g., after `gpu-nvidia-cuda-13-chatterbox`)

**Additional build types you may need:**
- ROCm/HIP: Use `build-type: 'hipblas'` with `base-image: "rocm/dev-ubuntu-24.04:7.2.1"`
- ROCm/HIP: Use `build-type: 'hipblas'` with `base-image: "rocm/dev-ubuntu-24.04:7.14.0-full"`
- Intel/SYCL: Use `build-type: 'intel'` or `build-type: 'sycl_f16'`/`sycl_f32` with `base-image: "intel/oneapi-basekit:2025.3.2-0-devel-ubuntu24.04"`
- L4T (ARM): Use `build-type: 'l4t'` with `platforms: 'linux/arm64'` and `runs-on: 'ubuntu-24.04-arm'`

Expand Down
2 changes: 1 addition & 1 deletion .agents/building-and-testing.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,7 @@ Let's say the user wants to build a particular backend for a given platform. For
- Use `.github/backend-matrix.yml` as a reference — it's the data-only YAML that lists every backend variant's `build-type`, `base-image`, `platforms`, etc. (`backend.yml` and `backend_pr.yml` consume it via `scripts/changed-backends.js`).
- l4t and cublas also require the CUDA major and minor version.
- For llama-cpp / ik-llama-cpp / turboquant the matrix also sets `builder-base-image` pointing at a prebuilt `quay.io/go-skynet/ci-cache:base-grpc-*` tag. Local `make backends/<name>` defaults to `BUILDER_TARGET=builder-fromsource` and doesn't need it — the Dockerfile's from-source stage installs everything itself.
- You can pretty print a command like `DOCKER_MAKEFLAGS=-j$(nproc --ignore=1) BUILD_TYPE=hipblas BASE_IMAGE=rocm/dev-ubuntu-24.04:7.2.1 make docker-build-coqui`
- You can pretty print a command like `DOCKER_MAKEFLAGS=-j$(nproc --ignore=1) BUILD_TYPE=hipblas BASE_IMAGE=rocm/dev-ubuntu-24.04:7.14.0-full make docker-build-coqui`
- Unless the user specifies that they want you to run the command, then just print it because not all agent frontends handle long running jobs well and the output may overflow your context
- The user may say they want to build AMD or ROCM instead of hipblas, or Intel instead of SYCL or NVIDIA insted of l4t or cublas. Ask for confirmation if there is ambiguity.
- Sometimes the user may need extra parameters to be added to `docker build` (e.g. `--platform` for cross-platform builds or `--progress` to view the full logs), in which case you can generate the `docker build` command directly.
Expand Down
2 changes: 1 addition & 1 deletion .agents/ci-caching.md
Original file line number Diff line number Diff line change
Expand Up @@ -51,7 +51,7 @@ The C++ backend Dockerfiles (`Dockerfile.{llama-cpp,ik-llama-cpp,turboquant}`) c
| `base-grpc-cuda-13-amd64` | the above + CUDA 13.0 toolkit (Ubuntu 22.04 base) |
| `base-grpc-cuda-13-arm64` | the above + CUDA 13.0 sbsa toolkit (Ubuntu 24.04 base) |
| `base-grpc-l4t-cuda-12-arm64` | JetPack r36.4.0 base (CUDA preinstalled, `SKIP_DRIVERS=true`) + gRPC |
| `base-grpc-rocm-amd64` | rocm/dev-ubuntu-24.04:7.2.1 base + hipblas/hipblaslt/rocblas + gRPC |
| `base-grpc-rocm-amd64` | rocm/dev-ubuntu-24.04:7.14.0-full base + hipblas/hipblaslt/rocblas + gRPC |
| `base-grpc-vulkan-amd64` / `base-grpc-vulkan-arm64` | Ubuntu 24.04 + Vulkan SDK 1.4.335 + gRPC |
| `base-grpc-intel-amd64` | intel/oneapi-basekit:2025.3.2 base + gRPC |

Expand Down
22 changes: 12 additions & 10 deletions .docker/install-base-deps.sh
Original file line number Diff line number Diff line change
Expand Up @@ -226,16 +226,18 @@ fi

# --- 6. ROCm / HIP build deps (BUILD_TYPE=hipblas) ---
if [ "${BUILD_TYPE:-}" = "hipblas" ] && [ "${SKIP_DRIVERS:-false}" = "false" ]; then
apt-get update
apt-get install -y --no-install-recommends \
hipblas-dev \
hipblaslt-dev \
rocblas-dev
apt-get clean
rm -rf /var/lib/apt/lists/*
# I have no idea why, but the ROCM lib packages don't trigger ldconfig after they install,
# which results in local-ai and others not being able to locate the libraries.
# We run ldconfig ourselves to work around this packaging deficiency.
# ROCm 7.x ("TheRock" packaging): the legacy hipblas-dev / hipblaslt-dev /
# rocblas-dev metapackages were removed — BLAS is consolidated and split per
# GPU arch (amdrocm-blas<ver>-gfx*, amdrocm-blas-dev, amdrocm-blas-host). The
# rocm/dev-ubuntu-*:*-full base image already ships those BLAS dev libs and
# headers, so no apt install is needed here.
# TheRock ships the ROCm libs under /opt/rocm/core*/lib (reached via the
# /opt/rocm/lib alternatives symlink) but registers no ld.so.conf.d entry, so
# the dynamic linker can't resolve them. Without this the ldd-based GPU-lib
# packaging (scripts/build/package-gpu-libs.sh) silently skips librocm_kpack,
# rocBLAS, hipBLASLt, the bundled rocm_sysdeps (zlib/zstd/elf) and the LLVM
# runtime, and the built backend fails to load at runtime.
printf '/opt/rocm/lib\n/opt/rocm/lib/rocm_sysdeps/lib\n/opt/rocm/llvm/lib\n' > /etc/ld.so.conf.d/rocm.conf
ldconfig
# Log which GPU architectures have rocBLAS kernel support
echo "rocBLAS library data architectures:"
Expand Down
98 changes: 68 additions & 30 deletions .github/backend-matrix.yml
Original file line number Diff line number Diff line change
Expand Up @@ -2288,7 +2288,7 @@ include:
tag-latest: 'auto'
tag-suffix: '-gpu-rocm-hipblas-rerankers'
runs-on: 'ubuntu-latest'
base-image: "rocm/dev-ubuntu-24.04:7.2.1"
base-image: "rocm/dev-ubuntu-24.04:7.14.0-full"
skip-drivers: 'false'
backend: "rerankers"
dockerfile: "./backend/Dockerfile.python"
Expand All @@ -2302,7 +2302,7 @@ include:
tag-suffix: '-gpu-rocm-hipblas-llama-cpp'
builder-base-image: 'quay.io/go-skynet/ci-cache:base-grpc-rocm-amd64'
runs-on: 'ubuntu-latest'
base-image: "rocm/dev-ubuntu-24.04:7.2.1"
base-image: "rocm/dev-ubuntu-24.04:7.14.0-full"
skip-drivers: 'false'
backend: "llama-cpp"
dockerfile: "./backend/Dockerfile.llama-cpp"
Expand All @@ -2316,7 +2316,7 @@ include:
tag-suffix: '-gpu-rocm-hipblas-bonsai'
builder-base-image: 'quay.io/go-skynet/ci-cache:base-grpc-rocm-amd64'
runs-on: 'ubuntu-latest'
base-image: "rocm/dev-ubuntu-24.04:7.2.1"
base-image: "rocm/dev-ubuntu-24.04:7.14.0-full"
skip-drivers: 'false'
backend: "bonsai"
dockerfile: "./backend/Dockerfile.bonsai"
Expand All @@ -2329,9 +2329,32 @@ include:
tag-latest: 'auto'
tag-suffix: '-gpu-rocm-hipblas-vllm'
runs-on: 'ubuntu-latest'
base-image: "rocm/dev-ubuntu-24.04:7.2.1"
base-image: "rocm/dev-ubuntu-24.04:7.14.0-full"
skip-drivers: 'false'
backend: "vllm"
# CDNA data-center arches -- the only ones wheels.vllm.ai/rocm publishes a
# wheel for. Pinned rather than inherited so this image keeps installing that
# prebuilt wheel; the repo-wide default list contains consumer arches, which
# would silently turn this into a multi-hour source build.
amdgpu-targets: 'gfx942,gfx950'
dockerfile: "./backend/Dockerfile.python"
context: "./"
ubuntu-version: '2404'
# Consumer/RDNA AMD GPUs. No prebuilt vllm wheel exists for these arches, so
# install.sh builds vllm from source -- a separate image because the two install
# paths cannot share one venv. One arch per image: each AMD torch device package
# is ~1.6 GB installed, so bundling arches is not free.
- build-type: 'hipblas'
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-rocm-hipblas-gfx1151-vllm'
runs-on: 'ubuntu-latest'
base-image: "rocm/dev-ubuntu-24.04:7.14.0-full"
skip-drivers: 'false'
backend: "vllm"
amdgpu-targets: 'gfx1151'
dockerfile: "./backend/Dockerfile.python"
context: "./"
ubuntu-version: '2404'
Expand All @@ -2342,7 +2365,7 @@ include:
tag-latest: 'auto'
tag-suffix: '-gpu-rocm-hipblas-vllm-omni'
runs-on: 'ubuntu-latest'
base-image: "rocm/dev-ubuntu-24.04:7.2.1"
base-image: "rocm/dev-ubuntu-24.04:7.14.0-full"
skip-drivers: 'false'
backend: "vllm-omni"
dockerfile: "./backend/Dockerfile.python"
Expand All @@ -2355,7 +2378,7 @@ include:
tag-latest: 'auto'
tag-suffix: '-gpu-rocm-hipblas-sglang'
runs-on: 'ubuntu-latest'
base-image: "rocm/dev-ubuntu-24.04:7.2.1"
base-image: "rocm/dev-ubuntu-24.04:7.14.0-full"
skip-drivers: 'false'
backend: "sglang"
dockerfile: "./backend/Dockerfile.python"
Expand All @@ -2368,9 +2391,14 @@ include:
tag-latest: 'auto'
tag-suffix: '-gpu-rocm-hipblas-transformers'
runs-on: 'ubuntu-latest'
base-image: "rocm/dev-ubuntu-24.04:7.2.1"
base-image: "rocm/dev-ubuntu-24.04:7.14.0-full"
skip-drivers: 'false'
backend: "transformers"
# AMD torch device packages are per-GPU-arch and ~1.6 GB installed each, so
# this image carries a deliberate set instead of the repo-wide 11-arch default
# (which would add ~17 GB): both CDNA data-center arches plus Strix Halo.
# Adding an arch is one entry here at ~1.6 GB.
amdgpu-targets: 'gfx942,gfx950,gfx1151'
dockerfile: "./backend/Dockerfile.python"
context: "./"
ubuntu-version: '2404'
Expand All @@ -2381,9 +2409,14 @@ include:
tag-latest: 'auto'
tag-suffix: '-gpu-rocm-hipblas-diffusers'
runs-on: 'ubuntu-latest'
base-image: "rocm/dev-ubuntu-24.04:7.2.1"
base-image: "rocm/dev-ubuntu-24.04:7.14.0-full"
skip-drivers: 'false'
backend: "diffusers"
# AMD torch device packages are per-GPU-arch and ~1.6 GB installed each, so
# this image carries a deliberate set instead of the repo-wide 11-arch default
# (which would add ~17 GB): both CDNA data-center arches plus Strix Halo.
# Adding an arch is one entry here at ~1.6 GB.
amdgpu-targets: 'gfx942,gfx950,gfx1151'
dockerfile: "./backend/Dockerfile.python"
context: "./"
ubuntu-version: '2404'
Expand All @@ -2394,7 +2427,7 @@ include:
tag-latest: 'auto'
tag-suffix: '-gpu-rocm-hipblas-ace-step'
runs-on: 'ubuntu-latest'
base-image: "rocm/dev-ubuntu-24.04:7.2.1"
base-image: "rocm/dev-ubuntu-24.04:7.14.0-full"
skip-drivers: 'false'
backend: "ace-step"
dockerfile: "./backend/Dockerfile.python"
Expand All @@ -2408,9 +2441,14 @@ include:
tag-latest: 'auto'
tag-suffix: '-gpu-rocm-hipblas-kokoro'
runs-on: 'ubuntu-latest'
base-image: "rocm/dev-ubuntu-24.04:7.2.1"
base-image: "rocm/dev-ubuntu-24.04:7.14.0-full"
skip-drivers: 'false'
backend: "kokoro"
# AMD torch device packages are per-GPU-arch and ~1.6 GB installed each, so
# this image carries a deliberate set instead of the repo-wide 11-arch default
# (which would add ~17 GB): both CDNA data-center arches plus Strix Halo.
# Adding an arch is one entry here at ~1.6 GB.
amdgpu-targets: 'gfx942,gfx950,gfx1151'
dockerfile: "./backend/Dockerfile.python"
context: "./"
ubuntu-version: '2404'
Expand All @@ -2421,7 +2459,7 @@ include:
tag-latest: 'auto'
tag-suffix: '-gpu-rocm-hipblas-vibevoice'
runs-on: 'ubuntu-latest'
base-image: "rocm/dev-ubuntu-24.04:7.2.1"
base-image: "rocm/dev-ubuntu-24.04:7.14.0-full"
skip-drivers: 'false'
backend: "vibevoice"
dockerfile: "./backend/Dockerfile.python"
Expand All @@ -2434,7 +2472,7 @@ include:
tag-latest: 'auto'
tag-suffix: '-gpu-rocm-hipblas-liquid-audio'
runs-on: 'ubuntu-latest'
base-image: "rocm/dev-ubuntu-24.04:7.2.1"
base-image: "rocm/dev-ubuntu-24.04:7.14.0-full"
skip-drivers: 'false'
backend: "liquid-audio"
dockerfile: "./backend/Dockerfile.python"
Expand All @@ -2447,7 +2485,7 @@ include:
tag-latest: 'auto'
tag-suffix: '-gpu-rocm-hipblas-qwen-asr'
runs-on: 'ubuntu-latest'
base-image: "rocm/dev-ubuntu-24.04:7.2.1"
base-image: "rocm/dev-ubuntu-24.04:7.14.0-full"
skip-drivers: 'false'
backend: "qwen-asr"
dockerfile: "./backend/Dockerfile.python"
Expand All @@ -2473,7 +2511,7 @@ include:
tag-latest: 'auto'
tag-suffix: '-gpu-rocm-hipblas-nemo'
runs-on: 'ubuntu-latest'
base-image: "rocm/dev-ubuntu-24.04:7.2.1"
base-image: "rocm/dev-ubuntu-24.04:7.14.0-full"
skip-drivers: 'false'
backend: "nemo"
dockerfile: "./backend/Dockerfile.python"
Expand All @@ -2486,7 +2524,7 @@ include:
tag-latest: 'auto'
tag-suffix: '-gpu-rocm-hipblas-qwen-tts'
runs-on: 'ubuntu-latest'
base-image: "rocm/dev-ubuntu-24.04:7.2.1"
base-image: "rocm/dev-ubuntu-24.04:7.14.0-full"
skip-drivers: 'false'
backend: "qwen-tts"
dockerfile: "./backend/Dockerfile.python"
Expand All @@ -2499,7 +2537,7 @@ include:
tag-latest: 'auto'
tag-suffix: '-gpu-rocm-hipblas-fish-speech'
runs-on: 'ubuntu-latest'
base-image: "rocm/dev-ubuntu-24.04:7.2.1"
base-image: "rocm/dev-ubuntu-24.04:7.14.0-full"
skip-drivers: 'false'
backend: "fish-speech"
dockerfile: "./backend/Dockerfile.python"
Expand All @@ -2512,7 +2550,7 @@ include:
tag-latest: 'auto'
tag-suffix: '-gpu-rocm-hipblas-voxcpm'
runs-on: 'ubuntu-latest'
base-image: "rocm/dev-ubuntu-24.04:7.2.1"
base-image: "rocm/dev-ubuntu-24.04:7.14.0-full"
skip-drivers: 'false'
backend: "voxcpm"
dockerfile: "./backend/Dockerfile.python"
Expand All @@ -2525,7 +2563,7 @@ include:
tag-latest: 'auto'
tag-suffix: '-gpu-rocm-hipblas-pocket-tts'
runs-on: 'ubuntu-latest'
base-image: "rocm/dev-ubuntu-24.04:7.2.1"
base-image: "rocm/dev-ubuntu-24.04:7.14.0-full"
skip-drivers: 'false'
backend: "pocket-tts"
dockerfile: "./backend/Dockerfile.python"
Expand All @@ -2538,7 +2576,7 @@ include:
tag-latest: 'auto'
tag-suffix: '-gpu-rocm-hipblas-faster-whisper'
runs-on: 'ubuntu-latest'
base-image: "rocm/dev-ubuntu-24.04:7.2.1"
base-image: "rocm/dev-ubuntu-24.04:7.14.0-full"
skip-drivers: 'false'
backend: "faster-whisper"
dockerfile: "./backend/Dockerfile.python"
Expand All @@ -2551,7 +2589,7 @@ include:
tag-latest: 'auto'
tag-suffix: '-gpu-rocm-hipblas-coqui'
runs-on: 'ubuntu-latest'
base-image: "rocm/dev-ubuntu-24.04:7.2.1"
base-image: "rocm/dev-ubuntu-24.04:7.14.0-full"
skip-drivers: 'false'
backend: "coqui"
dockerfile: "./backend/Dockerfile.python"
Expand Down Expand Up @@ -4229,7 +4267,7 @@ include:
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-rocm-hipblas-whisper'
base-image: "rocm/dev-ubuntu-24.04:7.2.1"
base-image: "rocm/dev-ubuntu-24.04:7.14.0-full"
runs-on: 'ubuntu-latest'
skip-drivers: 'false'
backend: "whisper"
Expand All @@ -4242,7 +4280,7 @@ include:
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-rocm-hipblas-crispasr'
base-image: "rocm/dev-ubuntu-24.04:7.2.1"
base-image: "rocm/dev-ubuntu-24.04:7.14.0-full"
runs-on: 'ubuntu-latest'
skip-drivers: 'false'
backend: "crispasr"
Expand Down Expand Up @@ -4351,7 +4389,7 @@ include:
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-rocm-hipblas-parakeet-cpp'
base-image: "rocm/dev-ubuntu-24.04:7.2.1"
base-image: "rocm/dev-ubuntu-24.04:7.14.0-full"
runs-on: 'ubuntu-latest'
skip-drivers: 'false'
backend: "parakeet-cpp"
Expand Down Expand Up @@ -4540,7 +4578,7 @@ include:
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-rocm-hipblas-moss-transcribe-cpp'
base-image: "rocm/dev-ubuntu-24.04:7.2.1"
base-image: "rocm/dev-ubuntu-24.04:7.14.0-full"
runs-on: 'ubuntu-latest'
skip-drivers: 'false'
backend: "moss-transcribe-cpp"
Expand Down Expand Up @@ -4688,7 +4726,7 @@ include:
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-rocm-hipblas-ced'
base-image: "rocm/dev-ubuntu-24.04:7.2.1"
base-image: "rocm/dev-ubuntu-24.04:7.14.0-full"
runs-on: 'ubuntu-latest'
skip-drivers: 'false'
backend: "ced"
Expand Down Expand Up @@ -4836,7 +4874,7 @@ include:
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-rocm-hipblas-voice-detect'
base-image: "rocm/dev-ubuntu-24.04:7.2.1"
base-image: "rocm/dev-ubuntu-24.04:7.14.0-full"
runs-on: 'ubuntu-latest'
skip-drivers: 'false'
backend: "voice-detect"
Expand Down Expand Up @@ -4984,7 +5022,7 @@ include:
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-rocm-hipblas-face-detect'
base-image: "rocm/dev-ubuntu-24.04:7.2.1"
base-image: "rocm/dev-ubuntu-24.04:7.14.0-full"
runs-on: 'ubuntu-latest'
skip-drivers: 'false'
backend: "face-detect"
Expand Down Expand Up @@ -5093,7 +5131,7 @@ include:
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-rocm-hipblas-acestep-cpp'
base-image: "rocm/dev-ubuntu-24.04:7.2.1"
base-image: "rocm/dev-ubuntu-24.04:7.14.0-full"
runs-on: 'ubuntu-latest'
skip-drivers: 'false'
backend: "acestep-cpp"
Expand Down Expand Up @@ -6055,7 +6093,7 @@ include:
# platforms: 'linux/amd64'
# tag-latest: 'auto'
# tag-suffix: '-gpu-hipblas-rfdetr'
# base-image: "rocm/dev-ubuntu-24.04:7.2.1"
# base-image: "rocm/dev-ubuntu-24.04:7.14.0-full"
# runs-on: 'ubuntu-latest'
# skip-drivers: 'false'
# backend: "rfdetr"
Expand Down Expand Up @@ -6127,7 +6165,7 @@ include:
tag-latest: 'auto'
tag-suffix: '-gpu-rocm-hipblas-neutts'
runs-on: 'ubuntu-latest'
base-image: "rocm/dev-ubuntu-24.04:7.2.1"
base-image: "rocm/dev-ubuntu-24.04:7.14.0-full"
skip-drivers: 'false'
backend: "neutts"
dockerfile: "./backend/Dockerfile.python"
Expand Down
2 changes: 1 addition & 1 deletion .github/workflows/base-images.yml
Original file line number Diff line number Diff line change
Expand Up @@ -83,7 +83,7 @@ jobs:
ubuntu-version: '2404'
- tag: 'base-grpc-rocm-amd64'
runs-on: 'ubuntu-latest'
base-image: 'rocm/dev-ubuntu-24.04:7.2.1'
base-image: 'rocm/dev-ubuntu-24.04:7.14.0-full'
build-type: 'hipblas'
cuda-major-version: ''
cuda-minor-version: ''
Expand Down
2 changes: 1 addition & 1 deletion .github/workflows/image-pr.yml
Original file line number Diff line number Diff line change
Expand Up @@ -80,7 +80,7 @@
platforms: 'linux/amd64'
tag-latest: 'false'
tag-suffix: '-hipblas'
base-image: "rocm/dev-ubuntu-24.04:7.2.1"
base-image: "rocm/dev-ubuntu-24.04:7.14.0-full"
runs-on: 'ubuntu-latest'
makeflags: "--jobs=3 --output-sync=target"
ubuntu-version: '2404'
Expand Down
2 changes: 1 addition & 1 deletion .github/workflows/image.yml
Original file line number Diff line number Diff line change
Expand Up @@ -85,7 +85,7 @@
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-hipblas'
base-image: "rocm/dev-ubuntu-24.04:7.2.1"
base-image: "rocm/dev-ubuntu-24.04:7.14.0-full"
runs-on: 'ubuntu-latest'
makeflags: "--jobs=3 --output-sync=target"
ubuntu-version: '2404'
Expand Down
Loading