ROCm 6.3.3 / 7.2.4 / 7.14 / 10.0 bench
ROCm 6.3.3|7.2.4|7.14|10.0 * split tensor|layer * GPUs 1|2|4 * ctx 0|16K|32K|64K- llama.cpp
v0.5.0(7fe450e),0.5.0-dev (build 1) - toolchain:
ML-gfx906@7a0d586— imagesregistry.arkprojects.space/apps/llama.cpp-gfx906:v0.5.0-rocm-<ver>-7a0d586-pre - mobo:
imb760 - 4× MI50/MI60 32 GiB behind PCIe bridges: GPU↔bridge 16 GT/s x16, bridge↔CPU x8 (CPU root port is x8),
HIP_VISIBLE_DEVICES=1,0,3,2,GGML_HIP_GRAPHS=ON - numbers below are the average of 5 llama-bench repetitions
The previous run (llama.cpp b9180, ROCm 6.3.3/7.2.3, hip graphs off/on) is kept on the ROCm _ graphs _ GPUs bench page.
All GPUs in all tests used with HBM2 mem OC (freq + timings). OC profile profile-mi50-113-D1631700-111-mem-oc.yaml
Results
Full per-model tables: Qwen3.8-27B, Qwen3.8-Flash-Next, gemma-4-26B-A4B-it, gemma-4-31B-it.
Avg tok/s for Prompt 2048 / Gen 0 is prompt processing, for Prompt 0 / Gen 256 is generation at the given Depth.
Summary vs ROCm 6.3.3
Median change across configurations (rows = one config, one ROCm version), relative to 6.3.3. In parentheses: number of configs with a statistically significant change (more than 2σ of the combined measurement noise).
| Model | Test | 7.2.4 | 7.14 | 10.0 |
|---|---|---|---|---|
| Qwen3.8-27B | pp2048 | -0.2% (2/2) | +11.2% (8/0) | +15.6% (8/0) |
| Qwen3.8-27B | tg256 | +24.6% (8/0) | +28.9% (8/0) | +28.7% (7/0) |
| Qwen3.8-Flash-Next | pp2048 | +1.4% (0/1) | +8.9% (4/0) | +8.6% (4/0) |
| Qwen3.8-Flash-Next | tg256 | +7.5% (4/0) | +9.0% (4/0) | +7.9% (4/0) |
| gemma-4-26B-A4B-it | pp2048 | -5.1% (3/8) | +6.6% (9/2) | +5.5% (9/1) |
| gemma-4-26B-A4B-it | tg256 | +17.2% (12/0) | +19.6% (12/0) | +18.5% (12/0) |
| gemma-4-31B-it | pp2048 | -5.0% (2/5) | +7.7% (5/0) | +3.8% (4/1) |
| gemma-4-31B-it | tg256 | +19.0% (8/0) | +19.8% (8/0) | +20.6% (8/0) |
Change by context depth
Median across all models and configurations.
| Depth | 7.2.4 | 7.14 | 10.0 |
|---|---|---|---|
| 0 | +10.2% | +17.7% | +19.9% |
| 16384 | +7.6% | +15.2% | +16.3% |
| 32768 | +8.3% | +15.7% | +15.6% |
| 65536 | +5.4% | +13.9% | +15.0% |
Takeaways
- Generation: +17…+29% on every 7.x ROCm in almost every configuration; the largest gains are at short context (
+20…+37%depending on model and GPU count). MoE Qwen3.8-Flash-Next gains the least (+7…+10%), dense models gain the most. - Prompt processing on 4 GPUs (tensor): consistent +8…+25%, 10.0 is the fastest (Qwen3.8-27B
+22.5%, gemma-4-26B+23.4%). - Prompt processing on 2 GPUs (tensor): mixed. Qwen3.8-27B +8…+19% on 7.14/10.0, gemma-4-26B +3…+11%, but 7.2.4 regresses on gemma (
−5…−9%) and gemma-4-31B (−7…−14%). - Single GPU (layer, gemma-4-26B) at 32K/64K context: the only stable regression,
−2…−4%on 7.14/10.0 and−8…−11%on 7.2.4; at 0/16K the new versions are equal or faster. - Gain shrinks with context depth: median
+20%at depth 0 down to+15%at 64K (10.0); 7.2.4 drops from+10%to+5%. - 7.2.4 is the weakest of the 7.x for prompt processing. 7.14 and 10.0 are near parity overall; 10.0 wins the Qwen3.8-27B 4-GPU prompt case by
+6…+12%.
Bench
The command is rendered from the YAML profiles by
build-bench-command.py,
copied to the container together with the profiles. ROCM_VER selects the result directory,
so one loop over the presets covers all ROCm versions.
PRESETS=(
'gemma-4-26B-A4B-it[tensor]'
'gemma-4-26B-A4B-it[layer]'
'gemma-4-31B-it[tensor]'
'Qwen3.8-27B[tensor]'
'Qwen3.8-Flash-Next[layer]'
)
for (( i=0; i<${#PRESETS[@]}; i++ )); do
echo "======================= ${PRESETS[$i]} ======================="
python3 ./bench-profiles/build-bench-command.py --source-dir ./bench-profiles --bench "${PRESETS[$i]}" --output - | bash
done
- global.yaml
- Qwen3.8.yaml
- gemma-4-26b-a4b-it.yaml
- gemma-4-31b-it.yaml
- table format
"*":
command: |
#!/usr/bin/env bash
set -eo pipefail
mkdir -p ./bench-results/rocm-$ROCM_VER
echo $ROCM_VER > ./bench-results/rocm-$ROCM_VER/rocm-ver
{env} ./llama-bench {args} | tee ./bench-results/rocm-$ROCM_VER/{profile}.jsonl
env:
HIP_VISIBLE_DEVICES: "1,0,3,2"
HSA_FORCE_FINE_GRAIN_PCIE: "1"
HSA_XNACK: "0"
args:
lazy-mode: "off"
offline: true
output: jsonl
verbose: false
progress: true
Qwen3.8-27B[tensor]:
args:
hf-repo: bartowski/Qwen3.8-27B-GGUF:Q8_0
flash-attn: true
split-mode: tensor
device:
- [ROCm0, ROCm1, ROCm2, ROCm3]
- [ROCm0, ROCm1]
n-prompt: 2048
ubatch-size: 2048
n-gen: 256
n-depth: [0, 16384, 32768, 65536]
Qwen3.8-Flash-Next[layer]:
args:
hf-repo: bartowski/Qwen3.8-Flash-Next-GGUF:Q4_K_M
flash-attn: auto
split-mode: layer
device:
- [ROCm0, ROCm1, ROCm2, ROCm3]
n-prompt: 2048
ubatch-size: 2048
n-gen: 256
n-depth: [0, 16384, 32768, 65536]
gemma-4-26B-A4B-it[tensor]:
args:
hf-repo: bartowski/google_gemma-4-26B-A4B-it-GGUF:Q8_0
flash-attn: true
split-mode: tensor
device:
- [ROCm0, ROCm1, ROCm2, ROCm3]
- [ROCm0, ROCm1]
n-prompt: 2048
ubatch-size: 2048
n-gen: 256
n-depth: [0, 16384, 32768, 65536]
gemma-4-26B-A4B-it[layer]:
args:
hf-repo: bartowski/google_gemma-4-26B-A4B-it-GGUF:Q8_0
flash-attn: true
split-mode: layer
device:
- [ROCm0]
n-prompt: 2048
ubatch-size: 2048
n-gen: 256
n-depth: [0, 16384, 32768, 65536]
gemma-4-31B-it[tensor]:
args:
hf-repo: unsloth/gemma-4-31B-it-GGUF:Q8_K_XL
flash-attn: true
split-mode: [tensor]
device:
- [ROCm0, ROCm1, ROCm2, ROCm3]
- [ROCm0, ROCm1]
n-prompt: 2048
ubatch-size: 2048
n-gen: 256
n-depth: [0, 16384, 32768, 65536]
version: 1
columns:
rocm_version:
name: ROCm
split_mode:
name: Split
device_count:
field: devices
name: GPUs
format: count
separator: "/"
flash_attn:
name: FA
format: map
map:
-1: auto
0: "off"
1: "on"
n_prompt:
name: Prompt
n_gen:
name: Gen
n_depth:
name: Depth
avg_ts:
name: Avg tok/s
precision: 2
stddev_ts:
name: ±
precision: 2
Software
Build config and version
version: 0.5.0-dev (build 1, commit 7fe450e)
built with GNU 13.3.0 for Linux x86_64
-DCMAKE_BUILD_TYPE=Release
-DLLAMA_BUILD_TESTS=OFF
-DGGML_BACKEND_DL=ON
-DGGML_HIP=ON
-DGGML_HIP_GRAPHS=ON
-DGGML_HIP_RCCL=ON
-DAMDGPU_TARGETS=gfx906
-DGGML_RPC=ON
-DGGML_CPU_ALL_VARIANTS=ON
-DGGML_AVX512=ON
-DGGML_AVX512_VBMI=ON
-DGGML_AVX512_VNNI=ON
-DGGML_AVX512_BF16=ON
Extra HIP flags (LLAMA_CMAKE_HIP_FLAGS):
| ROCm | HIP flags |
|---|---|
| 6.3.3, 7.2.4 | (none) |
| 7.14, 10.0 | -mllvm -amdgpu-sched-strategy=max-ilp |
Topology
================================ Weight between two GPUs =================================
GPU0 GPU1 GPU2 GPU3
GPU0 0 40 40 40
GPU1 40 0 40 40
GPU2 40 40 0 40
GPU3 40 40 40 0
================================= Hops between two GPUs ==================================
GPU0 GPU1 GPU2 GPU3
GPU0 0 2 2 2
GPU1 2 0 2 2
GPU2 2 2 0 2
GPU3 2 2 2 0
=============================== Link Type between two GPUs ===============================
GPU0 GPU1 GPU2 GPU3
GPU0 0 PCIE PCIE PCIE
GPU1 PCIE 0 PCIE PCIE
GPU2 PCIE PCIE 0 PCIE
GPU3 PCIE PCIE PCIE 0
GPU[0..3]: (Topology) Numa Node: 0, Numa Affinity: 0
31:00.0
LnkCap: Port #0, Speed 16GT/s, Width x16, ASPM L1, Exit Latency L1 <64us
LnkSta: Speed 16GT/s, Width x8 (downgraded)
34:00.0
LnkCap: Port #2, Speed 16GT/s, Width x16, ASPM L1, Exit Latency L1 <64us
LnkSta: Speed 16GT/s, Width x8 (downgraded)
4b:00.0
LnkCap: Port #0, Speed 16GT/s, Width x16, ASPM L1, Exit Latency L1 <64us
LnkSta: Speed 16GT/s, Width x8 (downgraded)
4e:00.0
LnkCap: Port #2, Speed 16GT/s, Width x16, ASPM L1, Exit Latency L1 <64us
LnkSta: Speed 16GT/s, Width x8 (downgraded)
ggml_cuda_init: found 4 ROCm devices (Total VRAM: 131008 MiB):
Device 0: AMD Instinct MI60 / MI50, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64, VRAM: 32752 MiB
Device 1: AMD Instinct MI60 / MI50, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64, VRAM: 32752 MiB
Device 2: AMD Instinct MI60 / MI50, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64, VRAM: 32752 MiB
Device 3: AMD Instinct MI60 / MI50, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64, VRAM: 32752 MiB
model name: Genuine Intel(R) CPU 0000%@
threads: 76 (2 NUMA nodes, GPUs on node 0)
Qwen3.8-27B
bartowski/Qwen3.8-27B-GGUF:Q8_0 — dense 27B, Q8_0, 27.1 GiB. tensor split over 4 and 2 GPUs, flash-attn on.
| 6.3.3 | tensor | 2 | on | 2048 | 0 | 0 | 455.53 | 50.18 |
| 7.2.4 | tensor | 2 | on | 2048 | 0 | 0 | 456.64 | 1.62 |
| 7.14 | tensor | 2 | on | 2048 | 0 | 0 | 538.06 | 0.34 |
| 10.0 | tensor | 2 | on | 2048 | 0 | 0 | 517.84 | 13.39 |
| 6.3.3 | tensor | 2 | on | 2048 | 0 | 16384 | 408.98 | 1.79 |
| 7.2.4 | tensor | 2 | on | 2048 | 0 | 16384 | 389.50 | 1.75 |
| 7.14 | tensor | 2 | on | 2048 | 0 | 16384 | 454.96 | 1.12 |
| 10.0 | tensor | 2 | on | 2048 | 0 | 16384 | 441.31 | 11.61 |
| 6.3.3 | tensor | 2 | on | 2048 | 0 | 32768 | 355.32 | 0.93 |
| 7.2.4 | tensor | 2 | on | 2048 | 0 | 32768 | 339.85 | 0.74 |
| 7.14 | tensor | 2 | on | 2048 | 0 | 32768 | 395.14 | 1.53 |
| 10.0 | tensor | 2 | on | 2048 | 0 | 32768 | 393.69 | 1.76 |
| 6.3.3 | tensor | 2 | on | 2048 | 0 | 65536 | 254.94 | 47.47 |
| 7.2.4 | tensor | 2 | on | 2048 | 0 | 65536 | 263.91 | 3.04 |
| 7.14 | tensor | 2 | on | 2048 | 0 | 65536 | 306.86 | 4.62 |
| 10.0 | tensor | 2 | on | 2048 | 0 | 65536 | 303.13 | 5.67 |
| 6.3.3 | tensor | 2 | on | 0 | 256 | 0 | 27.10 | 0.81 |
| 7.2.4 | tensor | 2 | on | 0 | 256 | 0 | 33.02 | 0.28 |
| 7.14 | tensor | 2 | on | 0 | 256 | 0 | 33.83 | 0.13 |
| 10.0 | tensor | 2 | on | 0 | 256 | 0 | 34.34 | 0.33 |
| 6.3.3 | tensor | 2 | on | 0 | 256 | 16384 | 26.82 | 0.32 |
| 7.2.4 | tensor | 2 | on | 0 | 256 | 16384 | 31.37 | 0.16 |
| 7.14 | tensor | 2 | on | 0 | 256 | 16384 | 32.13 | 0.16 |
| 10.0 | tensor | 2 | on | 0 | 256 | 16384 | 31.94 | 1.72 |
| 6.3.3 | tensor | 2 | on | 0 | 256 | 32768 | 25.90 | 0.29 |
| 7.2.4 | tensor | 2 | on | 0 | 256 | 32768 | 29.65 | 0.31 |
| 7.14 | tensor | 2 | on | 0 | 256 | 32768 | 30.79 | 0.17 |
| 10.0 | tensor | 2 | on | 0 | 256 | 32768 | 30.75 | 0.80 |
| 6.3.3 | tensor | 2 | on | 0 | 256 | 65536 | 20.50 | 2.39 |
| 7.2.4 | tensor | 2 | on | 0 | 256 | 65536 | 27.49 | 0.27 |
| 7.14 | tensor | 2 | on | 0 | 256 | 65536 | 28.04 | 0.61 |
| 10.0 | tensor | 2 | on | 0 | 256 | 65536 | 28.00 | 0.69 |
| 6.3.3 | tensor | 4 | on | 2048 | 0 | 0 | 704.48 | 6.04 |
| 7.2.4 | tensor | 4 | on | 2048 | 0 | 0 | 726.73 | 2.43 |
| 7.14 | tensor | 4 | on | 2048 | 0 | 0 | 767.75 | 0.29 |
| 10.0 | tensor | 4 | on | 2048 | 0 | 0 | 862.82 | 0.77 |
| 6.3.3 | tensor | 4 | on | 2048 | 0 | 16384 | 635.69 | 7.68 |
| 7.2.4 | tensor | 4 | on | 2048 | 0 | 16384 | 631.04 | 5.38 |
| 7.14 | tensor | 4 | on | 2048 | 0 | 16384 | 667.45 | 21.03 |
| 10.0 | tensor | 4 | on | 2048 | 0 | 16384 | 739.79 | 7.42 |
| 6.3.3 | tensor | 4 | on | 2048 | 0 | 32768 | 521.95 | 40.57 |
| 7.2.4 | tensor | 4 | on | 2048 | 0 | 32768 | 559.99 | 3.11 |
| 7.14 | tensor | 4 | on | 2048 | 0 | 32768 | 605.99 | 4.00 |
| 10.0 | tensor | 4 | on | 2048 | 0 | 32768 | 652.02 | 4.08 |
| 6.3.3 | tensor | 4 | on | 2048 | 0 | 65536 | 457.83 | 6.54 |
| 7.2.4 | tensor | 4 | on | 2048 | 0 | 65536 | 452.04 | 4.69 |
| 7.14 | tensor | 4 | on | 2048 | 0 | 65536 | 495.69 | 3.96 |
| 10.0 | tensor | 4 | on | 2048 | 0 | 65536 | 525.52 | 4.51 |
| 6.3.3 | tensor | 4 | on | 0 | 256 | 0 | 30.35 | 2.10 |
| 7.2.4 | tensor | 4 | on | 0 | 256 | 0 | 38.64 | 1.18 |
| 7.14 | tensor | 4 | on | 0 | 256 | 0 | 40.04 | 0.93 |
| 10.0 | tensor | 4 | on | 0 | 256 | 0 | 40.06 | 1.27 |
| 6.3.3 | tensor | 4 | on | 0 | 256 | 16384 | 28.71 | 1.89 |
| 7.2.4 | tensor | 4 | on | 0 | 256 | 16384 | 36.70 | 1.07 |
| 7.14 | tensor | 4 | on | 0 | 256 | 16384 | 38.74 | 0.57 |
| 10.0 | tensor | 4 | on | 0 | 256 | 16384 | 38.83 | 0.68 |
| 6.3.3 | tensor | 4 | on | 0 | 256 | 32768 | 30.01 | 1.92 |
| 7.2.4 | tensor | 4 | on | 0 | 256 | 32768 | 36.21 | 0.47 |
| 7.14 | tensor | 4 | on | 0 | 256 | 32768 | 37.78 | 0.58 |
| 10.0 | tensor | 4 | on | 0 | 256 | 32768 | 34.49 | 6.77 |
| 6.3.3 | tensor | 4 | on | 0 | 256 | 65536 | 25.75 | 3.93 |
| 7.2.4 | tensor | 4 | on | 0 | 256 | 65536 | 33.27 | 0.44 |
| 7.14 | tensor | 4 | on | 0 | 256 | 65536 | 34.75 | 2.06 |
| 10.0 | tensor | 4 | on | 0 | 256 | 65536 | 33.67 | 1.84 |
Qwen3.8-Flash-Next
bartowski/Qwen3.8-Flash-Next-GGUF:Q4_K_M — MoE A3B, Q4_K_M, ~111 GiB. layer split over 4 GPUs, flash-attn auto.
| 6.3.3 | layer | 4 | auto | 2048 | 0 | 0 | 513.24 | 1.71 |
| 7.2.4 | layer | 4 | auto | 2048 | 0 | 0 | 507.29 | 1.34 |
| 7.14 | layer | 4 | auto | 2048 | 0 | 0 | 554.28 | 1.16 |
| 10.0 | layer | 4 | auto | 2048 | 0 | 0 | 553.18 | 2.70 |
| 6.3.3 | layer | 4 | auto | 2048 | 0 | 16384 | 383.61 | 8.49 |
| 7.2.4 | layer | 4 | auto | 2048 | 0 | 16384 | 387.25 | 5.88 |
| 7.14 | layer | 4 | auto | 2048 | 0 | 16384 | 423.51 | 4.31 |
| 10.0 | layer | 4 | auto | 2048 | 0 | 16384 | 417.79 | 5.63 |
| 6.3.3 | layer | 4 | auto | 2048 | 0 | 32768 | 309.66 | 7.54 |
| 7.2.4 | layer | 4 | auto | 2048 | 0 | 32768 | 315.27 | 9.58 |
| 7.14 | layer | 4 | auto | 2048 | 0 | 32768 | 333.23 | 8.95 |
| 10.0 | layer | 4 | auto | 2048 | 0 | 32768 | 338.30 | 2.55 |
| 6.3.3 | layer | 4 | auto | 2048 | 0 | 65536 | 218.48 | 7.45 |
| 7.2.4 | layer | 4 | auto | 2048 | 0 | 65536 | 225.26 | 6.80 |
| 7.14 | layer | 4 | auto | 2048 | 0 | 65536 | 239.86 | 4.02 |
| 10.0 | layer | 4 | auto | 2048 | 0 | 65536 | 236.51 | 6.09 |
| 6.3.3 | layer | 4 | auto | 0 | 256 | 0 | 32.96 | 0.13 |
| 7.2.4 | layer | 4 | auto | 0 | 256 | 0 | 35.47 | 0.06 |
| 7.14 | layer | 4 | auto | 0 | 256 | 0 | 35.73 | 0.08 |
| 10.0 | layer | 4 | auto | 0 | 256 | 0 | 35.17 | 0.16 |
| 6.3.3 | layer | 4 | auto | 0 | 256 | 16384 | 26.59 | 0.35 |
| 7.2.4 | layer | 4 | auto | 0 | 256 | 16384 | 28.58 | 0.38 |
| 7.14 | layer | 4 | auto | 0 | 256 | 16384 | 29.14 | 0.28 |
| 10.0 | layer | 4 | auto | 0 | 256 | 16384 | 28.63 | 0.32 |
| 6.3.3 | layer | 4 | auto | 0 | 256 | 32768 | 22.23 | 0.96 |
| 7.2.4 | layer | 4 | auto | 0 | 256 | 32768 | 24.29 | 0.36 |
| 7.14 | layer | 4 | auto | 0 | 256 | 32768 | 24.39 | 0.37 |
| 10.0 | layer | 4 | auto | 0 | 256 | 32768 | 24.86 | 0.27 |
| 6.3.3 | layer | 4 | auto | 0 | 256 | 65536 | 18.56 | 0.18 |
| 7.2.4 | layer | 4 | auto | 0 | 256 | 65536 | 19.91 | 0.31 |
| 7.14 | layer | 4 | auto | 0 | 256 | 65536 | 20.13 | 0.31 |
| 10.0 | layer | 4 | auto | 0 | 256 | 65536 | 20.07 | 0.36 |
gemma-4-26B-A4B-it
bartowski/google_gemma-4-26B-A4B-it-GGUF:Q8_0 — MoE A4B, Q8_0, 25.0 GiB. tensor split over 4 and 2 GPUs plus layer split on 1 GPU, flash-attn on. Both presets share this table, Split/GPUs columns tell them apart.
| 6.3.3 | layer | 1 | on | 2048 | 0 | 0 | 1497.42 | 14.93 |
| 7.2.4 | layer | 1 | on | 2048 | 0 | 0 | 1418.10 | 14.17 |
| 7.14 | layer | 1 | on | 2048 | 0 | 0 | 1596.72 | 25.76 |
| 10.0 | layer | 1 | on | 2048 | 0 | 0 | 1580.72 | 44.38 |
| 6.3.3 | layer | 1 | on | 2048 | 0 | 16384 | 951.00 | 4.18 |
| 7.2.4 | layer | 1 | on | 2048 | 0 | 16384 | 863.03 | 5.87 |
| 7.14 | layer | 1 | on | 2048 | 0 | 16384 | 955.87 | 8.37 |
| 10.0 | layer | 1 | on | 2048 | 0 | 16384 | 961.58 | 16.00 |
| 6.3.3 | layer | 1 | on | 2048 | 0 | 32768 | 726.47 | 9.11 |
| 7.2.4 | layer | 1 | on | 2048 | 0 | 32768 | 649.55 | 19.81 |
| 7.14 | layer | 1 | on | 2048 | 0 | 32768 | 702.92 | 9.26 |
| 10.0 | layer | 1 | on | 2048 | 0 | 32768 | 714.15 | 10.77 |
| 6.3.3 | layer | 1 | on | 2048 | 0 | 65536 | 492.73 | 12.85 |
| 7.2.4 | layer | 1 | on | 2048 | 0 | 65536 | 450.29 | 13.00 |
| 7.14 | layer | 1 | on | 2048 | 0 | 65536 | 473.39 | 8.10 |
| 10.0 | layer | 1 | on | 2048 | 0 | 65536 | 473.19 | 8.29 |
| 6.3.3 | layer | 1 | on | 0 | 256 | 0 | 78.58 | 0.14 |
| 7.2.4 | layer | 1 | on | 0 | 256 | 0 | 94.50 | 0.22 |
| 7.14 | layer | 1 | on | 0 | 256 | 0 | 95.60 | 0.51 |
| 10.0 | layer | 1 | on | 0 | 256 | 0 | 95.19 | 0.82 |
| 6.3.3 | layer | 1 | on | 0 | 256 | 16384 | 72.27 | 0.21 |
| 7.2.4 | layer | 1 | on | 0 | 256 | 16384 | 85.29 | 0.19 |
| 7.14 | layer | 1 | on | 0 | 256 | 16384 | 86.29 | 0.37 |
| 10.0 | layer | 1 | on | 0 | 256 | 16384 | 85.74 | 0.66 |
| 6.3.3 | layer | 1 | on | 0 | 256 | 32768 | 69.47 | 1.27 |
| 7.2.4 | layer | 1 | on | 0 | 256 | 32768 | 81.55 | 0.17 |
| 7.14 | layer | 1 | on | 0 | 256 | 32768 | 82.69 | 0.35 |
| 10.0 | layer | 1 | on | 0 | 256 | 32768 | 81.38 | 1.24 |
| 6.3.3 | layer | 1 | on | 0 | 256 | 65536 | 63.43 | 1.46 |
| 7.2.4 | layer | 1 | on | 0 | 256 | 65536 | 74.44 | 0.25 |
| 7.14 | layer | 1 | on | 0 | 256 | 65536 | 75.41 | 0.54 |
| 10.0 | layer | 1 | on | 0 | 256 | 65536 | 75.04 | 0.92 |
| 6.3.3 | tensor | 2 | on | 2048 | 0 | 0 | 1997.36 | 11.71 |
| 7.2.4 | tensor | 2 | on | 2048 | 0 | 0 | 1899.30 | 14.05 |
| 7.14 | tensor | 2 | on | 2048 | 0 | 0 | 2128.68 | 12.74 |
| 10.0 | tensor | 2 | on | 2048 | 0 | 0 | 2106.77 | 13.45 |
| 6.3.3 | tensor | 2 | on | 2048 | 0 | 16384 | 1525.91 | 10.10 |
| 7.2.4 | tensor | 2 | on | 2048 | 0 | 16384 | 1398.98 | 29.50 |
| 7.14 | tensor | 2 | on | 2048 | 0 | 16384 | 1589.93 | 19.62 |
| 10.0 | tensor | 2 | on | 2048 | 0 | 16384 | 1581.45 | 9.53 |
| 6.3.3 | tensor | 2 | on | 2048 | 0 | 32768 | 1184.74 | 34.45 |
| 7.2.4 | tensor | 2 | on | 2048 | 0 | 32768 | 1157.55 | 15.26 |
| 7.14 | tensor | 2 | on | 2048 | 0 | 32768 | 1308.59 | 16.95 |
| 10.0 | tensor | 2 | on | 2048 | 0 | 32768 | 1299.56 | 16.82 |
| 6.3.3 | tensor | 2 | on | 2048 | 0 | 65536 | 927.28 | 25.62 |
| 7.2.4 | tensor | 2 | on | 2048 | 0 | 65536 | 852.72 | 3.70 |
| 7.14 | tensor | 2 | on | 2048 | 0 | 65536 | 965.35 | 4.44 |
| 10.0 | tensor | 2 | on | 2048 | 0 | 65536 | 960.41 | 4.07 |
| 6.3.3 | tensor | 2 | on | 0 | 256 | 0 | 64.37 | 0.95 |
| 7.2.4 | tensor | 2 | on | 0 | 256 | 0 | 72.60 | 0.52 |
| 7.14 | tensor | 2 | on | 0 | 256 | 0 | 75.43 | 1.77 |
| 10.0 | tensor | 2 | on | 0 | 256 | 0 | 76.42 | 0.80 |
| 6.3.3 | tensor | 2 | on | 0 | 256 | 16384 | 60.70 | 0.78 |
| 7.2.4 | tensor | 2 | on | 0 | 256 | 16384 | 68.21 | 0.92 |
| 7.14 | tensor | 2 | on | 0 | 256 | 16384 | 70.23 | 0.67 |
| 10.0 | tensor | 2 | on | 0 | 256 | 16384 | 70.52 | 0.75 |
| 6.3.3 | tensor | 2 | on | 0 | 256 | 32768 | 59.53 | 1.01 |
| 7.2.4 | tensor | 2 | on | 0 | 256 | 32768 | 67.10 | 0.71 |
| 7.14 | tensor | 2 | on | 0 | 256 | 32768 | 69.14 | 0.47 |
| 10.0 | tensor | 2 | on | 0 | 256 | 32768 | 69.16 | 0.84 |
| 6.3.3 | tensor | 2 | on | 0 | 256 | 65536 | 55.01 | 1.65 |
| 7.2.4 | tensor | 2 | on | 0 | 256 | 65536 | 64.39 | 0.90 |
| 7.14 | tensor | 2 | on | 0 | 256 | 65536 | 66.37 | 0.64 |
| 10.0 | tensor | 2 | on | 0 | 256 | 65536 | 64.11 | 4.85 |
| 6.3.3 | tensor | 4 | on | 2048 | 0 | 0 | 2215.81 | 47.42 |
| 7.2.4 | tensor | 4 | on | 2048 | 0 | 0 | 2516.70 | 4.04 |
| 7.14 | tensor | 4 | on | 2048 | 0 | 0 | 2725.89 | 14.64 |
| 10.0 | tensor | 4 | on | 2048 | 0 | 0 | 2734.33 | 14.09 |
| 6.3.3 | tensor | 4 | on | 2048 | 0 | 16384 | 1629.52 | 13.16 |
| 7.2.4 | tensor | 4 | on | 2048 | 0 | 16384 | 1756.24 | 15.21 |
| 7.14 | tensor | 4 | on | 2048 | 0 | 16384 | 1913.80 | 34.39 |
| 10.0 | tensor | 4 | on | 2048 | 0 | 16384 | 1930.41 | 17.83 |
| 6.3.3 | tensor | 4 | on | 2048 | 0 | 32768 | 1269.39 | 86.38 |
| 7.2.4 | tensor | 4 | on | 2048 | 0 | 32768 | 1363.28 | 20.32 |
| 7.14 | tensor | 4 | on | 2048 | 0 | 32768 | 1515.96 | 16.34 |
| 10.0 | tensor | 4 | on | 2048 | 0 | 32768 | 1513.67 | 16.06 |
| 6.3.3 | tensor | 4 | on | 2048 | 0 | 65536 | 972.36 | 11.10 |
| 7.2.4 | tensor | 4 | on | 2048 | 0 | 65536 | 949.67 | 9.65 |
| 7.14 | tensor | 4 | on | 2048 | 0 | 65536 | 1054.84 | 9.65 |
| 10.0 | tensor | 4 | on | 2048 | 0 | 65536 | 1058.20 | 8.35 |
| 6.3.3 | tensor | 4 | on | 0 | 256 | 0 | 53.82 | 1.46 |
| 7.2.4 | tensor | 4 | on | 0 | 256 | 0 | 63.64 | 1.22 |
| 7.14 | tensor | 4 | on | 0 | 256 | 0 | 67.28 | 0.93 |
| 10.0 | tensor | 4 | on | 0 | 256 | 0 | 67.02 | 1.31 |
| 6.3.3 | tensor | 4 | on | 0 | 256 | 16384 | 51.25 | 1.65 |
| 7.2.4 | tensor | 4 | on | 0 | 256 | 16384 | 59.74 | 1.66 |
| 7.14 | tensor | 4 | on | 0 | 256 | 16384 | 61.38 | 1.58 |
| 10.0 | tensor | 4 | on | 0 | 256 | 16384 | 61.94 | 1.11 |
| 6.3.3 | tensor | 4 | on | 0 | 256 | 32768 | 50.44 | 1.22 |
| 7.2.4 | tensor | 4 | on | 0 | 256 | 32768 | 59.27 | 0.96 |
| 7.14 | tensor | 4 | on | 0 | 256 | 32768 | 60.64 | 1.39 |
| 10.0 | tensor | 4 | on | 0 | 256 | 32768 | 61.00 | 1.04 |
| 6.3.3 | tensor | 4 | on | 0 | 256 | 65536 | 49.30 | 1.17 |
| 7.2.4 | tensor | 4 | on | 0 | 256 | 65536 | 57.25 | 1.06 |
| 7.14 | tensor | 4 | on | 0 | 256 | 65536 | 59.22 | 1.03 |
| 10.0 | tensor | 4 | on | 0 | 256 | 65536 | 57.81 | 1.63 |
gemma-4-31B-it
unsloth/gemma-4-31B-it-GGUF:Q8_K_XL — dense 31B, Q8_K_XL, 32.6 GiB. tensor split over 4 and 2 GPUs, flash-attn on.
| 6.3.3 | tensor | 2 | on | 2048 | 0 | 0 | 420.57 | 0.44 |
| 7.2.4 | tensor | 2 | on | 2048 | 0 | 0 | 378.59 | 19.45 |
| 7.14 | tensor | 2 | on | 2048 | 0 | 0 | 434.27 | 34.41 |
| 10.0 | tensor | 2 | on | 2048 | 0 | 0 | 419.95 | 11.35 |
| 6.3.3 | tensor | 2 | on | 2048 | 0 | 16384 | 308.92 | 9.60 |
| 7.2.4 | tensor | 2 | on | 2048 | 0 | 16384 | 288.61 | 1.68 |
| 7.14 | tensor | 2 | on | 2048 | 0 | 16384 | 324.70 | 1.34 |
| 10.0 | tensor | 2 | on | 2048 | 0 | 16384 | 326.50 | 0.57 |
| 6.3.3 | tensor | 2 | on | 2048 | 0 | 32768 | 263.83 | 2.78 |
| 7.2.4 | tensor | 2 | on | 2048 | 0 | 32768 | 227.26 | 4.35 |
| 7.14 | tensor | 2 | on | 2048 | 0 | 32768 | 262.26 | 3.36 |
| 10.0 | tensor | 2 | on | 2048 | 0 | 32768 | 254.97 | 4.22 |
| 6.3.3 | tensor | 2 | on | 2048 | 0 | 65536 | 187.83 | 12.63 |
| 7.2.4 | tensor | 2 | on | 2048 | 0 | 65536 | 167.17 | 3.30 |
| 7.14 | tensor | 2 | on | 2048 | 0 | 65536 | 185.60 | 1.76 |
| 10.0 | tensor | 2 | on | 2048 | 0 | 65536 | 186.24 | 1.72 |
| 6.3.3 | tensor | 2 | on | 0 | 256 | 0 | 23.76 | 0.63 |
| 7.2.4 | tensor | 2 | on | 0 | 256 | 0 | 28.21 | 0.14 |
| 7.14 | tensor | 2 | on | 0 | 256 | 0 | 28.68 | 0.05 |
| 10.0 | tensor | 2 | on | 0 | 256 | 0 | 28.98 | 0.15 |
| 6.3.3 | tensor | 2 | on | 0 | 256 | 16384 | 22.90 | 0.45 |
| 7.2.4 | tensor | 2 | on | 0 | 256 | 16384 | 26.63 | 0.07 |
| 7.14 | tensor | 2 | on | 0 | 256 | 16384 | 26.47 | 0.10 |
| 10.0 | tensor | 2 | on | 0 | 256 | 16384 | 24.90 | 1.51 |
| 6.3.3 | tensor | 2 | on | 0 | 256 | 32768 | 21.68 | 1.64 |
| 7.2.4 | tensor | 2 | on | 0 | 256 | 32768 | 24.69 | 0.98 |
| 7.14 | tensor | 2 | on | 0 | 256 | 32768 | 25.76 | 0.10 |
| 10.0 | tensor | 2 | on | 0 | 256 | 32768 | 26.12 | 0.08 |
| 6.3.3 | tensor | 2 | on | 0 | 256 | 65536 | 20.94 | 1.48 |
| 7.2.4 | tensor | 2 | on | 0 | 256 | 65536 | 24.09 | 0.42 |
| 7.14 | tensor | 2 | on | 0 | 256 | 65536 | 23.41 | 0.31 |
| 10.0 | tensor | 2 | on | 0 | 256 | 65536 | 23.93 | 0.16 |
| 6.3.3 | tensor | 4 | on | 2048 | 0 | 0 | 679.78 | 2.62 |
| 7.2.4 | tensor | 4 | on | 2048 | 0 | 0 | 704.02 | 1.08 |
| 7.14 | tensor | 4 | on | 2048 | 0 | 0 | 789.14 | 1.24 |
| 10.0 | tensor | 4 | on | 2048 | 0 | 0 | 787.36 | 8.55 |
| 6.3.3 | tensor | 4 | on | 2048 | 0 | 16384 | 539.22 | 4.64 |
| 7.2.4 | tensor | 4 | on | 2048 | 0 | 16384 | 555.62 | 3.84 |
| 7.14 | tensor | 4 | on | 2048 | 0 | 16384 | 619.50 | 37.49 |
| 10.0 | tensor | 4 | on | 2048 | 0 | 16384 | 629.57 | 5.47 |
| 6.3.3 | tensor | 4 | on | 2048 | 0 | 32768 | 489.98 | 10.48 |
| 7.2.4 | tensor | 4 | on | 2048 | 0 | 32768 | 473.07 | 6.80 |
| 7.14 | tensor | 4 | on | 2048 | 0 | 32768 | 540.40 | 4.60 |
| 10.0 | tensor | 4 | on | 2048 | 0 | 32768 | 499.25 | 39.31 |
| 6.3.3 | tensor | 4 | on | 2048 | 0 | 65536 | 361.75 | 27.51 |
| 7.2.4 | tensor | 4 | on | 2048 | 0 | 65536 | 368.77 | 0.95 |
| 7.14 | tensor | 4 | on | 2048 | 0 | 65536 | 419.38 | 3.23 |
| 10.0 | tensor | 4 | on | 2048 | 0 | 65536 | 417.11 | 4.94 |
| 6.3.3 | tensor | 4 | on | 0 | 256 | 0 | 28.98 | 0.38 |
| 7.2.4 | tensor | 4 | on | 0 | 256 | 0 | 34.82 | 0.42 |
| 7.14 | tensor | 4 | on | 0 | 256 | 0 | 35.75 | 0.28 |
| 10.0 | tensor | 4 | on | 0 | 256 | 0 | 35.70 | 0.71 |
| 6.3.3 | tensor | 4 | on | 0 | 256 | 16384 | 26.87 | 0.79 |
| 7.2.4 | tensor | 4 | on | 0 | 256 | 16384 | 32.22 | 0.53 |
| 7.14 | tensor | 4 | on | 0 | 256 | 16384 | 33.01 | 0.52 |
| 10.0 | tensor | 4 | on | 0 | 256 | 16384 | 33.37 | 0.42 |
| 6.3.3 | tensor | 4 | on | 0 | 256 | 32768 | 27.06 | 0.44 |
| 7.2.4 | tensor | 4 | on | 0 | 256 | 32768 | 32.29 | 0.43 |
| 7.14 | tensor | 4 | on | 0 | 256 | 32768 | 31.18 | 0.71 |
| 10.0 | tensor | 4 | on | 0 | 256 | 32768 | 32.51 | 0.90 |
| 6.3.3 | tensor | 4 | on | 0 | 256 | 65536 | 24.48 | 3.70 |
| 7.2.4 | tensor | 4 | on | 0 | 256 | 65536 | 30.09 | 0.89 |
| 7.14 | tensor | 4 | on | 0 | 256 | 65536 | 30.10 | 0.30 |
| 10.0 | tensor | 4 | on | 0 | 256 | 65536 | 29.58 | 0.13 |