Skip to main content

ROCm 6.3.3 / 7.2.4 / 7.14 / 10.0 bench

  • ROCm 6.3.3|7.2.4|7.14|10.0 * split tensor|layer * GPUs 1|2|4 * ctx 0|16K|32K|64K
  • llama.cpp v0.5.0 (7fe450e), 0.5.0-dev (build 1)
  • toolchain: ML-gfx906 @ 7a0d586 — images registry.arkprojects.space/apps/llama.cpp-gfx906:v0.5.0-rocm-<ver>-7a0d586-pre
  • mobo: imb760
  • 4× MI50/MI60 32 GiB behind PCIe bridges: GPU↔bridge 16 GT/s x16, bridge↔CPU x8 (CPU root port is x8), HIP_VISIBLE_DEVICES=1,0,3,2, GGML_HIP_GRAPHS=ON
  • numbers below are the average of 5 llama-bench repetitions
info

The previous run (llama.cpp b9180, ROCm 6.3.3/7.2.3, hip graphs off/on) is kept on the ROCm _ graphs _ GPUs bench page.

note

All GPUs in all tests used with HBM2 mem OC (freq + timings). OC profile profile-mi50-113-D1631700-111-mem-oc.yaml

Results​

Full per-model tables: Qwen3.8-27B, Qwen3.8-Flash-Next, gemma-4-26B-A4B-it, gemma-4-31B-it.

Avg tok/s for Prompt 2048 / Gen 0 is prompt processing, for Prompt 0 / Gen 256 is generation at the given Depth.

Summary vs ROCm 6.3.3​

Median change across configurations (rows = one config, one ROCm version), relative to 6.3.3. In parentheses: number of configs with a statistically significant change (more than 2σ of the combined measurement noise).

ModelTest7.2.47.1410.0
Qwen3.8-27Bpp2048-0.2% (2/2)+11.2% (8/0)+15.6% (8/0)
Qwen3.8-27Btg256+24.6% (8/0)+28.9% (8/0)+28.7% (7/0)
Qwen3.8-Flash-Nextpp2048+1.4% (0/1)+8.9% (4/0)+8.6% (4/0)
Qwen3.8-Flash-Nexttg256+7.5% (4/0)+9.0% (4/0)+7.9% (4/0)
gemma-4-26B-A4B-itpp2048-5.1% (3/8)+6.6% (9/2)+5.5% (9/1)
gemma-4-26B-A4B-ittg256+17.2% (12/0)+19.6% (12/0)+18.5% (12/0)
gemma-4-31B-itpp2048-5.0% (2/5)+7.7% (5/0)+3.8% (4/1)
gemma-4-31B-ittg256+19.0% (8/0)+19.8% (8/0)+20.6% (8/0)

Change by context depth​

Median across all models and configurations.

Depth7.2.47.1410.0
0+10.2%+17.7%+19.9%
16384+7.6%+15.2%+16.3%
32768+8.3%+15.7%+15.6%
65536+5.4%+13.9%+15.0%

Takeaways​

  • Generation: +17…+29% on every 7.x ROCm in almost every configuration; the largest gains are at short context (+20…+37% depending on model and GPU count). MoE Qwen3.8-Flash-Next gains the least (+7…+10%), dense models gain the most.
  • Prompt processing on 4 GPUs (tensor): consistent +8…+25%, 10.0 is the fastest (Qwen3.8-27B +22.5%, gemma-4-26B +23.4%).
  • Prompt processing on 2 GPUs (tensor): mixed. Qwen3.8-27B +8…+19% on 7.14/10.0, gemma-4-26B +3…+11%, but 7.2.4 regresses on gemma (−5…−9%) and gemma-4-31B (−7…−14%).
  • Single GPU (layer, gemma-4-26B) at 32K/64K context: the only stable regression, −2…−4% on 7.14/10.0 and −8…−11% on 7.2.4; at 0/16K the new versions are equal or faster.
  • Gain shrinks with context depth: median +20% at depth 0 down to +15% at 64K (10.0); 7.2.4 drops from +10% to +5%.
  • 7.2.4 is the weakest of the 7.x for prompt processing. 7.14 and 10.0 are near parity overall; 10.0 wins the Qwen3.8-27B 4-GPU prompt case by +6…+12%.

Bench​

The command is rendered from the YAML profiles by build-bench-command.py, copied to the container together with the profiles. ROCM_VER selects the result directory, so one loop over the presets covers all ROCm versions.

run.sh
PRESETS=(
'gemma-4-26B-A4B-it[tensor]'
'gemma-4-26B-A4B-it[layer]'
'gemma-4-31B-it[tensor]'
'Qwen3.8-27B[tensor]'
'Qwen3.8-Flash-Next[layer]'
)
for (( i=0; i<${#PRESETS[@]}; i++ )); do
echo "======================= ${PRESETS[$i]} ======================="
python3 ./bench-profiles/build-bench-command.py --source-dir ./bench-profiles --bench "${PRESETS[$i]}" --output - | bash
done
bench-profiles/global.yaml
"*":
command: |
#!/usr/bin/env bash
set -eo pipefail
mkdir -p ./bench-results/rocm-$ROCM_VER
echo $ROCM_VER > ./bench-results/rocm-$ROCM_VER/rocm-ver
{env} ./llama-bench {args} | tee ./bench-results/rocm-$ROCM_VER/{profile}.jsonl
env:
HIP_VISIBLE_DEVICES: "1,0,3,2"
HSA_FORCE_FINE_GRAIN_PCIE: "1"
HSA_XNACK: "0"
args:
lazy-mode: "off"
offline: true
output: jsonl
verbose: false
progress: true

Software​

Build config and version
llama-cli --version
version: 0.5.0-dev (build 1, commit 7fe450e)
built with GNU 13.3.0 for Linux x86_64
cmake
-DCMAKE_BUILD_TYPE=Release
-DLLAMA_BUILD_TESTS=OFF
-DGGML_BACKEND_DL=ON
-DGGML_HIP=ON
-DGGML_HIP_GRAPHS=ON
-DGGML_HIP_RCCL=ON
-DAMDGPU_TARGETS=gfx906
-DGGML_RPC=ON
-DGGML_CPU_ALL_VARIANTS=ON
-DGGML_AVX512=ON
-DGGML_AVX512_VBMI=ON
-DGGML_AVX512_VNNI=ON
-DGGML_AVX512_BF16=ON

Extra HIP flags (LLAMA_CMAKE_HIP_FLAGS):

ROCmHIP flags
6.3.3, 7.2.4(none)
7.14, 10.0-mllvm -amdgpu-sched-strategy=max-ilp
Topology
rocm-smi --showtopo
================================ Weight between two GPUs =================================
GPU0 GPU1 GPU2 GPU3
GPU0 0 40 40 40
GPU1 40 0 40 40
GPU2 40 40 0 40
GPU3 40 40 40 0

================================= Hops between two GPUs ==================================
GPU0 GPU1 GPU2 GPU3
GPU0 0 2 2 2
GPU1 2 0 2 2
GPU2 2 2 0 2
GPU3 2 2 2 0

=============================== Link Type between two GPUs ===============================
GPU0 GPU1 GPU2 GPU3
GPU0 0 PCIE PCIE PCIE
GPU1 PCIE 0 PCIE PCIE
GPU2 PCIE PCIE 0 PCIE
GPU3 PCIE PCIE PCIE 0

GPU[0..3]: (Topology) Numa Node: 0, Numa Affinity: 0
lspci
31:00.0
LnkCap: Port #0, Speed 16GT/s, Width x16, ASPM L1, Exit Latency L1 <64us
LnkSta: Speed 16GT/s, Width x8 (downgraded)
34:00.0
LnkCap: Port #2, Speed 16GT/s, Width x16, ASPM L1, Exit Latency L1 <64us
LnkSta: Speed 16GT/s, Width x8 (downgraded)
4b:00.0
LnkCap: Port #0, Speed 16GT/s, Width x16, ASPM L1, Exit Latency L1 <64us
LnkSta: Speed 16GT/s, Width x8 (downgraded)
4e:00.0
LnkCap: Port #2, Speed 16GT/s, Width x16, ASPM L1, Exit Latency L1 <64us
LnkSta: Speed 16GT/s, Width x8 (downgraded)
llama-bench --list-devices
ggml_cuda_init: found 4 ROCm devices (Total VRAM: 131008 MiB):
Device 0: AMD Instinct MI60 / MI50, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64, VRAM: 32752 MiB
Device 1: AMD Instinct MI60 / MI50, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64, VRAM: 32752 MiB
Device 2: AMD Instinct MI60 / MI50, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64, VRAM: 32752 MiB
Device 3: AMD Instinct MI60 / MI50, gfx906:sramecc+:xnack- (0x906), VMM: no, Wave Size: 64, VRAM: 32752 MiB
cpu / numa
model name: Genuine Intel(R) CPU 0000%@
threads: 76 (2 NUMA nodes, GPUs on node 0)

Qwen3.8-27B​

bartowski/Qwen3.8-27B-GGUF:Q8_0 — dense 27B, Q8_0, 27.1 GiB. tensor split over 4 and 2 GPUs, flash-attn on.

6.3.3tensor2on204800455.5350.18
7.2.4tensor2on204800456.641.62
7.14tensor2on204800538.060.34
10.0tensor2on204800517.8413.39
6.3.3tensor2on2048016384408.981.79
7.2.4tensor2on2048016384389.501.75
7.14tensor2on2048016384454.961.12
10.0tensor2on2048016384441.3111.61
6.3.3tensor2on2048032768355.320.93
7.2.4tensor2on2048032768339.850.74
7.14tensor2on2048032768395.141.53
10.0tensor2on2048032768393.691.76
6.3.3tensor2on2048065536254.9447.47
7.2.4tensor2on2048065536263.913.04
7.14tensor2on2048065536306.864.62
10.0tensor2on2048065536303.135.67
6.3.3tensor2on0256027.100.81
7.2.4tensor2on0256033.020.28
7.14tensor2on0256033.830.13
10.0tensor2on0256034.340.33
6.3.3tensor2on02561638426.820.32
7.2.4tensor2on02561638431.370.16
7.14tensor2on02561638432.130.16
10.0tensor2on02561638431.941.72
6.3.3tensor2on02563276825.900.29
7.2.4tensor2on02563276829.650.31
7.14tensor2on02563276830.790.17
10.0tensor2on02563276830.750.80
6.3.3tensor2on02566553620.502.39
7.2.4tensor2on02566553627.490.27
7.14tensor2on02566553628.040.61
10.0tensor2on02566553628.000.69
6.3.3tensor4on204800704.486.04
7.2.4tensor4on204800726.732.43
7.14tensor4on204800767.750.29
10.0tensor4on204800862.820.77
6.3.3tensor4on2048016384635.697.68
7.2.4tensor4on2048016384631.045.38
7.14tensor4on2048016384667.4521.03
10.0tensor4on2048016384739.797.42
6.3.3tensor4on2048032768521.9540.57
7.2.4tensor4on2048032768559.993.11
7.14tensor4on2048032768605.994.00
10.0tensor4on2048032768652.024.08
6.3.3tensor4on2048065536457.836.54
7.2.4tensor4on2048065536452.044.69
7.14tensor4on2048065536495.693.96
10.0tensor4on2048065536525.524.51
6.3.3tensor4on0256030.352.10
7.2.4tensor4on0256038.641.18
7.14tensor4on0256040.040.93
10.0tensor4on0256040.061.27
6.3.3tensor4on02561638428.711.89
7.2.4tensor4on02561638436.701.07
7.14tensor4on02561638438.740.57
10.0tensor4on02561638438.830.68
6.3.3tensor4on02563276830.011.92
7.2.4tensor4on02563276836.210.47
7.14tensor4on02563276837.780.58
10.0tensor4on02563276834.496.77
6.3.3tensor4on02566553625.753.93
7.2.4tensor4on02566553633.270.44
7.14tensor4on02566553634.752.06
10.0tensor4on02566553633.671.84

Qwen3.8-Flash-Next​

bartowski/Qwen3.8-Flash-Next-GGUF:Q4_K_M — MoE A3B, Q4_K_M, ~111 GiB. layer split over 4 GPUs, flash-attn auto.

6.3.3layer4auto204800513.241.71
7.2.4layer4auto204800507.291.34
7.14layer4auto204800554.281.16
10.0layer4auto204800553.182.70
6.3.3layer4auto2048016384383.618.49
7.2.4layer4auto2048016384387.255.88
7.14layer4auto2048016384423.514.31
10.0layer4auto2048016384417.795.63
6.3.3layer4auto2048032768309.667.54
7.2.4layer4auto2048032768315.279.58
7.14layer4auto2048032768333.238.95
10.0layer4auto2048032768338.302.55
6.3.3layer4auto2048065536218.487.45
7.2.4layer4auto2048065536225.266.80
7.14layer4auto2048065536239.864.02
10.0layer4auto2048065536236.516.09
6.3.3layer4auto0256032.960.13
7.2.4layer4auto0256035.470.06
7.14layer4auto0256035.730.08
10.0layer4auto0256035.170.16
6.3.3layer4auto02561638426.590.35
7.2.4layer4auto02561638428.580.38
7.14layer4auto02561638429.140.28
10.0layer4auto02561638428.630.32
6.3.3layer4auto02563276822.230.96
7.2.4layer4auto02563276824.290.36
7.14layer4auto02563276824.390.37
10.0layer4auto02563276824.860.27
6.3.3layer4auto02566553618.560.18
7.2.4layer4auto02566553619.910.31
7.14layer4auto02566553620.130.31
10.0layer4auto02566553620.070.36

gemma-4-26B-A4B-it​

bartowski/google_gemma-4-26B-A4B-it-GGUF:Q8_0 — MoE A4B, Q8_0, 25.0 GiB. tensor split over 4 and 2 GPUs plus layer split on 1 GPU, flash-attn on. Both presets share this table, Split/GPUs columns tell them apart.

6.3.3layer1on2048001497.4214.93
7.2.4layer1on2048001418.1014.17
7.14layer1on2048001596.7225.76
10.0layer1on2048001580.7244.38
6.3.3layer1on2048016384951.004.18
7.2.4layer1on2048016384863.035.87
7.14layer1on2048016384955.878.37
10.0layer1on2048016384961.5816.00
6.3.3layer1on2048032768726.479.11
7.2.4layer1on2048032768649.5519.81
7.14layer1on2048032768702.929.26
10.0layer1on2048032768714.1510.77
6.3.3layer1on2048065536492.7312.85
7.2.4layer1on2048065536450.2913.00
7.14layer1on2048065536473.398.10
10.0layer1on2048065536473.198.29
6.3.3layer1on0256078.580.14
7.2.4layer1on0256094.500.22
7.14layer1on0256095.600.51
10.0layer1on0256095.190.82
6.3.3layer1on02561638472.270.21
7.2.4layer1on02561638485.290.19
7.14layer1on02561638486.290.37
10.0layer1on02561638485.740.66
6.3.3layer1on02563276869.471.27
7.2.4layer1on02563276881.550.17
7.14layer1on02563276882.690.35
10.0layer1on02563276881.381.24
6.3.3layer1on02566553663.431.46
7.2.4layer1on02566553674.440.25
7.14layer1on02566553675.410.54
10.0layer1on02566553675.040.92
6.3.3tensor2on2048001997.3611.71
7.2.4tensor2on2048001899.3014.05
7.14tensor2on2048002128.6812.74
10.0tensor2on2048002106.7713.45
6.3.3tensor2on20480163841525.9110.10
7.2.4tensor2on20480163841398.9829.50
7.14tensor2on20480163841589.9319.62
10.0tensor2on20480163841581.459.53
6.3.3tensor2on20480327681184.7434.45
7.2.4tensor2on20480327681157.5515.26
7.14tensor2on20480327681308.5916.95
10.0tensor2on20480327681299.5616.82
6.3.3tensor2on2048065536927.2825.62
7.2.4tensor2on2048065536852.723.70
7.14tensor2on2048065536965.354.44
10.0tensor2on2048065536960.414.07
6.3.3tensor2on0256064.370.95
7.2.4tensor2on0256072.600.52
7.14tensor2on0256075.431.77
10.0tensor2on0256076.420.80
6.3.3tensor2on02561638460.700.78
7.2.4tensor2on02561638468.210.92
7.14tensor2on02561638470.230.67
10.0tensor2on02561638470.520.75
6.3.3tensor2on02563276859.531.01
7.2.4tensor2on02563276867.100.71
7.14tensor2on02563276869.140.47
10.0tensor2on02563276869.160.84
6.3.3tensor2on02566553655.011.65
7.2.4tensor2on02566553664.390.90
7.14tensor2on02566553666.370.64
10.0tensor2on02566553664.114.85
6.3.3tensor4on2048002215.8147.42
7.2.4tensor4on2048002516.704.04
7.14tensor4on2048002725.8914.64
10.0tensor4on2048002734.3314.09
6.3.3tensor4on20480163841629.5213.16
7.2.4tensor4on20480163841756.2415.21
7.14tensor4on20480163841913.8034.39
10.0tensor4on20480163841930.4117.83
6.3.3tensor4on20480327681269.3986.38
7.2.4tensor4on20480327681363.2820.32
7.14tensor4on20480327681515.9616.34
10.0tensor4on20480327681513.6716.06
6.3.3tensor4on2048065536972.3611.10
7.2.4tensor4on2048065536949.679.65
7.14tensor4on20480655361054.849.65
10.0tensor4on20480655361058.208.35
6.3.3tensor4on0256053.821.46
7.2.4tensor4on0256063.641.22
7.14tensor4on0256067.280.93
10.0tensor4on0256067.021.31
6.3.3tensor4on02561638451.251.65
7.2.4tensor4on02561638459.741.66
7.14tensor4on02561638461.381.58
10.0tensor4on02561638461.941.11
6.3.3tensor4on02563276850.441.22
7.2.4tensor4on02563276859.270.96
7.14tensor4on02563276860.641.39
10.0tensor4on02563276861.001.04
6.3.3tensor4on02566553649.301.17
7.2.4tensor4on02566553657.251.06
7.14tensor4on02566553659.221.03
10.0tensor4on02566553657.811.63

gemma-4-31B-it​

unsloth/gemma-4-31B-it-GGUF:Q8_K_XL — dense 31B, Q8_K_XL, 32.6 GiB. tensor split over 4 and 2 GPUs, flash-attn on.

6.3.3tensor2on204800420.570.44
7.2.4tensor2on204800378.5919.45
7.14tensor2on204800434.2734.41
10.0tensor2on204800419.9511.35
6.3.3tensor2on2048016384308.929.60
7.2.4tensor2on2048016384288.611.68
7.14tensor2on2048016384324.701.34
10.0tensor2on2048016384326.500.57
6.3.3tensor2on2048032768263.832.78
7.2.4tensor2on2048032768227.264.35
7.14tensor2on2048032768262.263.36
10.0tensor2on2048032768254.974.22
6.3.3tensor2on2048065536187.8312.63
7.2.4tensor2on2048065536167.173.30
7.14tensor2on2048065536185.601.76
10.0tensor2on2048065536186.241.72
6.3.3tensor2on0256023.760.63
7.2.4tensor2on0256028.210.14
7.14tensor2on0256028.680.05
10.0tensor2on0256028.980.15
6.3.3tensor2on02561638422.900.45
7.2.4tensor2on02561638426.630.07
7.14tensor2on02561638426.470.10
10.0tensor2on02561638424.901.51
6.3.3tensor2on02563276821.681.64
7.2.4tensor2on02563276824.690.98
7.14tensor2on02563276825.760.10
10.0tensor2on02563276826.120.08
6.3.3tensor2on02566553620.941.48
7.2.4tensor2on02566553624.090.42
7.14tensor2on02566553623.410.31
10.0tensor2on02566553623.930.16
6.3.3tensor4on204800679.782.62
7.2.4tensor4on204800704.021.08
7.14tensor4on204800789.141.24
10.0tensor4on204800787.368.55
6.3.3tensor4on2048016384539.224.64
7.2.4tensor4on2048016384555.623.84
7.14tensor4on2048016384619.5037.49
10.0tensor4on2048016384629.575.47
6.3.3tensor4on2048032768489.9810.48
7.2.4tensor4on2048032768473.076.80
7.14tensor4on2048032768540.404.60
10.0tensor4on2048032768499.2539.31
6.3.3tensor4on2048065536361.7527.51
7.2.4tensor4on2048065536368.770.95
7.14tensor4on2048065536419.383.23
10.0tensor4on2048065536417.114.94
6.3.3tensor4on0256028.980.38
7.2.4tensor4on0256034.820.42
7.14tensor4on0256035.750.28
10.0tensor4on0256035.700.71
6.3.3tensor4on02561638426.870.79
7.2.4tensor4on02561638432.220.53
7.14tensor4on02561638433.010.52
10.0tensor4on02561638433.370.42
6.3.3tensor4on02563276827.060.44
7.2.4tensor4on02563276832.290.43
7.14tensor4on02563276831.180.71
10.0tensor4on02563276832.510.90
6.3.3tensor4on02566553624.483.70
7.2.4tensor4on02566553630.090.89
7.14tensor4on02566553630.100.30
10.0tensor4on02566553629.580.13