Skip to main content

mxxm fork vs ggml upstream

base skew

The fork trails upstream: b10951 is upstream 5542318 (0.3.0-dev) plus one fork commit, while ggml is v0.5.0 (0.5.0-dev) three days later. Part of the delta below is upstream drift, not the fork.

Results​

Δ is mxxm vs ggml (positive = fork faster). Full per-model tables: Qwen3.8-27B, Qwen3.8-Flash-Next, gemma-4-26B-A4B-it, gemma-4-31B-it.

Prompt 2048 / Gen 0 is prompt processing, Prompt 0 / Gen 256 is generation at the given Depth.

Fork vs upstream​

Median Δ across configurations, per ROCm version. In parentheses: number of configs with a statistically significant change (more than 2σ of the combined measurement noise).

ModelTest6.3.310.0
Qwen3.8-27Bpp2048+6.7% (4/1)+14.5% (7/0)
Qwen3.8-27Btg256+47.7% (8/0)+26.5% (7/0)
Qwen3.8-Flash-Nextpp2048-2.1% (0/1)-16.6% (0/4)
Qwen3.8-Flash-Nexttg256-2.6% (0/3)-3.7% (0/4)
gemma-4-26B-A4B-itpp2048+1.5% (4/4)+9.8% (9/0)
gemma-4-26B-A4B-ittg256+49.0% (12/0)+42.2% (12/0)
gemma-4-31B-itpp2048+0.7% (3/1)+17.2% (7/0)
gemma-4-31B-ittg256+41.2% (8/0)+31.7% (8/0)

ROCm under the fork​

Median Δ of mxxm on 10.0 relative to mxxm on 6.3.3.

ModelTestΔ 10.0 vs 6.3.3
Qwen3.8-27Bpp2048+23.8% (7/0)
Qwen3.8-27Btg256+7.9% (4/0)
Qwen3.8-Flash-Nextpp2048-7.1% (0/4)
Qwen3.8-Flash-Nexttg256+6.9% (4/0)
gemma-4-26B-A4B-itpp2048+13.7% (11/0)
gemma-4-26B-A4B-ittg256+12.2% (12/0)
gemma-4-31B-itpp2048+20.4% (7/0)
gemma-4-31B-ittg256+11.8% (7/0)

Takeaways​

  • Q8_0 models love the fork. Generation is much faster on both ROCm versions: Qwen3.8-27B +34…+63% on 6.3.3 and +8…+47% on 10.0; gemma-4-26B +14…+85%; gemma-4-31B +13…+57%. The biggest gains are tensor split on 4 GPUs, single-GPU layer split gains the least (+4…+18%).
  • Prompt processing wins mostly on 10.0: Qwen3.8-27B +8…+29%, gemma-4-31B +2…+33%, gemma-4-26B +10…+28% (except 64K context, where tensor models are flat or slightly negative). On 6.3.3 prefill is mixed, including a −9% regression for single-GPU gemma-4-26B layer.
  • Qwen3.8-Flash-Next (Q4_K MoE) regresses: prompt processing −11…−21% on 10.0 and −3…+1% on 6.3.3, generation −2.6…−4.6% on both. Known issue, the fork is not a win for this model yet.
  • Under the fork, ROCm 10.0 is the better base: +14…+24% prompt and +7…+12% generation vs 6.3.3 (Flash-Next prompt is the exception, −7%). This is a stronger preference than with ggml, where 7.14/10.0 are near parity.

Builds and env​

BuildRepoRefVersionImage
ggmlggml-org/llama.cppv0.5.0 7fe450e0.5.0-devv0.5.0-rocm-<ver>-7a0d586-pre
mxxmmxxm-t/mx-llama.cppb10951 a2780ef0.3.0-devb10951-rocm-<ver>-801bf2e-pre

Both are built with the same Dockerfile and cmake flags (see Software); HIP flags differ per ROCm only: 6.3.3 has none, 10.0 uses -mllvm -amdgpu-sched-strategy=max-ilp.

Environment declared in the bench profiles and passed to llama-bench:

VariableggmlmxxmNote
HIP_VISIBLE_DEVICES1,0,3,21,0,3,2
HSA_FORCE_FINE_GRAIN_PCIE11
HSA_XNACK00
GGML_ENABLE_CUSTOM_AR—1mxxm only
GGML_CUDA_REPACK_Q8_0—1mxxm only
LLAMA_ENABLE_MTP_OPT—1mxxm only
GGML_TP_AR_TWOSHOT—1mxxm only
LLAMA_PLE_SHARD—0mxxm only
GPU_MAX_HW_QUEUES—8mxxm only

The mxxm-only variables are no-ops in ggml.

Bench​

Same run.sh for both builds: ROCM_VER selects the result directory, PRESETS lists the models.

run.sh
PRESETS=(
'gemma-4-26B-A4B-it[tensor]'
'gemma-4-26B-A4B-it[layer]'
'gemma-4-31B-it[tensor]'
'Qwen3.8-27B[tensor]'
'Qwen3.8-Flash-Next[layer]'
)

for (( i=0; i<${#PRESETS[@]}; i++ )); do
echo "======================= ${PRESETS[$i]} ======================="
python3 ./bench-profiles/build-bench-command.py --source-dir ./bench-profiles --bench "${PRESETS[$i]}" --output - | bash
done
bench-profiles/global.yaml
"*":
command: |
#!/usr/bin/env bash
set -eo pipefail
mkdir -p ./bench-results/rocm-$ROCM_VER
echo $ROCM_VER > ./bench-results/rocm-$ROCM_VER/rocm-ver
{env} ./llama-bench {args} | tee ./bench-results/rocm-$ROCM_VER/{profile}.jsonl
env:
HIP_VISIBLE_DEVICES: "1,0,3,2"
HSA_FORCE_FINE_GRAIN_PCIE: "1"
HSA_XNACK: "0"
args:
lazy-mode: "off"
offline: true
output: jsonl
verbose: false
progress: true

Model profiles are identical for both builds, only global.yaml differs.

Qwen3.8-27B​

bartowski/Qwen3.8-27B-GGUF:Q8_0 — dense 27B, Q8_0, 27.1 GiB. tensor split over 4 and 2 GPUs, flash-attn on.

6.3.3tensor2204800455.53495.27+8.72
10.0tensor2204800517.84668.62+29.12
6.3.3tensor22048016384408.98417.47+2.07
10.0tensor22048016384441.31510.03+15.57
6.3.3tensor22048032768355.32346.05-2.61
10.0tensor22048032768393.69433.94+10.23
6.3.3tensor22048065536254.94266.77+4.64
10.0tensor22048065536303.13327.73+8.11
6.3.3tensor20256027.1038.68+42.74
10.0tensor20256034.3440.92+19.19
6.3.3tensor202561638426.8236.81+37.22
10.0tensor202561638431.9438.51+20.54
6.3.3tensor202563276825.9034.69+33.93
10.0tensor202563276830.7533.33+8.38
6.3.3tensor202566553620.5030.45+48.55
10.0tensor202566553628.0031.95+14.11
6.3.3tensor4204800704.48809.53+14.91
10.0tensor4204800862.821064.62+23.39
6.3.3tensor42048016384635.69691.03+8.71
10.0tensor42048016384739.79861.84+16.50
6.3.3tensor42048032768521.95604.72+15.86
10.0tensor42048032768652.02739.12+13.36
6.3.3tensor42048065536457.83444.80-2.84
10.0tensor42048065536525.52495.72-5.67
6.3.3tensor40256030.3548.41+59.53
10.0tensor40256040.0657.93+44.62
6.3.3tensor402561638428.7146.80+63.01
10.0tensor402561638438.8351.45+32.50
6.3.3tensor402563276830.0144.05+46.79
10.0tensor402563276834.4950.77+47.19
6.3.3tensor402566553625.7540.90+58.84
10.0tensor402566553633.6746.14+37.03

Qwen3.8-Flash-Next​

bartowski/Qwen3.8-Flash-Next-GGUF:Q4_K_M — MoE A3B, Q4_K_M, ~111 GiB. layer split over 4 GPUs, flash-attn auto. The regression case: the fork is slower on every point.

6.3.3layer4204800513.24497.37-3.09
10.0layer4204800553.18438.35-20.76
6.3.3layer42048016384383.61378.17-1.42
10.0layer42048016384417.79341.27-18.31
6.3.3layer42048032768309.66301.33-2.69
10.0layer42048032768338.30287.90-14.90
6.3.3layer42048065536218.48220.64+0.99
10.0layer42048065536236.51210.86-10.85
6.3.3layer40256032.9632.15-2.44
10.0layer40256035.1733.86-3.72
6.3.3layer402561638426.5925.86-2.75
10.0layer402561638428.6327.55-3.78
6.3.3layer402563276822.2322.34+0.48
10.0layer402563276824.8623.98-3.54
6.3.3layer402566553618.5617.71-4.58
10.0layer402566553620.0719.17-4.49

gemma-4-26B-A4B-it​

bartowski/google_gemma-4-26B-A4B-it-GGUF:Q8_0 — MoE A4B, Q8_0, 25.0 GiB. tensor split over 4 and 2 GPUs plus layer split on 1 GPU, flash-attn on. Both presets share this table, Split/GPUs columns tell them apart.

6.3.3layer12048001497.421361.21-9.10
10.0layer12048001580.722018.69+27.71
6.3.3layer12048016384951.00880.62-7.40
10.0layer12048016384961.581088.64+13.21
6.3.3layer12048032768726.47692.65-4.66
10.0layer12048032768714.15779.75+9.19
6.3.3layer12048065536492.73480.41-2.50
10.0layer12048065536473.19507.88+7.33
6.3.3layer10256078.5892.58+17.81
10.0layer10256095.1998.72+3.72
6.3.3layer102561638472.2782.69+14.42
10.0layer102561638485.7489.11+3.92
6.3.3layer102563276869.4780.18+15.42
10.0layer102563276881.3885.22+4.72
6.3.3layer102566553663.4372.63+14.50
10.0layer102566553675.0477.89+3.79
6.3.3tensor22048001997.362082.76+4.28
10.0tensor22048002106.772641.43+25.38
6.3.3tensor220480163841525.911534.88+0.59
10.0tensor220480163841581.451762.09+11.42
6.3.3tensor220480327681184.741213.81+2.45
10.0tensor220480327681299.561424.05+9.58
6.3.3tensor22048065536927.28884.88-4.57
10.0tensor22048065536960.41965.85+0.57
6.3.3tensor20256064.3798.87+53.60
10.0tensor20256076.42112.58+47.33
6.3.3tensor202561638460.7090.55+49.18
10.0tensor202561638470.52100.39+42.36
6.3.3tensor202563276859.5388.57+48.79
10.0tensor202563276869.1697.53+41.01
6.3.3tensor202566553655.0180.27+45.91
10.0tensor202566553664.1191.09+42.10
6.3.3tensor42048002215.812649.01+19.55
10.0tensor42048002734.333224.73+17.94
6.3.3tensor420480163841629.521891.07+16.05
10.0tensor420480163841930.412124.30+10.04
6.3.3tensor420480327681269.391455.68+14.68
10.0tensor420480327681513.671603.07+5.91
6.3.3tensor42048065536972.361031.39+6.07
10.0tensor420480655361058.201039.21-1.79
6.3.3tensor40256053.8288.52+64.48
10.0tensor40256067.02123.68+84.53
6.3.3tensor402561638451.2584.10+64.10
10.0tensor402561638461.94110.09+77.74
6.3.3tensor402563276850.4487.29+73.05
10.0tensor402563276861.00106.54+74.66
6.3.3tensor402566553649.3078.45+59.14
10.0tensor402566553657.8194.09+62.76

gemma-4-31B-it​

unsloth/gemma-4-31B-it-GGUF:Q8_K_XL — dense 31B, Q8_K_XL, 32.6 GiB. tensor split over 4 and 2 GPUs, flash-attn on.

6.3.3tensor2204800420.57411.79-2.09
10.0tensor2204800419.95560.08+33.37
6.3.3tensor22048016384308.92312.41+1.13
10.0tensor22048016384326.50359.76+10.19
6.3.3tensor22048032768263.83259.04-1.81
10.0tensor22048032768254.97295.43+15.87
6.3.3tensor22048065536187.83188.43+0.32
10.0tensor22048065536186.24201.35+8.11
6.3.3tensor20256023.7632.79+38.00
10.0tensor20256028.9832.82+13.26
6.3.3tensor202561638422.9030.84+34.70
10.0tensor202561638424.9031.71+27.36
6.3.3tensor202563276821.6827.56+27.10
10.0tensor202563276826.1230.76+17.76
6.3.3tensor202566553620.9424.73+18.06
10.0tensor202566553623.9328.94+20.94
6.3.3tensor4204800679.78717.19+5.50
10.0tensor4204800787.36977.56+24.16
6.3.3tensor42048016384539.22594.53+10.26
10.0tensor42048016384629.57746.45+18.56
6.3.3tensor42048032768489.98451.27-7.90
10.0tensor42048032768499.25619.21+24.03
6.3.3tensor42048065536361.75401.71+11.05
10.0tensor42048065536417.11424.87+1.86
6.3.3tensor40256028.9845.56+57.21
10.0tensor40256035.7052.32+46.55
6.3.3tensor402561638426.8742.17+56.96
10.0tensor402561638433.3745.80+37.22
6.3.3tensor402563276827.0639.47+45.84
10.0tensor402563276832.5144.20+35.96
6.3.3tensor402566553624.4835.35+44.39
10.0tensor402566553629.5842.72+44.44