mxxm fork vs ggml upstream
- fork:
mxxm-t/mx-llama.cppb10951(a2780ef),0.3.0-dev - upstream:
ggml-org/llama.cppv0.5.0(7fe450e),0.5.0-dev - toolchain:
ML-gfx906—7a0d586-pre - ROCm 6.3.3 and 10.0; the fork was not measured on 7.x
- same hardware, topology, HBM2 mem OC, models and profiles as ROCm 6.3.3 … 10.0 bench
- 5 llama-bench repetitions per point
- raw JSONL:
_data/ggmland_data/mxxm
The fork trails upstream: b10951 is upstream 5542318 (0.3.0-dev) plus one fork commit, while ggml is v0.5.0 (0.5.0-dev) three days later. Part of the delta below is upstream drift, not the fork.
Results
Δ is mxxm vs ggml (positive = fork faster). Full per-model tables: Qwen3.8-27B, Qwen3.8-Flash-Next, gemma-4-26B-A4B-it, gemma-4-31B-it.
Prompt 2048 / Gen 0 is prompt processing, Prompt 0 / Gen 256 is generation at the given Depth.
Fork vs upstream
Median Δ across configurations, per ROCm version. In parentheses: number of configs with a statistically significant change (more than 2σ of the combined measurement noise).
| Model | Test | 6.3.3 | 10.0 |
|---|---|---|---|
| Qwen3.8-27B | pp2048 | +6.7% (4/1) | +14.5% (7/0) |
| Qwen3.8-27B | tg256 | +47.7% (8/0) | +26.5% (7/0) |
| Qwen3.8-Flash-Next | pp2048 | -2.1% (0/1) | -16.6% (0/4) |
| Qwen3.8-Flash-Next | tg256 | -2.6% (0/3) | -3.7% (0/4) |
| gemma-4-26B-A4B-it | pp2048 | +1.5% (4/4) | +9.8% (9/0) |
| gemma-4-26B-A4B-it | tg256 | +49.0% (12/0) | +42.2% (12/0) |
| gemma-4-31B-it | pp2048 | +0.7% (3/1) | +17.2% (7/0) |
| gemma-4-31B-it | tg256 | +41.2% (8/0) | +31.7% (8/0) |
ROCm under the fork
Median Δ of mxxm on 10.0 relative to mxxm on 6.3.3.
| Model | Test | Δ 10.0 vs 6.3.3 |
|---|---|---|
| Qwen3.8-27B | pp2048 | +23.8% (7/0) |
| Qwen3.8-27B | tg256 | +7.9% (4/0) |
| Qwen3.8-Flash-Next | pp2048 | -7.1% (0/4) |
| Qwen3.8-Flash-Next | tg256 | +6.9% (4/0) |
| gemma-4-26B-A4B-it | pp2048 | +13.7% (11/0) |
| gemma-4-26B-A4B-it | tg256 | +12.2% (12/0) |
| gemma-4-31B-it | pp2048 | +20.4% (7/0) |
| gemma-4-31B-it | tg256 | +11.8% (7/0) |
Takeaways
- Q8_0 models love the fork. Generation is much faster on both ROCm versions: Qwen3.8-27B
+34…+63%on 6.3.3 and+8…+47%on 10.0; gemma-4-26B+14…+85%; gemma-4-31B+13…+57%. The biggest gains are tensor split on 4 GPUs, single-GPU layer split gains the least (+4…+18%). - Prompt processing wins mostly on 10.0: Qwen3.8-27B
+8…+29%, gemma-4-31B+2…+33%, gemma-4-26B+10…+28%(except 64K context, where tensor models are flat or slightly negative). On 6.3.3 prefill is mixed, including a−9%regression for single-GPU gemma-4-26B layer. - Qwen3.8-Flash-Next (Q4_K MoE) regresses: prompt processing
−11…−21%on 10.0 and−3…+1%on 6.3.3, generation−2.6…−4.6%on both. Known issue, the fork is not a win for this model yet. - Under the fork, ROCm 10.0 is the better base:
+14…+24%prompt and+7…+12%generation vs 6.3.3 (Flash-Next prompt is the exception,−7%). This is a stronger preference than with ggml, where 7.14/10.0 are near parity.
Builds and env
| Build | Repo | Ref | Version | Image |
|---|---|---|---|---|
| ggml | ggml-org/llama.cpp | v0.5.0 7fe450e | 0.5.0-dev | v0.5.0-rocm-<ver>-7a0d586-pre |
| mxxm | mxxm-t/mx-llama.cpp | b10951 a2780ef | 0.3.0-dev | b10951-rocm-<ver>-801bf2e-pre |
Both are built with the same Dockerfile and cmake flags (see Software); HIP flags differ per ROCm only: 6.3.3 has none, 10.0 uses -mllvm -amdgpu-sched-strategy=max-ilp.
Environment declared in the bench profiles and passed to llama-bench:
| Variable | ggml | mxxm | Note |
|---|---|---|---|
HIP_VISIBLE_DEVICES | 1,0,3,2 | 1,0,3,2 | |
HSA_FORCE_FINE_GRAIN_PCIE | 1 | 1 | |
HSA_XNACK | 0 | 0 | |
GGML_ENABLE_CUSTOM_AR | — | 1 | mxxm only |
GGML_CUDA_REPACK_Q8_0 | — | 1 | mxxm only |
LLAMA_ENABLE_MTP_OPT | — | 1 | mxxm only |
GGML_TP_AR_TWOSHOT | — | 1 | mxxm only |
LLAMA_PLE_SHARD | — | 0 | mxxm only |
GPU_MAX_HW_QUEUES | — | 8 | mxxm only |
The mxxm-only variables are no-ops in ggml.
Bench
Same run.sh for both builds: ROCM_VER selects the result directory, PRESETS lists the models.
PRESETS=(
'gemma-4-26B-A4B-it[tensor]'
'gemma-4-26B-A4B-it[layer]'
'gemma-4-31B-it[tensor]'
'Qwen3.8-27B[tensor]'
'Qwen3.8-Flash-Next[layer]'
)
for (( i=0; i<${#PRESETS[@]}; i++ )); do
echo "======================= ${PRESETS[$i]} ======================="
python3 ./bench-profiles/build-bench-command.py --source-dir ./bench-profiles --bench "${PRESETS[$i]}" --output - | bash
done
- global.yaml (ggml)
- global.yaml (mxxm)
- Qwen3.8.yaml
- gemma-4-26b-a4b-it.yaml
- gemma-4-31b-it.yaml
"*":
command: |
#!/usr/bin/env bash
set -eo pipefail
mkdir -p ./bench-results/rocm-$ROCM_VER
echo $ROCM_VER > ./bench-results/rocm-$ROCM_VER/rocm-ver
{env} ./llama-bench {args} | tee ./bench-results/rocm-$ROCM_VER/{profile}.jsonl
env:
HIP_VISIBLE_DEVICES: "1,0,3,2"
HSA_FORCE_FINE_GRAIN_PCIE: "1"
HSA_XNACK: "0"
args:
lazy-mode: "off"
offline: true
output: jsonl
verbose: false
progress: true
"*":
command: |
#!/usr/bin/env bash
set -eo pipefail
mkdir -p ./bench-results/rocm-$ROCM_VER
echo $ROCM_VER > ./bench-results/rocm-$ROCM_VER/rocm-ver
{env} ./llama-bench {args} | tee ./bench-results/rocm-$ROCM_VER/{profile}.jsonl
env:
HIP_VISIBLE_DEVICES: "1,0,3,2"
HSA_FORCE_FINE_GRAIN_PCIE: "1"
HSA_XNACK: "0"
GGML_ENABLE_CUSTOM_AR: "1"
GGML_CUDA_REPACK_Q8_0: "1"
LLAMA_ENABLE_MTP_OPT: "1"
GGML_TP_AR_TWOSHOT: "1"
LLAMA_PLE_SHARD: "0"
GPU_MAX_HW_QUEUES: "8"
args:
lazy-mode: "off"
offline: true
output: jsonl
verbose: false
progress: true
Qwen3.8-27B[tensor]:
args:
hf-repo: bartowski/Qwen3.8-27B-GGUF:Q8_0
flash-attn: true
split-mode: tensor
device:
- [ROCm0, ROCm1, ROCm2, ROCm3]
- [ROCm0, ROCm1]
n-prompt: 2048
ubatch-size: 2048
n-gen: 256
n-depth: [0, 16384, 32768, 65536]
Qwen3.8-Flash-Next[layer]:
args:
hf-repo: bartowski/Qwen3.8-Flash-Next-GGUF:Q4_K_M
flash-attn: auto
split-mode: layer
device:
- [ROCm0, ROCm1, ROCm2, ROCm3]
n-prompt: 2048
ubatch-size: 2048
n-gen: 256
n-depth: [0, 16384, 32768, 65536]
gemma-4-26B-A4B-it[tensor]:
args:
hf-repo: bartowski/google_gemma-4-26B-A4B-it-GGUF:Q8_0
flash-attn: true
split-mode: tensor
device:
- [ROCm0, ROCm1, ROCm2, ROCm3]
- [ROCm0, ROCm1]
n-prompt: 2048
ubatch-size: 2048
n-gen: 256
n-depth: [0, 16384, 32768, 65536]
gemma-4-26B-A4B-it[layer]:
args:
hf-repo: bartowski/google_gemma-4-26B-A4B-it-GGUF:Q8_0
flash-attn: true
split-mode: layer
device:
- [ROCm0]
n-prompt: 2048
ubatch-size: 2048
n-gen: 256
n-depth: [0, 16384, 32768, 65536]
gemma-4-26B-A4B-it[cpu]:
env:
HIP_VISIBLE_DEVICES: ""
args:
hf-repo: bartowski/google_gemma-4-26B-A4B-it-GGUF:Q8_0
split-mode: layer
device: none
n-prompt: 2048
ubatch-size: 2048
n-gen: 256
n-depth: [0, 16384]
threads: 32
gemma-4-31B-it[tensor]:
args:
hf-repo: unsloth/gemma-4-31B-it-GGUF:Q8_K_XL
flash-attn: true
split-mode: [tensor]
device:
- [ROCm0, ROCm1, ROCm2, ROCm3]
- [ROCm0, ROCm1]
n-prompt: 2048
ubatch-size: 2048
n-gen: 256
n-depth: [0, 16384, 32768, 65536]
Model profiles are identical for both builds, only global.yaml differs.
Qwen3.8-27B
bartowski/Qwen3.8-27B-GGUF:Q8_0 — dense 27B, Q8_0, 27.1 GiB. tensor split over 4 and 2 GPUs, flash-attn on.
| 6.3.3 | tensor | 2 | 2048 | 0 | 0 | 455.53 | 495.27 | +8.72 |
| 10.0 | tensor | 2 | 2048 | 0 | 0 | 517.84 | 668.62 | +29.12 |
| 6.3.3 | tensor | 2 | 2048 | 0 | 16384 | 408.98 | 417.47 | +2.07 |
| 10.0 | tensor | 2 | 2048 | 0 | 16384 | 441.31 | 510.03 | +15.57 |
| 6.3.3 | tensor | 2 | 2048 | 0 | 32768 | 355.32 | 346.05 | -2.61 |
| 10.0 | tensor | 2 | 2048 | 0 | 32768 | 393.69 | 433.94 | +10.23 |
| 6.3.3 | tensor | 2 | 2048 | 0 | 65536 | 254.94 | 266.77 | +4.64 |
| 10.0 | tensor | 2 | 2048 | 0 | 65536 | 303.13 | 327.73 | +8.11 |
| 6.3.3 | tensor | 2 | 0 | 256 | 0 | 27.10 | 38.68 | +42.74 |
| 10.0 | tensor | 2 | 0 | 256 | 0 | 34.34 | 40.92 | +19.19 |
| 6.3.3 | tensor | 2 | 0 | 256 | 16384 | 26.82 | 36.81 | +37.22 |
| 10.0 | tensor | 2 | 0 | 256 | 16384 | 31.94 | 38.51 | +20.54 |
| 6.3.3 | tensor | 2 | 0 | 256 | 32768 | 25.90 | 34.69 | +33.93 |
| 10.0 | tensor | 2 | 0 | 256 | 32768 | 30.75 | 33.33 | +8.38 |
| 6.3.3 | tensor | 2 | 0 | 256 | 65536 | 20.50 | 30.45 | +48.55 |
| 10.0 | tensor | 2 | 0 | 256 | 65536 | 28.00 | 31.95 | +14.11 |
| 6.3.3 | tensor | 4 | 2048 | 0 | 0 | 704.48 | 809.53 | +14.91 |
| 10.0 | tensor | 4 | 2048 | 0 | 0 | 862.82 | 1064.62 | +23.39 |
| 6.3.3 | tensor | 4 | 2048 | 0 | 16384 | 635.69 | 691.03 | +8.71 |
| 10.0 | tensor | 4 | 2048 | 0 | 16384 | 739.79 | 861.84 | +16.50 |
| 6.3.3 | tensor | 4 | 2048 | 0 | 32768 | 521.95 | 604.72 | +15.86 |
| 10.0 | tensor | 4 | 2048 | 0 | 32768 | 652.02 | 739.12 | +13.36 |
| 6.3.3 | tensor | 4 | 2048 | 0 | 65536 | 457.83 | 444.80 | -2.84 |
| 10.0 | tensor | 4 | 2048 | 0 | 65536 | 525.52 | 495.72 | -5.67 |
| 6.3.3 | tensor | 4 | 0 | 256 | 0 | 30.35 | 48.41 | +59.53 |
| 10.0 | tensor | 4 | 0 | 256 | 0 | 40.06 | 57.93 | +44.62 |
| 6.3.3 | tensor | 4 | 0 | 256 | 16384 | 28.71 | 46.80 | +63.01 |
| 10.0 | tensor | 4 | 0 | 256 | 16384 | 38.83 | 51.45 | +32.50 |
| 6.3.3 | tensor | 4 | 0 | 256 | 32768 | 30.01 | 44.05 | +46.79 |
| 10.0 | tensor | 4 | 0 | 256 | 32768 | 34.49 | 50.77 | +47.19 |
| 6.3.3 | tensor | 4 | 0 | 256 | 65536 | 25.75 | 40.90 | +58.84 |
| 10.0 | tensor | 4 | 0 | 256 | 65536 | 33.67 | 46.14 | +37.03 |
Qwen3.8-Flash-Next
bartowski/Qwen3.8-Flash-Next-GGUF:Q4_K_M — MoE A3B, Q4_K_M, ~111 GiB. layer split over 4 GPUs, flash-attn auto. The regression case: the fork is slower on every point.
| 6.3.3 | layer | 4 | 2048 | 0 | 0 | 513.24 | 497.37 | -3.09 |
| 10.0 | layer | 4 | 2048 | 0 | 0 | 553.18 | 438.35 | -20.76 |
| 6.3.3 | layer | 4 | 2048 | 0 | 16384 | 383.61 | 378.17 | -1.42 |
| 10.0 | layer | 4 | 2048 | 0 | 16384 | 417.79 | 341.27 | -18.31 |
| 6.3.3 | layer | 4 | 2048 | 0 | 32768 | 309.66 | 301.33 | -2.69 |
| 10.0 | layer | 4 | 2048 | 0 | 32768 | 338.30 | 287.90 | -14.90 |
| 6.3.3 | layer | 4 | 2048 | 0 | 65536 | 218.48 | 220.64 | +0.99 |
| 10.0 | layer | 4 | 2048 | 0 | 65536 | 236.51 | 210.86 | -10.85 |
| 6.3.3 | layer | 4 | 0 | 256 | 0 | 32.96 | 32.15 | -2.44 |
| 10.0 | layer | 4 | 0 | 256 | 0 | 35.17 | 33.86 | -3.72 |
| 6.3.3 | layer | 4 | 0 | 256 | 16384 | 26.59 | 25.86 | -2.75 |
| 10.0 | layer | 4 | 0 | 256 | 16384 | 28.63 | 27.55 | -3.78 |
| 6.3.3 | layer | 4 | 0 | 256 | 32768 | 22.23 | 22.34 | +0.48 |
| 10.0 | layer | 4 | 0 | 256 | 32768 | 24.86 | 23.98 | -3.54 |
| 6.3.3 | layer | 4 | 0 | 256 | 65536 | 18.56 | 17.71 | -4.58 |
| 10.0 | layer | 4 | 0 | 256 | 65536 | 20.07 | 19.17 | -4.49 |
gemma-4-26B-A4B-it
bartowski/google_gemma-4-26B-A4B-it-GGUF:Q8_0 — MoE A4B, Q8_0, 25.0 GiB. tensor split over 4 and 2 GPUs plus layer split on 1 GPU, flash-attn on. Both presets share this table, Split/GPUs columns tell them apart.
| 6.3.3 | layer | 1 | 2048 | 0 | 0 | 1497.42 | 1361.21 | -9.10 |
| 10.0 | layer | 1 | 2048 | 0 | 0 | 1580.72 | 2018.69 | +27.71 |
| 6.3.3 | layer | 1 | 2048 | 0 | 16384 | 951.00 | 880.62 | -7.40 |
| 10.0 | layer | 1 | 2048 | 0 | 16384 | 961.58 | 1088.64 | +13.21 |
| 6.3.3 | layer | 1 | 2048 | 0 | 32768 | 726.47 | 692.65 | -4.66 |
| 10.0 | layer | 1 | 2048 | 0 | 32768 | 714.15 | 779.75 | +9.19 |
| 6.3.3 | layer | 1 | 2048 | 0 | 65536 | 492.73 | 480.41 | -2.50 |
| 10.0 | layer | 1 | 2048 | 0 | 65536 | 473.19 | 507.88 | +7.33 |
| 6.3.3 | layer | 1 | 0 | 256 | 0 | 78.58 | 92.58 | +17.81 |
| 10.0 | layer | 1 | 0 | 256 | 0 | 95.19 | 98.72 | +3.72 |
| 6.3.3 | layer | 1 | 0 | 256 | 16384 | 72.27 | 82.69 | +14.42 |
| 10.0 | layer | 1 | 0 | 256 | 16384 | 85.74 | 89.11 | +3.92 |
| 6.3.3 | layer | 1 | 0 | 256 | 32768 | 69.47 | 80.18 | +15.42 |
| 10.0 | layer | 1 | 0 | 256 | 32768 | 81.38 | 85.22 | +4.72 |
| 6.3.3 | layer | 1 | 0 | 256 | 65536 | 63.43 | 72.63 | +14.50 |
| 10.0 | layer | 1 | 0 | 256 | 65536 | 75.04 | 77.89 | +3.79 |
| 6.3.3 | tensor | 2 | 2048 | 0 | 0 | 1997.36 | 2082.76 | +4.28 |
| 10.0 | tensor | 2 | 2048 | 0 | 0 | 2106.77 | 2641.43 | +25.38 |
| 6.3.3 | tensor | 2 | 2048 | 0 | 16384 | 1525.91 | 1534.88 | +0.59 |
| 10.0 | tensor | 2 | 2048 | 0 | 16384 | 1581.45 | 1762.09 | +11.42 |
| 6.3.3 | tensor | 2 | 2048 | 0 | 32768 | 1184.74 | 1213.81 | +2.45 |
| 10.0 | tensor | 2 | 2048 | 0 | 32768 | 1299.56 | 1424.05 | +9.58 |
| 6.3.3 | tensor | 2 | 2048 | 0 | 65536 | 927.28 | 884.88 | -4.57 |
| 10.0 | tensor | 2 | 2048 | 0 | 65536 | 960.41 | 965.85 | +0.57 |
| 6.3.3 | tensor | 2 | 0 | 256 | 0 | 64.37 | 98.87 | +53.60 |
| 10.0 | tensor | 2 | 0 | 256 | 0 | 76.42 | 112.58 | +47.33 |
| 6.3.3 | tensor | 2 | 0 | 256 | 16384 | 60.70 | 90.55 | +49.18 |
| 10.0 | tensor | 2 | 0 | 256 | 16384 | 70.52 | 100.39 | +42.36 |
| 6.3.3 | tensor | 2 | 0 | 256 | 32768 | 59.53 | 88.57 | +48.79 |
| 10.0 | tensor | 2 | 0 | 256 | 32768 | 69.16 | 97.53 | +41.01 |
| 6.3.3 | tensor | 2 | 0 | 256 | 65536 | 55.01 | 80.27 | +45.91 |
| 10.0 | tensor | 2 | 0 | 256 | 65536 | 64.11 | 91.09 | +42.10 |
| 6.3.3 | tensor | 4 | 2048 | 0 | 0 | 2215.81 | 2649.01 | +19.55 |
| 10.0 | tensor | 4 | 2048 | 0 | 0 | 2734.33 | 3224.73 | +17.94 |
| 6.3.3 | tensor | 4 | 2048 | 0 | 16384 | 1629.52 | 1891.07 | +16.05 |
| 10.0 | tensor | 4 | 2048 | 0 | 16384 | 1930.41 | 2124.30 | +10.04 |
| 6.3.3 | tensor | 4 | 2048 | 0 | 32768 | 1269.39 | 1455.68 | +14.68 |
| 10.0 | tensor | 4 | 2048 | 0 | 32768 | 1513.67 | 1603.07 | +5.91 |
| 6.3.3 | tensor | 4 | 2048 | 0 | 65536 | 972.36 | 1031.39 | +6.07 |
| 10.0 | tensor | 4 | 2048 | 0 | 65536 | 1058.20 | 1039.21 | -1.79 |
| 6.3.3 | tensor | 4 | 0 | 256 | 0 | 53.82 | 88.52 | +64.48 |
| 10.0 | tensor | 4 | 0 | 256 | 0 | 67.02 | 123.68 | +84.53 |
| 6.3.3 | tensor | 4 | 0 | 256 | 16384 | 51.25 | 84.10 | +64.10 |
| 10.0 | tensor | 4 | 0 | 256 | 16384 | 61.94 | 110.09 | +77.74 |
| 6.3.3 | tensor | 4 | 0 | 256 | 32768 | 50.44 | 87.29 | +73.05 |
| 10.0 | tensor | 4 | 0 | 256 | 32768 | 61.00 | 106.54 | +74.66 |
| 6.3.3 | tensor | 4 | 0 | 256 | 65536 | 49.30 | 78.45 | +59.14 |
| 10.0 | tensor | 4 | 0 | 256 | 65536 | 57.81 | 94.09 | +62.76 |
gemma-4-31B-it
unsloth/gemma-4-31B-it-GGUF:Q8_K_XL — dense 31B, Q8_K_XL, 32.6 GiB. tensor split over 4 and 2 GPUs, flash-attn on.
| 6.3.3 | tensor | 2 | 2048 | 0 | 0 | 420.57 | 411.79 | -2.09 |
| 10.0 | tensor | 2 | 2048 | 0 | 0 | 419.95 | 560.08 | +33.37 |
| 6.3.3 | tensor | 2 | 2048 | 0 | 16384 | 308.92 | 312.41 | +1.13 |
| 10.0 | tensor | 2 | 2048 | 0 | 16384 | 326.50 | 359.76 | +10.19 |
| 6.3.3 | tensor | 2 | 2048 | 0 | 32768 | 263.83 | 259.04 | -1.81 |
| 10.0 | tensor | 2 | 2048 | 0 | 32768 | 254.97 | 295.43 | +15.87 |
| 6.3.3 | tensor | 2 | 2048 | 0 | 65536 | 187.83 | 188.43 | +0.32 |
| 10.0 | tensor | 2 | 2048 | 0 | 65536 | 186.24 | 201.35 | +8.11 |
| 6.3.3 | tensor | 2 | 0 | 256 | 0 | 23.76 | 32.79 | +38.00 |
| 10.0 | tensor | 2 | 0 | 256 | 0 | 28.98 | 32.82 | +13.26 |
| 6.3.3 | tensor | 2 | 0 | 256 | 16384 | 22.90 | 30.84 | +34.70 |
| 10.0 | tensor | 2 | 0 | 256 | 16384 | 24.90 | 31.71 | +27.36 |
| 6.3.3 | tensor | 2 | 0 | 256 | 32768 | 21.68 | 27.56 | +27.10 |
| 10.0 | tensor | 2 | 0 | 256 | 32768 | 26.12 | 30.76 | +17.76 |
| 6.3.3 | tensor | 2 | 0 | 256 | 65536 | 20.94 | 24.73 | +18.06 |
| 10.0 | tensor | 2 | 0 | 256 | 65536 | 23.93 | 28.94 | +20.94 |
| 6.3.3 | tensor | 4 | 2048 | 0 | 0 | 679.78 | 717.19 | +5.50 |
| 10.0 | tensor | 4 | 2048 | 0 | 0 | 787.36 | 977.56 | +24.16 |
| 6.3.3 | tensor | 4 | 2048 | 0 | 16384 | 539.22 | 594.53 | +10.26 |
| 10.0 | tensor | 4 | 2048 | 0 | 16384 | 629.57 | 746.45 | +18.56 |
| 6.3.3 | tensor | 4 | 2048 | 0 | 32768 | 489.98 | 451.27 | -7.90 |
| 10.0 | tensor | 4 | 2048 | 0 | 32768 | 499.25 | 619.21 | +24.03 |
| 6.3.3 | tensor | 4 | 2048 | 0 | 65536 | 361.75 | 401.71 | +11.05 |
| 10.0 | tensor | 4 | 2048 | 0 | 65536 | 417.11 | 424.87 | +1.86 |
| 6.3.3 | tensor | 4 | 0 | 256 | 0 | 28.98 | 45.56 | +57.21 |
| 10.0 | tensor | 4 | 0 | 256 | 0 | 35.70 | 52.32 | +46.55 |
| 6.3.3 | tensor | 4 | 0 | 256 | 16384 | 26.87 | 42.17 | +56.96 |
| 10.0 | tensor | 4 | 0 | 256 | 16384 | 33.37 | 45.80 | +37.22 |
| 6.3.3 | tensor | 4 | 0 | 256 | 32768 | 27.06 | 39.47 | +45.84 |
| 10.0 | tensor | 4 | 0 | 256 | 32768 | 32.51 | 44.20 | +35.96 |
| 6.3.3 | tensor | 4 | 0 | 256 | 65536 | 24.48 | 35.35 | +44.39 |
| 10.0 | tensor | 4 | 0 | 256 | 65536 | 29.58 | 42.72 | +44.44 |