record(GEMMA4-ROCM-KEEP): measured SharedK-WMMA plateau on 2x R9700 - #676
record(GEMMA4-ROCM-KEEP): measured SharedK-WMMA plateau on 2x R9700#676bakon11 wants to merge 2 commits into
Conversation
Contributor closeout (PREFIX_CACHE=0, unique pads, 2026-08-13): 2014 t/s @~11k, 1099 t/s @~42k, decode ~55 t/s. Quality Paris/63/tool_calls. Vulkan Q8 on the same box is still ahead; this is the reliable recipe, not that bar. Speculative/ngram/FMHA/layer-split stay off. HIP cm1 spill is a residual, not a named LLVM defect. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: Hermes:grok-4.6 [Hermes]
Matched RADV pipeline stats on gfx1201 (Mesa 26.0.3): 0 spilled VGPR, 0 scratch at VGPR=256 for llama.cpp coopmat1 d=512. HIP cm1 still 235-339 spills. (a)-lean only; no component name, no 3.35x claim. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: Hermes:grok-4.6 [Hermes]
|
Thanks for this — and welcome. Reviewed as part of a sweep over the open external PRs. Before the findings, the part that matters: all eight checkers you list as passing locally do pass — I ran every one at your head SHA and confirmed it. Everything below is something no checker in this repo looks for, so none of it is a diligence failure on your part. Several of your calls are ones this project learned the hard way: disclosing the dropped 1170 first-rep with its cause rather than quietly taking the median, per-depth rep counts, explicitly writing "does not claim that bar", refusing to name an LLVM defect from a spill count, and not touching The blocker is that the recipe cannot be reproduced from Four of the decode knobs are read by no production code in this tree:
Meanwhile every prefill item — SharedK-WMMA on, FLASH off, I don't think the numbers are wrong. The likely story is honest measurements of your working branch written up against the wrong baseline. But a record whose whole value is reproducibility has to name the tree it describes, and the spec header still says Why this was easy to miss, and not your fault: Second blocker: no issue. AGENTS.md wants one linked in three places that agree — the roadmap issue table, the row's spec, and the PR body. The spec header cites On the denominator. The comparison arm is "Vulkan Q8" with the stack unnamed. The prior entry for this same box labelled it Your "does not claim that bar" framing is the right instinct and it is one sentence from being fully admissible. There's a template already in the tree at Placement. The measurement landed on
Two smaller things worth having:
Please keep the residual section as-is. The HIP cm1 339-spill vs ACO 0-spill contrast on the same silicon, with the explicit refusal to call it a named defect, is disciplined work and it should not be lost when the spec is compacted — I'd suggest an issue so it survives independently. Decode regime would help too: 55.5 / 49.1 t/s with no concurrency, depth or batch size can't be compared to anything, including a future vLLM leg. Your prefill depth curve is exactly right; decode just needs the same treatment. Nothing here needs new measurements. Happy to look again once the recipe names its tree. |
localai-org-maint-bot
left a comment
There was a problem hiding this comment.
The measured KEEP result needs the repository-required benchmark checkpoint before this can merge. This change publishes binding throughput numbers in FEATURES/ROCM/USAGE, but it does not update docs/BENCHMARKS.md and provides no committed command/log evidence anchor for reproducing the contributor-lab medians or the discarded outlier. Please add the exact workload/commands and evidence location to BENCHMARKS (or mark the numbers non-binding until that evidence exists), as required by AGENTS.md.
|
Following up on the review above — there is one question only you can answer, and it blocks the rest. The question: which tree produced these numbers?The KEEP recipe lists four decode knobs, and none of them is read by production code on
Meanwhile every prefill item — SharedK-WMMA on, FLASH off, I am not suggesting the measurements are wrong. The most likely explanation is that they are honest numbers from your working branch, written up against So: which revision were the 2112 / 2014 / 1705 / 1099 prefill figures and the 55.5 / 49.1 decode figures taken on? Once that is stated, there are two clean ways forward and either is fine:
The practical urgency is that Not your fault, and worth saying
The smaller items, unchanged from the earlier review
Please keep the residual section as-isThe HIP cm1 339-spill vs ACO 0-spill contrast on the same silicon, with the explicit refusal to call it a named LLVM defect, is the most reusable thing in this PR. I would suggest an issue for it so it survives the spec being compacted later. Nothing here needs new measurements — just the provenance question answered. |
What
Docs-only record of the contributor KEEP recipe for Gemma-4-26B FP8 on dual R9700 (gfx1201 / ROCm 7.2.4). No kernel or default-env change.
Fair protocol:
PREFIX_CACHE=0+ unique pads (2026-08-13).Decode stream ~55 t/s temp=0 / ~49 t/s temp=0.7. Paris / arith 63 /
gemma4tool_calls held.Same-box Vulkan Q8 unique-pad bar is still ~3503 @11k / ~2714 @42k. This PR names the reliable ROCm plateau; it does not claim that bar.
Out of the recipe
Speculative / ngram / FMHA / layer-split / Head-TP. Isolated P1 cm1 was ~1.13x KEEP (need ~3.35x isolated). HIP cm1 hsaco spills (339 vs KEEP 35) are a residual mechanism, not a named LLVM defect.
Files
.agents/specs/gemma4-rocm-fp8-moe.mdmeasured KEEP + rejected + residualdocs/USAGE.md/docs/ROCM.md/docs/FEATURES.md/docs/ENVIRONMENT.md(GEMM_Mlab note 2048)Local gates: device-leakage, doc-checkpoint, public-doc-tables, readme, env-doc, agent-record, pr-size, model-checklist.