Strata on RX 6800: faster Qwen3.8-Flash-Next prompt processing

rx6800local-llmcoding-agentsstrata

A block of layered rock on a dark desk, its teal and amber layers holding small glowing cubes, connected by a thin teal light to a graphics card lying beside it

Strata runs Qwen3.8-Flash-Next, a 125-billion-parameter mixture-of-experts model, using one graphics card and system RAM. I had been using my RX 6800 with llama.cpp and Qwen3.8-27B for coding work (the earlier post). I installed Strata on the card, repeated the repair benchmark, and investigated why its prompt processing fell short of Strata’s published figures for newer AMD cards.

Setup on a 32 GB PC

The RX 6800 sits in an OCuLink dock attached to a GMKtec NucBox K8 Plus (Ryzen 7 8845HS). Strata’s own link probe measured 7.1 GB/s from host to GPU. Windows sees 28.8 GB of the 32 GB of RAM because the Radeon 780M reserves the rest.

START-HERE.bat --check found the RX 6800 (gfx1030, which Strata lists as unvalidated) and rejected every size of the full model for lack of RAM. The Coder variant fit. ISTA-DASLab made it for coding by removing 256 of the model’s 512 experts; the downloaded version is labeled IQ1_M.

Its experts use a per-layer mix of IQ2_S, IQ3_XXS, IQ3_S and IQ4_NL for gate and up, and IQ4_NL or Q2_0 for down. In the low-RAM mode selected by setup, the GPU holds about 8.4 GiB of experts. RAM holds most of the rest, with the remainder read from the SSD.

Setup downloaded about 65 GB (the model’s two shards and Qwen’s multi-token-prediction layer) and a ready-made AMD engine with ROCm included. It finished in 28 minutes, and the server loaded in under a minute.

The repair benchmark

The benchmark seeds bugs into a tower-defense game that passes all 38 browser checks. The model returns search-and-replace patches as JSON, and a task counts as solved when the patched game passes all 38 checks again, within three rounds. One of the ten seeded bugs never fails a check, so nine count. Each prompt is about 12,000 tokens.

ModelSolvedFirst tryRepair timeDecode
Qwen3.8-27B, llama.cpp (two runs)9/9, 9/97, 91,513 s, 1,108 s22.5 tok/s
Strata Coder, json_schema9/96925 s29.6 tok/s
Strata Coder, schema in the prompt9/99664 s28.2 tok/s

Strata processed a 12,000-token prompt at about 310 tokens per second, twice the 27B’s rate on a similar prompt. Most of the decode speedup came from Qwen’s multi-token prediction (MTP), which drafts tokens for the model to check. About 88% of the drafts were accepted.

I requested JSON through json_schema in the first Strata run and put the schema in the prompt for the second. llama.cpp turns response_format: json_schema into a grammar that constrains every token. Strata describes the schema to the model and validates the answer afterwards.

Twice in the first run, the model answered with something other than a JSON object, costing one round each time. I replayed the two affected prompts four times each with every request style below, for eight requests per style.

Request styleValid JSON
json_schema, temperature 0.2 (the benchmark’s setting)6/8
Same, temperature 08/8
json_object with the schema in the system prompt8/8
No response_format, schema in the system prompt, JSON extracted by the client8/8
Anthropic endpoint with the reply prefilled with {8/8

Prefilling the reply with { failed when I ran the full repair benchmark. On a prompt the server had not cached, Strata echoed the { and the model ended its turn immediately. I put the schema in the system prompt for the final benchmark run because that request style also works with llama.cpp.

Where the prompt time went

Strata’s documentation reports 1,420 tokens per second of prompt processing for the Coder on an RX 9070 XT. I first suspected the SSD and the x4 link. Larger prompt chunks (--prefill 16384) and a locked RAM budget (--resident-budget-gib 12) changed the amount read from disk. Across these runs, prompt processing ranged from 273 to 320 tokens per second.

ConfigurationPromptDecodeRead from SSD per request
Setup’s default308 tok/s28.1 tok/s12 GB
--prefill 16384320 tok/s24.5 tok/s9 GB
--resident-budget-gib 12273 tok/s20.5 tok/s28 GB
Both307 tok/s21.4 tok/s29 GB

One request read nothing from the SSD and still took 36 seconds. STRATA_PREFILL_TIMING=1 splits the GPU time of a prompt into phases. The largest was the linear-attention layers at 9.4 of 38.9 seconds. Waiting for experts streamed from RAM took 4.8 seconds, and the hyper-connection mixers took 4.5.

Logging every rocBLAS call during one prompt (ROCBLAS_LAYER=2) showed 970 BF16 and 600 FP16 matrix multiplications (GEMMs), all with FP32 output. The rocBLAS in Strata’s ROCm build (10.2.0a20260930) has tuned gfx1030 kernels for FP16 with FP16 output, for int8, and for FP32. Every other combination uses a generic fallback kernel.

I first converted the BF16 inputs to FP16 before calling GEMM, following Strata’s approach for older NVIDIA cards. A PyTorch benchmark had shown FP16 running four times faster than BF16 at these shapes. Converting the inputs slowed the engine down because its calls still used a generic kernel with FP32 output. PyTorch had requested FP16 output, allowing rocBLAS to use a tuned kernel. A small HIP program making the engine’s exact call measured 10.85 ms with BF16 inputs and 10.69 ms with FP16 inputs for the largest BF16 shape.

rocBLAS has tuned FP32 kernels, and converting FP16 or BF16 inputs to FP32 is exact. I tried SGEMM, its routine for FP32 matrix multiplication. For one 12,000-token prompt, the FP16 GEMMs took 14.1 seconds with the engine’s existing calls and 4.65 seconds with SGEMM, including the conversion. The largest single shape (N 10240, T 8192, K 2560) went from 86.7 to 27.7 ms.

Running the GEMMs as SGEMMs

The change adds about 90 lines to Strata’s GEMM wrapper. On gfx103x cards, each 16-bit GEMM with at least 64 output columns converts its inputs to FP32 and calls SGEMM. Narrower GEMMs keep the native call because the conversion costs more than they gain. If the FP32 buffers cannot be allocated, the native call runs. STRATA_RDNA2_SGEMM=0 turns the route off, which let me compare both versions with one binary on four fresh 12,000-token prompts each.

Native GEMMsSGEMM route
Prompt processing312 tok/s452 tok/s
GPU time per prompt37.2 s25.5 s
Linear-attention layers9.4 s2.8 s
Attention projections3.6 s1.4 s
Hyper-connection mixers3.85 s2.84 s
Decode29–34 tok/s29–35 tok/s

On random inputs, the two routes differed by at most 4.3 millionths of the largest output. The SGEMM route was at least as close to a float64 reference on all ten shapes I checked. The repair benchmark still solved 9 of 9 on the first try, and model inference time fell from 458 to 374 seconds. The whole run went from 888 to 797 seconds; each task also spends about 20 seconds validating the game in a browser.

Strata’s HIP test suite gave the same results on the patched branch and on an unmodified build of main. Of the 67 tests, 59 passed, 2 were skipped and 6 failed. Four of the failing tests require model files I do not have or cover a CUDA-only feature.

The doorbell test expects mapped memory to become visible without a driver call; the engine already works around this issue. The remaining failure comes from the MMQ test harness, where every case fails on this card. I have not found the cause.

In the engine, switching MMQ on and off produced the same answer to a short prompt and equivalent answers to a 7,000-token prompt. Processing the longer prompt was about 9% faster with MMQ on. I submitted the SGEMM change upstream as Niko1221/Strata#1006.

An earlier fix for the same GEMMs

After I opened #1006, xjc10 pointed me to Niko1221/Strata#835, which they had opened earlier for the same slow GEMMs on RDNA2. It converts the BF16 weights to FP16 and has the activation kernels write FP16 directly. rocBLAS multiplies those inputs with FP16 output using tuned gfx1030 kernels, then the result is widened to FP32. FP16 holds values only up to 65504, so #835 also adds STRATA_F16_RANGE=1, which prints the largest values the prompt path produced.

I built #835 for gfx1030 and ran the three engines on fresh 12,000-token prompts in one session. The SGEMM route measured 439 tokens per second this time, against 452 in the earlier runs.

Native GEMMsSGEMM (#1006)#835
Prompt processing310 tok/s439 tok/s494 tok/s
GPU time per prompt38.1 s26.2 s23.2 s
Linear-attention layers9.4 s2.8 s1.9 s
Attention projections3.6 s1.4 s0.9 s

With STRATA_F16_RANGE=1, all four prompts reported the same maxima: 105.3 for the activations, 11.56 for the BF16 weights and 424.8 for the FP16 GEMM outputs. No value exceeded the FP16 range or became NaN. The repair benchmark solved 9 of 9 on the first try with #835. Model inference time was 307 seconds, compared with 374 for SGEMM, and the whole run took 740 seconds.

The range check covers this Coder model on these prompts. I did not compare #835 with a float64 reference the way I checked the SGEMM route.

xjc10 proposed making #835 the default on gfx103x and rebasing #1006 on top of it as an optional FP32 path (STRATA_RDNA2_SGEMM=1) for models whose values come closer to the FP16 limit. That path converts FP16 or BF16 inputs to FP32 without loss, then performs the multiplication with FP32 rounding. I agreed in the pull request. Strata 0.1.40 then added #835 as an opt-in setting, STRATA_HIP_PROMPT_F16=1, which is off by default; on gfx103x cards the engine prints a tip about it. The same day, the maintainers rewrote the repository’s history, and GitHub closed #1006 automatically when main was force-pushed. The maintainers said the closure was not a rejection.

My PC now runs the official 0.1.40 engine with STRATA_HIP_PROMPT_F16=1. It processed fresh 12,000-token prompts at about 560 tokens per second.

Decode

Decode stays near 40 tokens per second on a 700-token answer. STRATA_DECODE_TIMING=1 showed that the expert was already in VRAM for 61% of lookups. For the remaining lookups, the CPU computed the expert outputs using data in RAM, taking 33.6 of the 93.6 milliseconds in each decode step. Changing the CPU worker count or the share of misses sent over PCIe kept decode between 39.5 and 40.6 tokens per second. Raising the worker count to 15, more than the CPU’s eight cores, dropped it to 10.

With a card that holds all 23.4 GiB of the Coder’s experts, those misses would disappear. Strata’s documentation lists 45 to 60 tokens per second of decode for a 32 GB Radeon AI PRO R9700. I have not tested a larger card.

Hermes

Hermes now points at Strata’s OpenAI-compatible endpoint instead of the 27B. Tool calls came back in the standard format. Hermes sends about 14,000 tokens of system prompt and tool definitions with each request. Strata processed them once in 32 seconds and reused them from its cache for later turns, which took 5 to 7 seconds each. A four-turn file task finished in 93 seconds. The RX 6800 cannot run Strata and the 27B’s llama-server together, so switching back means stopping one of them.