Coding on the RX 6800 while the Radeon 780M generates sprites
Qwen3.8-27B was generating about 7 tokens per second on my Radeon 780M, with a six-minute wait before the first token from a 20K-token prompt (the earlier run). After setting up an RX 6800 for ComfyUI, I tried moving the coding model to its 16 GB of dedicated VRAM. That brought generation to roughly three times the 780M’s rate and left the integrated GPU available for other work.
I tried small coding models on the 780M, then gave it the task of selecting code for the RX model to repair. That reduced the repair prompts, but having the RX model select its own code was faster in these tests. The two GPUs also completed a game build and sprite generation at the same time, with RX generation speed close to the runs where it worked alone.
Running the 27B on the RX 6800
The model is the same Unsloth Qwen3.8-27B-UD-Q3_K_XL file, served by Unsloth’s build of llama.cpp with the Vulkan backend. All model layers ran on the RX 6800. The server had a 64K context window, a Q8 KV cache for its attention state, flash attention enabled, and a micro-batch of 256.
The longer prompts contained source code. As their length increased, prompt processing slowed and output generation stayed between 19.1 and 22.5 tokens per second.
| Prompt | Time to first token | Prompt processing | Decode |
|---|---|---|---|
| 14,012 tokens | 90 s | 155 tok/s | 22.5 tok/s |
| 27,151 tokens | 211 s | 129 tok/s | 21.0 tok/s |
| 48,511 tokens | 490 s | 99 tok/s | 19.1 tok/s |
Qwen’s built-in multi-token prediction, which gave the 780M a 67% decode speedup, held the RX 6800 at about 7 tokens per second in every configuration I tried. Without it, decode at short context was 23.5. Micro-batches of 384 or larger dropped prompt processing from about 190 to 12 tokens per second in llama-bench.
At 128K context, the KV cache filled the card’s 16 GB, and prompt processing fell to about 9 tokens per second. At 64K the server uses 13.8 GB and leaves some headroom.
Any other program holding VRAM on the RX 6800 caused a sharp drop. My idle RX ComfyUI server kept 2.58 GB after unloading its models. With it running, decode was 22.2 tokens per second at 12K context and 7.5 at 16K, as the growing KV cache spilled into shared system memory. Moving 8% of the layers to the 780M avoided the drop but halved decode, because every token then crossed between the two cards. I now stop the RX ComfyUI server while the LLM is in use.
With the RX available to the model alone, the Hermes agent finished a small file task in 190 seconds. The first attempt took 609 seconds with another program holding VRAM and a 128K context window. Those runs changed both memory conditions and context size. Hermes expected at least 64K context, while the 780M setup could provide only 24K.
Building the game from a revised specification
I gave Hermes the tower-defense specification from my earlier harness, the task on which the 27B had stalled on the 780M. On the RX 6800, it reasoned for 30 minutes before calling a tool, then wrote the three game files. The run took 6,022 seconds and passed 33 of the browser validator’s 38 checks.
Two of the five failures came from the validator. It read debug fields, enemyPositions and selectedTowerType, that the specification never asked for. It also assumed a fixed tower button order, an id named splash, phone and desktop canvas sizes, and specific points where it taps the canvas. I added every one of those assumptions to the specification.
For subsequent runs I used a small orchestration script to call the same 27B. It requested a plan and generated each file in a separate call. Repairs returned exact search-and-replace patches; after applying one, the script ran the browser validator and kept the change only if more checks passed.
| Specification | Process | Result | Time |
|---|---|---|---|
| Original | Hermes agent | 33/38 | 6,022 s |
| Original | Orchestrator | 33/38 | 2,004 s |
| Revised | Orchestrator | 38/38 | 907 s to build |
| Revised | Orchestrator, second run | 38/38 after two accepted repairs | 2,228 s |
The original specification still reached only 33/38 after six repair attempts. With the revised specification, one initial build passed all checks in 907 seconds. That run then spent another 1,208 seconds on a MiMo-9B review on the 780M, for a total of 2,137 seconds. The second revised run included a review by the 27B and four repair attempts, two of which were accepted. The table therefore includes different amounts of review and repair work.
Both reviewers exhausted their 12,000-token output budgets. MiMo-9B on the 780M reasoned through the code for about 20 minutes without reaching a verdict. The 27B filled a JSON list with entries marked “no fix needed”. Its one concrete finding incorrectly claimed that a splash tower costing 80 violated a requirement that the price exceed 70.
The validator had its own bug. It killed only the Edge process it launched, while the headless browser ran in a separate process tree. Each validation left several software-rendered browser processes running the game loop. After a few dozen runs they used all the CPU, and RX decode fell to about 3.5 tokens per second. The cleanup now kills every Edge process that uses the run’s temporary profile directory.
Testing coding helpers on the 780M
To test the 780M as a helper, I seeded ten realistic bugs into a game that passes all 38 checks. Examples are an off-by-one loop, an inverted loop condition, and gold spent from the score instead. The validator could not detect one of them, a smaller HUD font, so nine counted. The 27B on the RX 6800 fixed all nine, seven on the first attempt.
Ling-3.0-tiny on the 780M decodes at 39 tokens per second, faster than the 27B. Each round, both models proposed a fix, and the first one to pass all 38 checks would win. Ling won none of the nine. Each of its proposals took 88 to 103 seconds, and its diagnoses were wrong. In one case it explained a no-op damage line as enemy positions that never updated.
I also tried three tasks with other models that fit in the 780M’s available memory while the RX server was running. MiMo-9B fixed none. Qwen3-8B and Ornith-9B each fixed the desktop canvas-layout bug, the only CSS task in that set.
For code selection, the 780M model read an outline of script.js: about 800 tokens of function names and comments. It selected the functions to send to the repair model, reducing the 27B’s prompt from about 11,900 tokens to 7,000.
| Setup | Fixed | Median per task |
|---|---|---|
| 27B with the full file (two runs) | 9/9, 9/9 | 126 s, 123 s |
| Qwen3-8B navigator on the 780M (three runs) | 9/9, 9/9, 7/9 | 97 s, 100 s, 83 s |
| 27B navigating for itself | 9/9 | 78 s |
| Navigator on the 780M, one task ahead | 9/9 | 93 s |
The 780M navigator selected the function containing the bug in 25 of 27 first attempts. The 27B selected it every time, averaging 7.6 seconds per selection. Qwen3-8B on the 780M averaged between 5.8 and 8.7 seconds across four runs. Preparing the next task during the current repair overlapped about 32 seconds of work per task, including about 22 seconds of browser validation on the CPU. The selection call on the 780M took about 6–9 seconds. Timing varied substantially between runs: the 27B took 60 seconds to repair one bug in one run and 301 seconds in another. These measurements cover this small set of repairs.
Generating sprites during the build
The 780M already runs ComfyUI-TheRock and produced the sprites in my Qwen-Image 2.1 tests. I ran the orchestrator build on the RX 6800 while the 780M generated eight 512×512 sprites with Qwen-Image 2.1 and its W4A8 text encoder.
The memory guard stopped the first image-generation attempt when free RAM fell to 1.46 GB. The LLM server had started with --no-mmap and held about 8 GB of the model file in private system memory, although all model layers ran in the RX’s VRAM. The 780M uses system RAM as graphics memory, so those 8 GB were unavailable to it. With memory mapping, the server needs about 0.6 GB of system RAM during requests, and decode speed did not change.
On the second attempt, both jobs finished. The build passed 38 of 38 checks after one repair, in 1,107 seconds. The eight sprites took 537 seconds, 52 to 71 seconds each, about the same as on an idle PC. Decode on the RX 6800 stayed between 22.2 and 23.5 tokens per second, within 0.2 tokens per second of builds with nothing else running. Free RAM never fell below 5.2 GB.
![]()
Seven of the eight sprites came back on an opaque white background, although earlier runs with the larger text encoder produced transparent backgrounds. A short script now makes white areas connected to the image border transparent. One more 27B call added the sprites to the game, with the old vector shapes as a fallback, and the game still passed all 38 checks. In the browser, the towers and enemies now use the generated art.
When I played the game, I found a placement problem the checks had missed. The specification requires one tap point to be buildable ground, and in this build the path runs through that point, so a tower can sit on the path.