ComfyUI on the RX 6800: fixing INT8, GPU VAE decode, and VRAM spill

rx6800780mcomfyuirocm

A three-fan desktop graphics card edged in teal light, resting on a small computer with an amber glow behind it

I tried the workflows from my 780M music post and Qwen-Image comparison on a Radeon RX 6800 with 16 GB of dedicated VRAM. Some failed outright; others ran slowly or spilled into system memory. MiniMax Music 3 failed at its first INT8 matrix multiplication, and the floating-point fallback took about 7.2 seconds per token. At that rate, generating a 20-second song would have taken roughly an hour.

After the fixes, the 20-second Music workflow finished in 178 seconds. A full two-minute song took 19.3 minutes, compared with 48.1 minutes on the 780M. The final runs also covered H3 video, LTX, image generation, and a Hunyuan3D mesh.

Test setup

Both GPUs were connected to the same Windows PC, with about 28.8 GiB of usable system RAM. The RX 6800 is gfx1030; the 780M is gfx1103. I installed ComfyUI separately for the RX and pointed it at the existing model files with ComfyUI’s extra-model-paths configuration.

RX 6800Radeon 780M
ComfyUI / Python0.37.2 / 3.12.100.37.2 / 3.12.10
PyTorch2.12.0+rocm10.2.0a202609292.12.0+rocm7.15.0a20260728
Launch flags--cuda-device 1 --use-pytorch-cross-attention --reserve-vram 2--use-pytorch-cross-attention --reserve-vram 6
VAEGPUGPU
Local changesRX helper node, comfy_kitchen patches, and shared ComfyUI-GGUF loader changesExisting TheRock setup, plus shared H3 GGUF loader changes

The results compare these two working setups. They use different ROCm builds and reserve different amounts of VRAM; the RX also has the fixes described below.

Only one card generated at a time. I stopped the RX server for the 780M runs. Before the RX H3 runs, I also released the idle 780M server’s models through ComfyUI’s /free endpoint. Even when idle, the 780M server held enough RAM to cause paging during the RX runs. Without releasing those models, RX H3 jobs took 726–1127 seconds.

Matched workflow results

Each pair used the same model files, prompts, dimensions, and sampling settings. Seeds also matched except for the two MiniMax Music rows: the RX validation runs used 527121 for 20 seconds and 527124 for 120 seconds to avoid ComfyUI’s cache; both 780M runs used 527001. Both full-length runs produced 2986 frames, and the 20-second runs produced 501 tokens. The output folder also differed.

Times come from ComfyUI execution events over the websocket, from execution start through the final node; they exclude time waiting behind another job. Each table entry comes from a single run. Repeated RX runs during the work varied by up to about 10%: LTX 2.3 took 119.5 and 107.1 seconds, while H3 Q3 took 233.0 and 221.7 seconds.

WorkflowRX 6800780MSpeedup
MiniMax Music 3, 20 s, 30 diffusion steps178.2 s470.5 s2.6×
MiniMax Music 3, full 120 s, 30 diffusion steps1157.9 s2886.4 s2.5×
H3 Q3_K_XL, 640×384, 56 frames, 6 turbo steps233.0 s636.0 s2.7×
H3 Q2_K, 768×448, 56 frames, 6 turbo steps353.2 s825.1 s2.3×
Hunyuan3D-2mini, 30 steps, octree 384102.3 s171.7 s1.7×
Qwen-Image-2512 Q4_K_M, 1216×640, 20 steps, CFG 2.5345.7 s1093.2 s3.2×
Qwen-Image 2.1 INT8, 1216×640, 25 steps104.7 s208.8 s2.0×
LTX 2.3 Q2_K, 512×288, 97 frames with audio, 8 steps119.5 s288.9 s2.4×
LTX 2.5 Q3_K_S, 512×288, 97 frames with audio, 25 steps290.2 s784.5 s2.7×

H3 Q3 was the first job after a server start on both cards, followed by Q2. Each job reloaded the approximately 13 GB H3 text encoder. Q2 uses a larger canvas here, so the timing difference also includes the extra work at that resolution.

The rx6800-comfyui repository contains the API graphs for both GPUs, along with the RX helper node and two ComfyUI-GGUF patches. The link points to the v2026-10-01 snapshot; the patches may need changes for other ComfyUI or ROCm builds. The Music graphs retain the base seed 527001; measured-runs.json records the RX seed overrides used for the two table rows. LTX depends on saved conditioning tensors from the original runs, and the repository does not include those tensors or any model weights.

For the runs in the table, I patched comfy_kitchen directly in the RX virtual environment. The published helper node applies those changes at runtime. With that node and an unmodified comfy_kitchen, the 20-second Music graph took 183.6 seconds and Qwen-Image 2.1 took 105.6 seconds. The lab report retains the intermediate results as well as the final table.

Calling INT8 kernels through rocBLAS

The failing call was torch._int_mm, which uses hipBLASLt in this build. The error was HIPBLAS_STATUS_INVALID_VALUE, and the logs reported a missing TensileLibrary_lazy_gfx1030.dat. An FP32 fallback ran, but it made MiniMax impractically slow. Converting weights to FP16 improved things enough to complete songs and Qwen-Image 2.1 images.

Plain rocBLAS in the installed runtime also contained INT8×INT8→INT32 kernels for gfx1030, labeled generic “fallback” builds. I could call them through ctypes with rocblas_gemm_ex. The RX patch keeps a handle per device, binds it to PyTorch’s current stream, and supplies a fixed 64 MB workspace. HIP graph capture records GPU operations for replay, reducing the overhead of launching them repeatedly. The workspace prevents rocBLAS from allocating memory during that recording.

That restored the model’s original algorithm: quantize activations per row, multiply INT8 matrices, then apply the scales. The integer matrix product matched the reference exactly. Against a floating-point reference, the full linear layer had about 0.9% relative error from activation quantization, consistent with the model’s intended INT8 calculation. I did not verify the INT8 kernel path used on the 780M. The earlier FP16 path had lower numerical error, so switching back also changed the generated audio.

Large GEMMs reached about 30–31 TOPS with INT8, compared with 25–27 TFLOPS for FP16 in these measurements. These generic fallback kernels help explain why INT8 was only about 1.2× faster at these sizes. Multiplying a two-row input by a 24576×4096 weight matrix took 1.04 ms through native INT8, versus 2.11 ms through FP16. Qwen-Image 2.1’s end-to-end time fell from 137.0 to 105.3 seconds in that intermediate comparison.

VAE kernels and attention needed their own fixes

I initially forced VAE decode onto the CPU after GPU video decode failed with HIP unspecified launch failure. Profiling later showed that this build’s MIOpen paths for gfx1030 handled some convolutions poorly.

For stride-one, ungrouped 3D convolutions, the RX helper runs one 2D convolution per temporal kernel tap and sums the results. An isolated LTX 2.3 decode of 97 frames at 512×288 dropped from 44 seconds on CPU to 1.9 seconds on GPU, matching the CPU output within 3e-6. That is a decoder measurement, separate from the full workflow time in the table.

MiniMax’s audio VAE had another expensive path: dilated 1D convolutions. A representative layer with dilation 9 took 436 ms through the selected MIOpen kernel and 9 ms as matrix multiplications per kernel tap. On a fresh server, replacing these calls cut the 20-second song’s audio decode from 125.4 seconds to 2.5 seconds. Windows GPU timeout remains a plausible explanation for the earlier launch failures; these measurements do not establish it as the cause.

H3 and Hunyuan3D also used PyTorch’s scaled dot-product attention (SDPA), which weights and combines information according to how closely query and key vectors match. Its math backend allocated large attention score matrices. H3 used 14.5 GB of shared system memory and took roughly six minutes per step. Splitting queries into bounded chunks kept the score buffers small. On the initial Q2 preset this brought sampling down to about 34 seconds per step.

The chunk budget mattered even after the Q3 model loaded fully into VRAM. Its 9.3 GB of weights fit, but temporary buffers spilled about 1.2 GB into shared memory. Reducing the score budget from 256 million to 64 million elements cut the first step from about 300 seconds to 28 seconds in that experiment. The smaller budget added about four seconds to Hunyuan3D decode.

Models left in memory after a long song

MiniMax runs an autoregressive model, a diffusion model, and an audio VAE in sequence. On this Windows/ROCm setup, mem_get_info reported enough free memory for ComfyUI to leave previous models resident. After a long song, later decode buffers spilled into shared system memory. A 20-second decode that normally took two seconds could take 16–22 seconds; restarting the server cleared it.

The RX helper now unloads idle models before MiniMax diffusion and before audio decode. The same rule before video VAE decode also helped LTX. The exact owner of the extra residency was not identified, but unloading between stages kept subsequent jobs fast in the measured sequence.

Jobs in one server process, in orderARDiffusionDecodeTotal
20 s song115.0 s59.7 s2.2 s178.2 s
Full 120 s song743.5 s403.5 s9.9 s1157.9 s
20 s song afterwards113.3 s59.7 s2.0 s175.2 s
Another 20 s song113.2 s59.7 s2.1 s175.2 s

I also kept an FP16 copy of the small depth decoder’s quantized weights for each Music generation call. This used about 1.1 GB of extra VRAM. That decoder runs seven times per audio frame. Batching independent diffusion windows saved about 9% in the 20-second case. A separate weighted-sum workaround was needed for the long song because its original einsum triggered an invalid rocBLAS launch at roughly 3000 frames on gfx1030.

For the full-length benchmark I wrote enough lyrics for a two-minute structure. On both GPUs, the autoregressive stage generated 2986 frames of audio codes, covering about 119 seconds. Earlier attempts that simply increased the duration while retaining the short lyrics stopped early, between about 15 and 65 seconds. Checking only the saved file’s duration would have missed the early stop.

Generated samples

Qwen-Image-2512 output: an amber image panel above a compact computer, with a teal frame and dark space on the left.

The Qwen-Image-2512 scene follows the existing OGP workflow: 1216×640, dark slate, amber light, and space on the left for a title. The final images from the two GPUs had a mean pixel difference of 3.0 on a 0–255 scale. Qwen-Image 2.1’s corresponding difference was 1.1.

Qwen-Image 2.1 output: a compact PC and a red potion bottle on a glass pane, lit in amber and teal.

The H3 Q3 sample has 56 frames at 24 fps with audio. The H3 scenes differed between GPUs despite the same seed. Different compute dtypes and kernels can change sampling; I checked the videos visually, without a formal quality score. LTX 2.5 also uses euler_ancestral, which adds noise during sampling. Its audio was near-silent in both the old CPU-VAE run and the new GPU-VAE run.

The two-minute Music 3 sample is available for listening. I listened for problems with the overall sound and checked some sung lines against the written lyrics. I found no problems in those checks, though I did not compare every line. The short 780M music samples were quieter than the RX samples; the full-length pair had the same overall RMS, about 0.152.