Generate

Engineering

Open-VC Attention: an open-source, error-tunable variant of VC-Attention

By the MachGen team

Attention is the most expensive part of generating video with a diffusion transformer. We are now open-sourcing an attention variant built on top of VC-Attention for NVIDIA Blackwell: about 2× faster than BF16 FlashAttention-4 on B200.

Inference optimization comes down to two things: kernels and recipes. Great kernels push the hardware closer to its limits. Great recipes figure out how to turn that speed into real systems. When strong kernels are open, what might have been one company’s moat can become everyone else’s foundation.

Much of what we’ve built at MachGen stands on years of open research, papers, kernels, and ideas shared by people we may never meet. Open-sourcing this is our small way of continuing that cycle—not just publishing what we learned, but giving the next person a higher place to start.

Code at: github.com/MachGen/open-vc-attention

Background

On B200, FP8 doubles tensor-core throughput, but the special-function unit (MUFU) that computes exponentials did not get faster (NVIDIA). With one exponential per attention score, a kernel that sends them all through MUFU tops out near 2.53 PFLOP/s, about half the FP8 peak. VC-Attention (Li et al., 2026) removes that bottleneck with ExpCast: an FP8 (E4M3) byte is almost an affine function of the logarithm of the value it encodes, so the FP8 code of a softmax probability can be computed straight from the score with a single fused multiply-add, with no exponential and no FP32-to-FP8 conversion. The paper pairs ExpCast with V-Smooth, a value-clustering scheme that reduces V’s quantization error. Open-VC keeps ExpCast and builds everything around it differently.

Design and optimizations

ExpCast is the paper’s contribution; making it fast is ours. Once the exponential is nearly free, the cost moves to everything around it: scale handling, rescales of the running output, data movement and input preparation. This section covers what Open-VC does about each, and how it handles V’s quantization error without V-Smooth.

Scheduling around ExpCast

One instruction per score. Q and K are quantized to FP8 with one scale per 128-token block, and V with one scale per head. The kernel never materializes dequantized Q or K. Instead, the softmax scale, the Q and K dequantization scales, log2e, ExpCast’s factor of 8 and the running maximum all fold into the one fused multiply-add that produces each ExpCast code, and byte permutes pack four codes into a single register. Because K’s scale changes from tile to tile, raw score maxima from different tiles are not comparable, so the kernel tracks the running maximum on the true score:

Sij = τ · sQ(i) · sK(j) · 〈Q8,i, K8,j〉

A consistent denominator. The softmax row sum adds up the same decoded FP8 probabilities that feed the PV product. The output is therefore a weighted average of value rows under exactly the weights the kernel used, so ExpCast’s approximation shows up only as a mild, bounded reweighting of keys, never as a mismatch between numerator and denominator.

Mid-window key traversal. ExpCast has no headroom above the running maximum. BF16 FlashAttention-4 can skip a rescale when the maximum grows only slightly, but under ExpCast a score just 0.074 nats above a stale maximum would be clipped. Open-VC therefore applies every increase, and each one costs a rescale of the running output on the CUDA cores, which are already the bottleneck. The cheapest fix is to find the maximum early. Video tokens are flattened in space-time order, and keys near the query’s own position usually score highest, so the kernel starts a few blocks past the query’s diagonal block, sweeps down to block 0, then visits the tail. Every block is still visited exactly once; only the order changes.

j0 = min(diag + w, n − 1), w = 4

With diagonal block 3 and 10 key blocks, the order is 7 6 5 4 3 2 1 0 8 9 instead of 9 8 7 … 0.

A warp-specialized pipeline. Separate warps load K and V tiles, issue the tensor-core matrix multiplies, and run ExpCast and the softmax bookkeeping. QKT, score conversion, the denominator, PV and the output correction form one pipeline, coordinated by four hand-offs: scores ready, probabilities ready, PV issued and output rescaled. Each CTA holds two 128-row query tiles that alternate, so one tile’s softmax runs while the other’s matrix multiplies execute. On B200, each K tile’s scale is loaded early enough to overlap the wait for that tile’s scores.

Packed V. The PV product reduces over keys, so Open-VC stores V transposed and padded, as [H, 128, Spadded], with the key dimension contiguous. That matches the operand layout of the FP8 PV path, so the tensor cores read V with contiguous loads.

Fused input preparation. Activations change at every denoising step, so the BF16 inputs have to be quantized on every call. On B200 this takes three stream-ordered kernels: one quantizes Q and K per block, one reduces each head’s V maximum, and one converts V straight into the packed layout together with its scale. The result is byte-identical to the general preparation path and can be captured in a CUDA graph. At 73,397 tokens × 56 heads it adds 0.86 ms (1.5%) to the attention call, and since preparation grows linearly with sequence length while attention grows quadratically, that share shrinks for longer videos.

Why we left out V-Smooth

V-Smooth clusters value tokens with online k-means, permutes K and V so each 128-row block holds similar tokens, and subtracts the block mean before quantizing V. It is a sound idea for low-bit formats, but at FP8 it does not pay for itself.

The reason is how E4M3 error behaves. Error grows with magnitude, so subtracting a mean lowers V’s quantization error only by the share of energy the mean carried. The paper’s own energy shares predict its measured error ratios almost exactly (0.957 vs 0.961, and 0.696 vs 0.707).

The cost, meanwhile, lands right on the bottleneck. Restoring the means adds per-tile work on the CUDA cores, and a step that regroups with k-means takes 4.7 ms longer at 56 heads, or 7% of VC’s attention time. Permuting keys by value cluster also scrambles the positional order that the mid-window traversal depends on.

At 4 bits, things look different. With NVFP4, the same relative gain applies to a much larger base error, which is where VC-Attention reports V-Smooth’s speedups.

Why V residual repair

Dropping V-Smooth still leaves V’s quantization error, which is most of the total: VC-Attention attributes 82% of attention-output error to V on Wan2.2. With one scale per head, every value token j leaves a residual:

rj = Vj − sV V8,j

That residual is not spread evenly. A small number of tokens round badly and carry most of the error. So rather than reshaping the whole sequence to shrink every residual a little, as V-Smooth does, we fix the worst few exactly.

For each head, we rank tokens by residual energy Σc rj,c2 and keep the top ρ·S. Each selected residual becomes an extra FP8 value row, paired with a copy of its original key. The kernel treats these rows like any other key, so it adds the missing P:,jrj term without any change to its inner loop. Repair rows are visited last and the denominator is restored, so they only affect the numerator. Nothing is clustered or permuted, so the mid-window order stays intact, and the budget ρ sets how much extra time you spend for lower error.

Repair budgetRel. L2 errorTotal time (ms)Error reductionExtra time
0%2.980%58.14--
0.5%2.924%59.051.90%1.55%
1%2.919%59.572.07%2.46%
2%2.912%60.812.28%4.58%
4%2.905%62.762.53%7.95%
8%2.894%66.992.90%15.21%
V residual repair on the 56-head activations (B200). Total time is the complete call, including quantization and repair selection. The BF16 kernel takes 114.36 ms on the same inputs.

The first 0.5% of tokens deliver most of the gain at a small extra cost. Beyond that, cost grows linearly while the gain flattens, so 0.5% is the operating point we use.

Open-VC, no repairOpen-VC, 0.5% repair
An astronaut on Mars (10 s), same prompt and seed. Watch the visor close-up at about 6 s: without repair the face behind the visor collapses into a dark hollow; with a 0.5% V repair budget it renders correctly.

Results

On B200, with activations taken from a MiniMax-H3 video model, Open-VC sustains 2.70-2.76 PFLOP/s in the attention kernel. That is about 60% of the FP8 dense peak, and above the 2.53 PFLOP/s ceiling we derived earlier for any kernel that sends its exponentials through MUFU. It is about 2× faster than BF16 FlashAttention-4, still about 2× once FP8 input preparation is included, and about 1.2× faster than our implementation of the VC-Attention method. Output error against BF16 is about 3% relative L2, and slightly lower with V repair turned on.

We timed three implementations of the same operator in one harness on one B200: BF16, the unmodified upstream FlashAttention-4 4.0.0b33 release; VC, our efficient implementation of the published VC-Attention method (ExpCast plus V-Smooth value grouping) on the same kernel family; and Open-VC. The inputs are BF16 Q, K and V from a real MiniMax-H3 denoising step, shape [73,397, 56, 128]. Seven heads is the per-GPU shape when this model runs with 8-way Ulysses sequence parallelism. Every call is CUDA-graph captured, and we report medians of 12 rounds in randomized order.

Attention throughput on B200, PFLOP/s (higher is better)

MiniMax-H3 activations, 73,397 tokens × 56 heads × 128

BF16 FlashAttention-4
1.35
VC-Attention method
2.20
Open-VC Attention
2.70
  • FA4 published BF16 (1.61)
  • MUFU-only ceiling (2.53)
HeadsScopeBF16 (ms)VC (ms)Open-VC (ms)Open-VC speedup
7attention kernel14.188.566.982.03×
7complete call14.158.757.111.99×
14attention kernel28.6217.2014.162.02×
14complete call28.6617.5414.332.00×
56attention kernel114.3670.1657.262.00×
56complete call114.4671.6758.121.97×
B200, S = 73,397, D = 128. “Attention kernel” is one kernel launch on prepared inputs; “complete call” starts from BF16 inputs and includes all FP8 preparation. Speedup is relative to BF16 in the same run.

The speedup holds from 7 to 56 heads, so the sharded shape used in multi-GPU inference benefits as much as the full model. Both FP8 kernels use the same ExpCast probabilities, so the 1.21-1.23× between VC and Open-VC comes entirely from Open-VC’s scheduling and data layout.

HeadsVC rel. L2Open-VC rel. L2Open-VC + 0.5% repair
73.09%3.19%3.14%
142.86%2.99%2.92%
562.87%2.98%2.92%
Relative L2 error of the attention output against BF16 FlashAttention-4, on Q/K/V dumped from MiniMax-H3 (denoising step 24, block 20, 73,397 tokens). Dumps from steps 10 and 40 show the same ordering.

Accuracy is close between the two FP8 kernels: VC’s value smoothing is 3-4% better in relative terms, and Open-VC’s 0.5% V repair closes about half of that gap at a fraction of the cost.

About the baseline. On these inputs, upstream BF16 FlashAttention-4 runs at 1.35-1.36 PFLOP/s, 15-16% below the 1,605 TFLOP/s it reports for B200. Measured against that published figure instead, Open-VC still reaches 1.68-1.72×.

What it looks like in generated video

Kernel metrics only tell part of the story, so we also generated full videos with MiniMax-H3, swapping only the attention kernel. Each panel uses the same prompt and seed.

BF16 FlashAttention-4VC (ExpCast + V-Smooth)Open-VC (0.5% repair)
A steam train crossing a snowy mountain trestle at sunset (5 s). The locomotive detail, steam and snowy valley hold up across all three kernels.
BF16 FlashAttention-4VC (ExpCast + V-Smooth)Open-VC (0.5% repair)
A night endurance race with in-cockpit close-ups (15 s). Faces, helmet and livery text, rain reflections and headlights hold up in both FP8 kernels.

As with any change in numerics, the FP8 runs are not pixel-identical to BF16. Small differences early in denoising can change a background element or a camera path, while subject, layout and style stay with the prompt.

Scope

Open-VC Attention targets forward inference on long, non-causal video sequences: head dimension 128, multi-head attention, and at least 32,768 tokens for the packed fast path. Other calls fall back to a general path with the same semantics. Fused preparation and V repair currently run on B200, and B300 has its own schedule.

Open-VC is derived from the FlashAttention-4 CuTe DSL kernel and implements ExpCast-FP8 from VC-Attention (Li et al., 2026). Open-VC is an implementation inspired by the published paper.

Faster video models on MachGen

Kernels like this one are part of the inference stack behind the video models we serve. Talk to us if you want this stack for a private deployment in your VPC or dedicated cloud.

Curious how your models perform on MachGen? We’ll run a low-touch PoC in 2-4 weeks, handling the optimization end-to-end while your team stays focused. You’ll get a live endpoint to evaluate the results on your own workloads.