META-CACHED SPARSE ATTENTION

MC-Sparse.

Deconstructing and Closing the Dense–Sparse
Attention Gap in Diffusion Transformers

Jiarui Chen1,2,3 Zeqiang Lai2,4† Jiangshan Wang2,4 Ziheng Ouyang2,5 Ye Huang6 Xiangyu Yue4 Cewu Lu3,7* Chunchao Guo2*
1Fudan University2Tencent HY3Shanghai Innovation Institute
4MMLab, CUHK5Nankai University
6Peking University7Shanghai Jiao Tong University

† Project lead * Corresponding authors

A training-free, token-level sparse attention framework
for faster video and 3D generation.

1.80×Video denoising speedupMiniMax-H3-Base
2.32×3D denoising speedup3D generation
15%Attention densityFor both results above
Training-freeDrop-in accelerationNo model retraining
MiniMax-H3-Base

MiniMax-H3-Base · 15% attention density Drag the divider to compare.

01 / VIDEO GENERATION

Every frame. Every detail.

Explore motion, texture, and temporal consistency.
Selected MiniMax-H3 results generated with MC-Sparse.

Some prompts are borrowed from VDN-H3.

02 / 3D ASSET GENERATION

Less compute.
All the little things.

Fine structures, continuous surfaces, intricate geometry.
Inspect the original meshes from a shared viewpoint.

2.32× denoising speedup

15% attention density

96.33 F1 @ 0.001

3D generation · Table 2

03 / UNDERSTANDING THE GAP

Where does sparse
attention lose fidelity?

Three sources of error.
Controlled oracle comparisons reveal their effects.

THREE SOURCES OF ERROROriginal figure PDF
Three error sources: block constraints miss important individual tokens; approximate selection chooses the wrong blocks; sparse attention discards the nonzero contributions of unselected tokens.
Three sources of the dense–sparse attention gap. Figure 3 from the paper.
01 / STRUCTURE

Structural binding error

KV blocks tie important tokens to irrelevant neighbors. Shared query groups force queries with different attention patterns to use the same selection.

02 / SELECTION

Selection error

Approximate scores, such as mean-pooled query and key scores, can rank the wrong blocks above the interactions that matter.

03 / DISCARDED TAIL

Discarded tail error

Even an exact per-query selector omits tokens with nonzero contributions. Selecting the best tokens still leaves a gap to dense attention.

What the oracle comparisons reveal

At the same attention density, improve the selector and then relax the grouping constraints.

WAN2.1-1.3B-T2V · 480POriginal figure PDF
Oracle curves across attention densities 0.15 to 0.30: per-query oracle captures the most attention mass, gives the highest output cosine similarity, and has the lowest relative L1 error. Shared-query token oracle follows, then block oracle, then mean-pooled BSA. A gap to dense attention remains even for per-query oracle.
Attention mass captured ↑ · Output cosine similarity ↑ · Output relative L₁ error ↓. Figure 2 from the paper.
  1. Mean-pooled BSA → Block oracleExact scoring isolates selection error.
  2. Block oracle → Token oracleIndividual KV tokens remove KV-block binding.
  3. Token oracle → Per-query oracleIndependent selections remove shared-query binding.
  4. Per-query oracle → DenseThe remaining gap reflects discarded contributions.

Oracle selectors use exact dense-attention probabilities. These are diagnostic fidelity comparisons at matched densities, not runtime speedup measurements.

From these findings to MC-Sparse

04 / THE IDEA

Select tokens precisely.
Reuse across many steps.

Compute exact token selections once at each anchor step.
Reuse them over the following denoising steps.

The oracle comparisons motivate three design choices: finer-grained selection, accurate attention scoring, and compensation for discarded contributions.

MC-Sparse addresses all three: select individual KV tokens, group similar queries into GPU-aligned tiles, and reuse exact selections with a cached residual across denoising steps.

01
Group similar queries into equal-size GPU tiles Colored circles represent queries. They are reordered so similar queries share an equal-size group. Colors illustrate similarity, not measured attention values. QueriesEqual-size groups Color = query similarityOne row = one GPU tile

Group similar queries

Fast PDDP puts similar queries into equal-size, GPU-aligned groups. Each group shares a KV selection, reducing disagreement within a group.

Reduce query binding
02
Select individual KV tokens using exact attention mass Bars illustrate the exact attention mass aggregated over a query group. The purple bars identify the highest-scoring individual KV tokens. Their indices are cached for reuse. Exact attention mass per KV token Keep top-K tokensDiscard the rest

Select individual KV tokens

Rank individual KV tokens using exact attention mass at anchor steps. Select the top-K across block boundaries, then cache their indices for reuse.

Finer selection · Exact scores
03

Add the cached residual

Cache the dense–sparse output difference at an anchor step. Add it to later sparse outputs to compensate for discarded contributions. Q, K, and V stay fresh.

Compensate for the discarded tail

Schematic illustrations. O denotes an attention output; R is the cached dense–sparse residual.

THE MC-SPARSE PIPELINEOriginal figure PDF ↗
At an anchor step, MC-Sparse computes query groups, exact KV indices, and a dense-minus-sparse residual. Reuse steps use this metadata with fresh Q, K, and V, then add the residual.
Cache the metadata. Recompute the attention with current Q, K, and V.

05 / QUANTITATIVE RESULTS

The fidelity–efficiency balance.

Higher fidelity to dense-attention outputs.
Faster denoising across video and 3D generation.

MINIMAX-H3-BASE · DENOISING SPEEDUP

Same model.
Less waiting.

Dense attention1.00×
MC-Sparse · 25%1.61×

1.80× speedup Relative to dense attention (FA3)

MiniMax-H3-Base

Results reproduced from Tables 1–2 and Figure 1 of the paper. Speedup measures DiT denoising, including warm-up; condition encoding and VAE decoding are excluded. Video fidelity is measured against dense attention. See the paper for full settings.

First page of the MC-Sparse manuscript, including authors and affiliations

READ THE FULL STORY

MC-Sparse

Deconstructing and Closing the Dense–Sparse
Attention Gap in Diffusion Transformers

View paper ↗Download PDF ↓Code on GitHub