A collection of interactive, self-contained visualizations. Each opens in its own page.
Visualizations
- BMG Hardware Peak and MFU Reference B60.html
- Sparse Attention for Video Generation — Trends & Key Insights Sparse-Attn-Video.html
- Attention Matrix View — Make Only Frame-0 Query Rows Dense attention_matrix_frame0_dense_detail.html
- SignRound / AutoRound Algorithm Visualization auto-round-v1.html
- NVIDIA B200 Key Specs — Node / GPU / Die b200_key_specs_node_gpu_die.html
- Intel Arc Pro B70 vs B60 Comparison b60_vs_b70_comparison.html
- 1–8 Bit Packing Patterns — Humming Generic Template bit_packing_all_viz.html
- Cache-DiT Working Flow — No Cache vs Cache-DiT cache_dit_workflow_visual.html
- Cosmos 3 Super Architecture Visual Guide cosmos3_architecture_visual.html
- Cosmos 3 — How the Generator Uses Reasoner Context cosmos3_reasoner_generator_context_visual (2).html
- DeepSeek-V4 CSA — How 4 Tokens Become 1 Compressed KV Entry deepseek_v4_csa_4_to_1_visual.html
- Dense FlashAttention vs Block-Sparse Attention: KV Tile Prefetch dense_vs_blocksparse_kv_prefetch.html
- Low-Bit Quantization: Error, Scale, and Rounding e.html
- B200-specific fast_topk_clusters_exact Optimizations fast_topk_clusters_exact_b200_highlight.html
- GLM-5.2 IndexShare — Accurate Architecture Note glm-5.2-sharedindexer.html
- INT4 vs FP4 — Encoding Visualized int4-vs-fp4-encoding.html
- Intel GPU Software Stack — Hardware → PyTorch (with NVIDIA counterparts) intel_gpu_stack_clean.html
- ARK XPU joint_matrix — The Debug Journey joint_matrix_debug.html
- MXFP4 / MXFP8 Quantization Visualizer mxfp4.html
- NVIDIA cp.async vs Intel Xe Prefetch nvidia_cp_async_vs_intel_xe_prefetch_v2.html
- Low-Bit Quantization: Error, Scale, and Rounding quant-error.html
- Quantization Granularity Visualizer · 16×64 quant-granularity.html
- HND vs NHD Sparse KV Memory Access sparse_attention_layout_memory_access (2).html
- SSA Gist Tokens — Unified Visual Deck ssa_gist_tokens_unified_visual.html
- Step1 GEMM: Which Thread Owns Which Elements step1_thread_ownership.html
- TF32 vs FP32: Why Tensor Cores Use TF32 tf32_vs_fp32_explainer.html
- Radix Top-K Visualization topk_methods_comparison.html
- Top-K Methods Comparison: Real vs Approx Top-K topk_methods_comparison_real_exact.html
- Xe SDPA Forward Mainloop — Key Loops and Shapes xe_sdpa_sagev1_concrete_instance (1).html