← Back to explorer

DiffusionBlocks: Block-wise Neural Network Training via Diffusion Interpretation

Type
other
Venue
arXiv / Sakana AI / The University of Tokyo

Summary

Interprets residual updates as Euler steps of a probability-flow ODE, partitions layers into B blocks each assigned an equi-probability noise range, and trains one block at a time with score matching (Bx memory cut). Matches or beats end-to-end on ViT CIFAR-100 (59.30 vs 60.25), DiT CIFAR-10/ImageNet FID, MD4 text8 BPC 1.45 vs 1.56, Llama-2-style LM1B/OWT, and Huginn recurrent-depth (single-pass vs 32 iterations). Also beats NoProp variants on CIFAR-100. Code https://github.com/SakanaAI/DiffusionBlocks.

Keywords

diffusionblocks · block-wise-training · score-matching · sakana · iclr · backprop-free · huginn

Topics

block-wise training, diffusion, memory-efficient training

Research notes

  • Primary: arxiv abs (cs.LG/AI/stat.ML). Sakana AI / UTokyo; correspondence mkshing/takiba@sakana.ai, masanori.koyama@weblab.t.u-tokyo.ac.jp. Code Apache-2.0 https://github.com/SakanaAI/DiffusionBlocks (244 stars / 27 forks at check). HF paper page 5 upvotes; linked models are unofficial (joelhenwang/OdinNext-*), not recorded in hf_* fields. Discord posted abs. Uses public CIFAR/ImageNet/text8/LM1B/OWT; no new corpus, so no datasets_local row. ArXiv license widget not visible in converted abs HTML, so paper license left blank.