← Feed

From the source

[NeurIPS 2022] Transformers meet Stochastic Block Models: Attention with Data-Adaptive Sparsity and Cost — forck