about Transformer
updated
What Matters in Transformers? Not All Attention is Needed
Paper
• 2406.15786
• Published • 33
Breaking the Memory Barrier: Near Infinite Batch Size Scaling for
Contrastive Loss
Paper
• 2410.17243
• Published • 93
Forgetting Transformer: Softmax Attention with a Forget Gate
Paper
• 2503.02130
• Published • 32
Transformers without Normalization
Paper
• 2503.10622
• Published • 171
SageAttention2++: A More Efficient Implementation of SageAttention2
Paper
• 2505.21136
• Published • 45
How to train your ViT? Data, Augmentation, and Regularization in Vision
Transformers
Paper
• 2106.10270
• Published • 3
Scaling Vision Transformers
Paper
• 2106.04560
• Published
VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction
Paper
• 2608.26005
• Published • 178
Normalized Low-Rank Adaptation
Paper
• 2608.31036
• Published • 51
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability
Paper
• 2608.30320
• Published • 54
Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM
Paper
• 2609.04098
• Published • 78