From Representation Learning to World Modeling
2026-07-22
Joint Embedding Predictive Architecture (JEPA)
25 words
|
1 minute
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
2026-04-15
Three Part of FlashAttention: Part 1
14 words
|
1 minute
Optimizing the Softmax loss
2026-03-11
Distance-based loss function for deep feature space learning of convolutional neural networks & Pairwise Gaussian Loss for Convolutional Neural Networks
52 words
|
1 minute
Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
2025-12-26
Three Part of Mamba: Part 3
7 words
|
1 minute
Mamba: Linear-Time Sequence Modeling with Selective State Spaces
2025-11-07
Three Part of Mamba: Part 2
7 words
|
1 minute
Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
2025-09-03
Three Part of Mamba: Part 1
11 words
|
1 minute