ModernBERT takes inspiration from the Transformer++ (from Mamba), including using RoPE & GeGLU, removing unnecessary bias terms, & adding an extra norm layer after embeddings.
Then, we added Alternating Attention (very impactful!), Sequence Packing, & Hardware-Aware Design
本日のコミューンDS/ML勉強会では以下の内容を紹介しました!
- 2025年の年始に読み直したAIエージェントの設計原則とか実装パターン集
- rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking