← 返回专题广场
linear-attention
2 个项目 · ⭐ 20.9k1
RWKV (pronounced RwaKuv) is an RNN with great LLM performance, which can also be directly trained like a GPT transformer (parallelizable). We are at RWKV-7 "Goose". So it's combining the best of RNN and transformer - great performance, linear time, constant space (no kv-cache), fast training, infinite ctx_len, and free sentence embedding.
Python
⭐ 14.7k
⑂ 1.0k
Apache-2.0
· 1 天前推送
1 天前
最近推送
2
15 天前
最近推送