← 返回专题广场
flash-attention
11 个项目 · ⭐ 43.9k1
The official repo of Qwen (通义千问) chat & pretrained large language model proposed by Alibaba Cloud.
Python
⭐ 21.6k
⑂ 1.9k
Apache-2.0
· 2026-03-05推送
2026-03-05
最近推送
2
Official release of InternLM series (InternLM, InternLM2, InternLM2.5, InternLM3).
Python
⭐ 7.3k
⑂ 509
Apache-2.0
· 2025-10-30推送
2025-10-30
最近推送
3
中文LLaMA-2 & Alpaca-2大模型二期项目 + 64K超长上下文模型 (Chinese LLaMA-2 & Alpaca-2 LLMs with 64K long context models)
Python
⭐ 7.1k
⑂ 561
Apache-2.0
· 2026-04-19推送
2026-04-19
最近推送
4
8 天前
最近推送
5
MoBA: Mixture of Block Attention for Long-Context LLMs
Python
⭐ 2.2k
⑂ 158
MIT
· 2025-04-03推送
2025-04-03
最近推送
6
Emotion text classification using Llama3-8b with LoRA and FlashAttention. Based on LLaMA-Factory.
Python
⭐ 73
⑂ 10
Apache-2.0
· 2026-07-12推送
2026-07-12
最近推送
7
1 天前
最近推送
8
Automatically benchmark and optimize attention in diffusion models. 1.5-2x speedup on RTX 4090.
Python
⭐ 35
⑂ 7
MIT
· 2026-02-09推送
2026-02-09
最近推送
9
11 天前
最近推送
10
2026-04-22
最近推送
11
Lightweight, Self-Hosted AI Guardrails Model based on ModernBERT.
Jupyter Notebook
⭐ 12
⑂ 1
Apache-2.0
· 2026-03-25推送
2026-03-25
最近推送