2781
Agent-R1: Training Powerful LLM Agents with End-to-End Reinforcement Learning
Python
⭐ 1.6k
⑂ 113
MIT
· 7 天前推送
7 天前
最近推送
2782
2026-06-26
最近推送
2783
21 天前
最近推送
2784
4 天前
最近推送
2785
23 小时前
最近推送
2786
2026-03-18
最近推送
2787
A curated list of awesome resources, tools, and other shiny things for LLM prompt engineering.
Python
⭐ 1.6k
⑂ 199
NOASSERTION
· 2026-02-23推送
2026-02-23
最近推送
2788
Unified multimodal backend for AI data apps
Python
⭐ 1.6k
⑂ 219
Apache-2.0
· 13 小时前推送
13 小时前
最近推送
2790
WikiChat is an improved RAG. It stops the hallucination of large language models by retrieving data from a corpus.
Python
⭐ 1.6k
⑂ 146
Apache-2.0
· 2026-01-31推送
2026-01-31
最近推送
2792
Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback
Python
⭐ 1.6k
⑂ 133
Apache-2.0
· 2025-11-24推送
2025-11-24
最近推送
2794
2025-01-02
最近推送
2795
A Python stream processing engine modeled after Yahoo! Pipes
Python
⭐ 1.6k
⑂ 75
MIT
· 1 天前推送
1 天前
最近推送
2796
Automatically update running docker containers with newest available image
Python
⭐ 1.6k
⑂ 152
MIT
· 2023-02-10推送
2023-02-10
最近推送
2797
12 天前
最近推送
2799
WebGLM: An Efficient Web-enhanced Question Answering System (KDD 2023)
Python
⭐ 1.6k
⑂ 131
Apache-2.0
· 2025-03-25推送
2025-03-25
最近推送
2800
2026-07-25
最近推送
共 9118 条 · 第 140 / 456 页