Open-source deep research for AI agents: 40 channels, 10+ Chinese sources.
A1111, ComfyUI, Forge, Forge-Classic/Neo, ReForge, SD-UX - One NoteBook for Google Colab & Kaggle
Scalable group inference for generating high quality and diverse images with diffusion models.
An easy way to view the images and metadata generated by Stable Diffusion's Automatic1111 WebUI
Generating figures from research papers, using textual captions from the paper.
Visual Instruction Tuning for Qwen2 Base Model
An end-to-end image captioning project using a CNN encoder (ResNet-50) and LSTM decoder in PyTorch. Includes vocabulary building, preprocessing, training with BLEU evaluation, and inference. Generates natural language captions for images with saved metrics, model checkpoints, and visualization outputs.
ONNX-compatible Fast SeamlessM4T—Massively Multilingual & Multimodal Machine Translation
A Continual Learning Framework for Production LLM Agents
共 7264 条 · 第 302 / 364 页