ocr
157 个项目 · ⭐ 880.8kAndroid native AI inference library, bringing text, image, video, STT, TTS inference
Find your files with natural language and ask questions.
Windows 本地 AI 语音输入:Qwen3-ASR CPU/GPU 双运行形态、本地屏幕 OCR 上下文、热词纠错与可选 AI 润色。
Transform screenshots into searchable Obsidian notes using AI vision and text analysis
🚀 100% local RAG system with one-command setup. Your data never leaves your server. Apache-2.0
给 DeepSeek 补上「眼睛和耳朵」的多模态视觉插件:看图 / OCR / 物体检测 / 视频理解 / 语音转写 / 截图直读,一键安装(DSH 插件)。
Fine-tuned Qwen2-VL-7B for LaTeX OCR using LoRA and Unsloth on the LaTeX OCR dataset. Built augmentation pipeline (rotation, noise, contrast jitter), ran LoRA rank sweep (r=8/16/32), and evaluated across CER, Token F1, BLEU-4, and Exact Match. Deployed as a Gradio Space with live metric computation.
A character tokenizer for Hugging Face Transformers
High-accuracy Russian ANPR system built with YOLOv8 for detection and an optimized, custom-trained PyTorch CRNN for OCR.
Convert images or audio files to plain text on the command line
Advanced receipt OCR and analysis using PaddleOCR, GPT-3.5-turbo, Plotly, and Gradio for interactive visualizations.
共 157 条 · 第 7 / 8 页