The official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.
-
Updated
Mar 25, 2026 - Python
The official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.
A hub for various industry-specific schemas to be used with VLMs.
Agentic RAG Harness for long documents, Tree and Graph based reasoning. Cited answers down to the pixel
Yet another self-hosted AI voice assistant. GlaDOS' blazing fast pipeline with Kokoro TTS voice and vision.
Redact PDF/image-based documents, Word, or CSV/XLSX files using a graphical user interface. Demo: https://huggingface.co/spaces/seanpedrickcase/document_redaction or with try with VLMs: https://huggingface.co/spaces/seanpedrickcase/document_redaction_vlm
IFTG (ImageFromTextGenerator) is a Python package that simplifies creating robust datasets for OCR models. Generate images from text, apply over 10 built-in noise effects, and customize fonts and layouts. IFTG supports all languages and offers endless noise combinations, including custom noise creation.
Beautiful, minimal formula OCR desktop app powered by vision-language model APIs. Capture, edit, render, and copy LaTeX anywhere.
Production grade Document AI workflows run on Tensorlake
🏆 智能证书识别与保研加分系统 - End-to-end certificate recognition & structured extraction powered by FireRed-OCR + Qwen3-VL-4B
Document image retrieval via MCP or API for agentic systems using semantic embeddings, YOLO, and VLM classification.
The CyberTech VLM Detector is a computer vision system designed to run entirely on edge devices, without requiring cloud access. The system uses vision-language models (VLM) to detect and locate objects in images based on natural language commands and development, including my creation of HIM™ and MAIC™
DocuLingo is a powerful document parsing tool built with multimodal large language models to enhance RAG (Retrieval Augmented Generation) workflows.
Digitizing various Bantu languages' dictionaries using OCR and Vision LLMs following Cross-Linguistic Data Format (CLDF), as part of the SocioBaGS Project. Built with guidance from Jr. Prof. Annemarie Verkerk and Ph. D. Rita Popova.
A lightweight tool that converts handwritten Rnote papers into well-structured documents (Markdown/LaTeX/Typst) using VLMs
Hackathons 2026 du master Humanités numériques (Enc-PSL) - résultats du pôle HTR
To associate your repository with the vlm-ocr topic, visit your repo's landing page and select "manage topics."