Qari-OCR-v0.3-VL-2B-Instruct is a 2-billion-parameter vision-language model specialized for optical character recognition and visual document understanding tasks. It accepts both image and text inputs, enabling it to read text from scanned documents, screenshots, and complex layouts while following natural language instructions. The model is distributed on Hugging Face and can be used locally via Transformers or Docker, making it suitable for developers building offline or privacy-focused document processing applications.
Qari OCR V0.3 VL 2B Instruct sits in PulseGate's Foundation models & chat category. It focuses on extracting accurate text and structured information from complex document images without proprietary OCR services. Qari OCR V0.3 VL 2B Instruct is an open-source project aimed at developers. The project is open source (Open Source). It runs on the web and the command line, and it can be self-hosted.
Behind Qari OCR V0.3 VL 2B Instruct is NAMAA-Space, and it first shipped in 2025. Among its 5 catalogued features are Vision Language Model, OCR, and Document Understanding. It exposes integrations via a public API.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do