khocr-gen is an open-source Python package that generates synthetic OCR training data for mixed Khmer and English text. It creates realistic document-style images with varied fonts, backgrounds, distortions, and layouts to train deep learning models for text recognition. The tool is designed for researchers and engineers working on low-resource language OCR, particularly for Khmer script.
In the AI & ML space, khocr-gen takes a focused approach. It focuses on generating realistic synthetic training images for Khmer/English OCR models without needing large real-world labeled datasets. It is built as an open-source project for machine learning engineers and researchers. khocr-gen is open source under the MIT license. It ships for the command line.
Behind khocr-gen is LazyGreed, and it first shipped in 2026. The project is developed in the open on GitHub with 2 commits in the last 90 days. Key capabilities include Synthetic Image Generation, Khmer-English Support, and Deep Learning Ready.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Same category — not a similarity match