Photokin is an MIT-licensed Python package that uses vision LLMs to analyze photos and scanned documents. It produces verbatim transcriptions of text, descriptive captions, relevant keywords, and cautious estimates for dates and locations. The tool augments or replaces traditional EXIF/XMP metadata workflows and integrates with applications like Lightroom. Source code is available on GitHub.
In the Developer Tools space, photokin takes a focused approach. Automatically generating accurate metadata, transcriptions, captions, and tags from photo content using vision language models. It is built as an open-source project for developers and photographers. The project is open source (MIT). It ships for the command line.
It is developed by asielen, and it first shipped in 2026. Development happens publicly on GitHub with 14 commits in the last 90 days. Key capabilities include Photo Transcription, Metadata Extraction, and Caption Generation.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Same category — not a similarity match