page-vision-mcp is an open-source structured visual analysis pipeline designed for coding agents and autonomous AI systems. It enables agents to process, analyze, and interpret UI screenshots and visual data, supporting tasks like OCR and structured UI understanding. Ideal for developers building AI agents that require visual perception capabilities.
page-vision-mcp is an AI project. It focuses on enabling AI coding agents to analyze and interpret visual UI data and screenshots for automation and understanding. It is built as an open-source project for AI developers building coding agents or autonomous agents requiring visual UI analysis. page-vision-mcp is open source under the MIT license. page-vision-mcp is available on the command line and API, and it can be self-hosted.
page-vision-mcp first shipped in 2026. Development happens publicly on GitHub with 2 commits in the last 90 days. Among its 10 catalogued features are Visual UI analysis, OCR extraction, and screenshot processing. It exposes integrations via a public API and an MCP server.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do