page-vision-mcp is an open-source structured visual analysis pipeline designed for coding agents and autonomous AI systems. It enables agents to process, analyze, and interpret UI screenshots and visual data, supporting tasks like OCR and structured UI understanding. Ideal for developers building AI agents that require visual perception capabilities.
In the AI & ML space, page-vision-mcp takes a focused approach. It focuses on enabling AI coding agents to analyze and interpret visual UI data and screenshots for automation and understanding. page-vision-mcp is an open-source project aimed at AI developers building coding agents or autonomous agents requiring visual UI analysis. The project is open source (MIT). The product ships for the command line and API, and it can be self-hosted.
page-vision-mcp first shipped in 2026. The project is developed in the open on GitHub with 2 commits in the last 90 days. Among its 10 catalogued features are Visual UI analysis, OCR extraction, and screenshot processing. It exposes integrations via a public API and an MCP server.
Latest indexed changes and source events
page-vision-mcp verified by the PulseGate indexer
Other apps tracked under the same category.