MCP OCR Server is an MCP server that provides optical character recognition capabilities by connecting language models directly to the Tesseract OCR engine. Software engineers, automation specialists, and knowledge workers use this server to parse textual data trapped inside visual assets without needing external proprietary cloud services. The tool allows Model Context Protocol clients to handle various image inputs, including local files, web-hosted image URLs, and raw image bytes. Once connected, an assistant can automatically parse receipts, inspect diagrams, digitize document scans, or analyze interface screenshots during interactive conversations. It also exposes diagnostic commands to verify which OCR languages are installed on the host system. The server handles Tesseract dependencies automatically on supported environments such as macOS and Linux, ensuring that AI agents can reliably process mixed media files across automated workflows.
Category: Design, Media & Creative
Tags: image-processing, ocr, tesseract, text extraction
bash pip install mcp-ocr or bash uv pip install mcp-ocr Note: Tesseract installs automatically on macOS and supported Linux distributions. Windows users require manual Tesseract installation. 2. Open your Claude Desktop configuration file at ~/Library/Application Support/Claude/claude_desktop_config.json. 3. Add the server entry under the mcpServers object: json { "mcpServers": { "ocr": { "command": "python", "args": ["-m", "mcp_ocr"] } } } 4. Restart Claude Desktop to enable the OCR tools.Part of MCP Servers
You can install the server using standard Python package managers by running pip install mcp-ocr or uv pip install mcp-ocr. Tesseract will install automatically on macOS via Homebrew and on Linux via native package managers like apt, dnf, or pacman. Windows users must install Tesseract manually.
MCP OCR Server allows AI assistants to extract text from images using the Tesseract OCR engine. It supports inputs such as local file paths, image URLs, and raw image bytes, and includes a tool to list all available OCR languages.
The server exposes two primary tools: perform_ocr, which extracts text from an image path, URL, or raw bytes; and get_supported_languages, which returns the list of installed OCR languages recognized by Tesseract on the host system.
It works with any client that implements the Model Context Protocol, including Claude Desktop. You register it as a stdio server running the python -m mcp_ocr command in the client configuration file.
Yes, MCP OCR Server is open-source software distributed under the MIT License. The code and issue tracker are publicly available on GitHub.