MCP Knowledge Base is an MCP server that processes local documents and answers queries based on their contents through similarity search. It connects directly to files on local storage across PDF, DOCX, TXT, and HTML formats, making their textual data searchable for AI agents. The server utilizes text processing libraries including pdf-parse, mammoth, cheerio, and natural to extract text and generate document indices. Developers, technical writers, and researchers use this tool to interface desktop documents with MCP-compatible clients without relying on external cloud storage. MCP Knowledge Base enables automated document ingestion, directory-level indexing, similarity-based snippet retrieval, and file management through eight distinct tool primitives. It includes dedicated support for Chinese text processing and persists both original documents and indexing data in local directories. By exposing retrieval tools directly via standard Model Context Protocol schemas, connected AI assistants can cite local research, compare multi-page reports, and manage internal knowledge bases in real time.
Category: AI Memory & Context
Tags: document processing, docx, pdf, rag, similarity-search
bash npm install 2. Build the TypeScript files to JavaScript: bash npm run build 3. Add the server configuration to your MCP client configuration file (such as Claude Desktop or Cursor): json { "mcpServers": { "knowledge-base": { "command": "node", "args": ["dist/index.js"], "env": {} } } } Ensure the path in args resolves properly to the compiled dist/index.js file relative to your working directory or provide an absolute path.add_directory tool. - Querying local technical specifications with natural language to locate relevant passages and similarity scores using query_knowledge_base. - Listing, inspecting, and deleting stored reference documents through MCP client prompts using list_documents, get_document, and remove_document. - Inspecting knowledge base metrics, document counts, and stored file details by executing the get_stats tool. - Parsing local HTML documentation and plain text files for context injection without uploading data to external third-party cloud services.Part of MCP Servers
MCP Knowledge Base is a Model Context Protocol server that reads local files and answers user questions based on document content using similarity search. It processes formats such as PDF, DOCX, TXT, and HTML locally and provides retrieval tools directly to AI assistants.
The server supports local files in PDF, DOCX, TXT, and HTML formats. It utilizes specialized parsing libraries such as pdf-parse, mammoth, and cheerio to extract readable text before indexing and querying the content.
It provides tools to add individual files (add_document), index entire folders (add_directory), search indexed text (query_knowledge_base), list files (list_documents), view specific file details (get_document), remove documents (remove_document), clear the store (clear_knowledge_base), and view repository metrics (get_stats).
Yes, MCP Knowledge Base is open source software distributed under the MIT license. You can review the complete source code, contribute modifications, or inspect its development scripts directly on its official GitHub repository.