Skip to content

Repository files navigation

DocForge

License: MIT Python 3.11+ FastAPI MCP Railway

Document → Clean Markdown / Chunks for RAG pipelines, vector databases, and AI agents.

DocForge is an open-source document ingest platform that converts PDF, HTML, DOCX, and plain text into normalized Markdown plus deterministic chunked JSON with token estimates — ready for Pinecone, Weaviate, pgvector, LangChain, LlamaIndex, or any embedding pipeline.

Live demo: https://docforge.konsole.one
Repository: https://github.com/bdeeps/docforge


Table of contents


Why DocForge?

Every RAG and agent pipeline needs the same first step: turn messy documents into clean, chunkable text. DocForge solves that with:

Problem DocForge solution
PDFs with broken layout PyMuPDF text extraction → structured Markdown
HTML noise (scripts, styles) BeautifulSoup + markdownify cleanup
Word docs (.docx) python-docx → Markdown with headings & tables
Inconsistent chunk boundaries Deterministic chunking (same input → same chunks)
Unknown token counts tiktoken estimates per chunk and document
Agent integration Self-registration, MCP tools, machine-readable manifest
Production deploy Railway Postgres, Docker, OpenAPI

Search terms: RAG document ingest, PDF to markdown API, document chunking, MCP document server, agent self-registration, vector store preprocessing, deterministic chunking, tiktoken chunk size.


Features

  • Multi-format ingest — PDF · HTML · HTM · DOCX · TXT · Markdown
  • Clean Markdown output — normalized whitespace, preserved structure
  • Deterministic chunking
    • headings — split on H1–H6 with heading_path metadata
    • size — fixed token windows with configurable overlap
  • Stable chunk IDs — SHA256-derived, idempotent vector upserts
  • Token estimates — tiktoken (cl100k_base default)
  • Agent self-registration — REST + MCP, API keys (X-DocForge-Key)
  • MCP server — docforge-mcp for Cursor, Claude Desktop, agent tooling
  • Machine-readable discovery — AGENTS.md, llms.txt, /api/v1/agent-manifest
  • AEO FAQ — 24 Q&A pairs at docs/faq.md
  • Web UI — upload, chunk config, Markdown/JSON preview
  • OpenAPI — interactive docs at /api/docs
  • PostgreSQL — Railway-ready persistent agent registry

Who is this for?

  • ML / platform engineers building RAG ingestion pipelines
  • AI agent developers needing MCP-native document tools
  • DevOps teams deploying document preprocessing on Railway/Docker
  • Cursor / Claude users who want one command to ingest docs into a vector store
  • Answer engines & crawlers — structured docs, FAQ, JSON-LD schema

Quick start

One command (local)

git clone https://github.com/bdeeps/docforge.git
cd docforge
chmod +x scripts/dev.sh
./scripts/dev.sh

Open http://127.0.0.1:8787

Manual setup

python3 -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
cd web && npm install && npm run build && cd ..
docforge-api

Ingest your first document

curl -X POST http://127.0.0.1:8787/api/v1/ingest \
  -F "file=@document.pdf" \
  -F "strategy=headings" \
  -F "max_tokens=512"

Response includes markdown, chunks[], token_estimate, and document_id.


Architecture

flowchart LR
  subgraph Input
    PDF[PDF]
    HTML[HTML]
    DOCX[DOCX]
    TXT[Text]
  end

  subgraph DocForge
    API[FastAPI REST]
    MCP[docforge-mcp]
    CONV[Converters]
    CHK[Chunking Engine]
    AGT[Agent Registry]
    DB[(SQLite / Postgres)]
  end

  subgraph Output
    MD[Clean Markdown]
    JSON[Chunked JSON]
    TOK[Token Estimates]
  end

  PDF --> CONV
  HTML --> CONV
  DOCX --> CONV
  TXT --> CONV
  CONV --> CHK
  CHK --> MD
  CHK --> JSON
  CHK --> TOK
  API --> CONV
  MCP --> API
  AGT --> DB
  API --> AGT
Loading

Stack: Python 3.11+ · FastAPI · SQLAlchemy (async) · PyMuPDF · tiktoken · React + Vite + Tailwind


API reference

Base path: /api/v1 · OpenAPI: /api/docs

Method Endpoint Auth Description
GET /health — Health + DB status
GET /agent-manifest — Machine-readable agent discovery JSON
POST /ingest Optional key File → Markdown + chunks (full pipeline)
POST /convert Optional key File → clean Markdown only
POST /chunk Optional key Markdown → chunked JSON
POST /agents/register — Agent self-registration (API key once)
GET /agents — List registered agents
GET /agents/me Required key Authenticated agent profile

Example: full ingest

curl -sS -X POST 'http://127.0.0.1:8787/api/v1/ingest' \
  -H 'X-DocForge-Key: df_YOUR_KEY' \
  -F 'file=@report.pdf' \
  -F 'strategy=headings' \
  -F 'max_tokens=512' \
  -F 'overlap_tokens=64'

Chunk response shape

{
  "id": "a1b2c3d4e5f67890",
  "index": 0,
  "content": "## Section\n\nParagraph text…",
  "token_estimate": 142,
  "char_count": 580,
  "heading_path": ["Introduction", "Overview"],
  "metadata": { "strategy": "headings", "section_index": 0 }
}

Full REST reference: docs/agents/rest.md


Chunking strategies

Strategy Best for Key params
headings Manuals, wikis, structured reports min_heading_level, max_heading_level, max_tokens
size Logs, transcripts, fixed embedding windows max_tokens, overlap_tokens, token_model

Both strategies are deterministic — identical input and config produce identical chunk IDs and boundaries.

Guide: docs/agents/chunking.md


AI agents & MCP

DocForge is built for agent-native document ingest. Agents discover, self-register, and ingest without human UI.

Discovery (start here)

Resource Local Deployed
Agent guide AGENTS.md https://YOUR_HOST/AGENTS.md
JSON manifest /api/v1/agent-manifest https://YOUR_HOST/api/v1/agent-manifest
LLM index llms.txt https://YOUR_HOST/llms.txt
FAQ (AEO) docs/faq.md https://YOUR_HOST/docs/faq.md
Cursor skill .cursor/skills/docforge/SKILL.md https://YOUR_HOST/docs/skill.md

Self-register (60 seconds)

curl -sS -X POST 'https://docforge.konsole.one/api/v1/agents/register' \
  -H 'content-type: application/json' \
  -d '{"name":"my-rag-agent","capabilities":["ingest","chunk","convert"]}'

Save api_key, cursor_config, and mcp_env from the response (shown once).

MCP (Cursor / Claude Desktop)

{
  "mcpServers": {
    "docforge": {
      "command": "docforge-mcp",
      "env": {
        "DOCFORGE_API_URL": "https://docforge.konsole.one",
        "DOCFORGE_API_KEY": "df_YOUR_KEY"
      }
    }
  }
}

MCP tools: get_agent_manifest · register_agent · list_agents · convert_document · chunk_markdown · ingest_document · agent_whoami · health_check

See mcp-config.example.json and docs/agents/mcp.md.


Documentation index

Document Description
README.md This file — overview & quick start
AGENTS.md Primary guide for AI agents
docs/faq.md 24-question FAQ (AEO-optimized)
docs/agents/rest.md REST API reference
docs/agents/mcp.md MCP tool reference
docs/agents/chunking.md Chunking strategies & RAG tips
docs/skill.md Cursor Agent Skill (HTTP copy)
CONTRIBUTING.md How to contribute
SECURITY.md Security policy

Deployment (Railway)

DocForge deploys to Railway with PostgreSQL for persistent agent registrations.

Steps

  1. Fork / connect github.com/bdeeps/docforge
  2. Create Railway project → Deploy from GitHub
  3. Add PostgreSQL plugin → DATABASE_URL auto-set
  4. Set service env:
Variable Value
DOCFORGE_HOST 0.0.0.0
PORT (Railway auto)
DATABASE_URL (Postgres plugin auto)
  1. Deploy using included Dockerfile

DocForge normalizes postgres:// → postgresql+asyncpg:// automatically.

Health check: GET /api/v1/health (returns database: "connected" when Postgres is reachable)


Docker

docker compose up --build

Optional local Postgres:

docker compose --profile postgres up --build

Environment variables

Copy .env.example → .env

Variable Default Description
DOCFORGE_HOST 0.0.0.0 Bind address
DOCFORGE_PORT / PORT 8787 Server port (Railway sets PORT)
DOCFORGE_APP_ROOT . (cwd) Root for web/dist and agent docs
DOCFORGE_DATABASE_URL / DATABASE_URL SQLite file Postgres on Railway
DOCFORGE_MAX_UPLOAD_MB 50 Max upload size
DOCFORGE_API_KEY_HEADER X-DocForge-Key Agent auth header

Project structure

docforge/
├── AGENTS.md              # AI agent onboarding guide
├── llms.txt               # LLM / answer-engine discovery index
├── docs/
│   ├── faq.md             # AEO FAQ (24 Q&A)
│   ├── skill.md           # Cursor skill (HTTP)
│   └── agents/            # MCP, REST, chunking references
├── src/docforge/
│   ├── api/               # FastAPI routes + docs serving
│   ├── mcp/               # docforge-mcp stdio server
│   ├── converters.py      # PDF/HTML/DOCX → Markdown
│   ├── chunking.py        # Deterministic chunk engine
│   └── agent_manifest.py  # Machine-readable discovery JSON
├── web/                   # React UI (Vite + Tailwind)
├── Dockerfile
├── docker-compose.yml
└── .cursor/skills/docforge/SKILL.md

Development

pip install -e ".[dev]"
pytest                    # run tests
ruff check src tests      # lint
cd web && npm run dev     # UI dev server (proxies /api)

FAQ

Common questions: docs/faq.md — formats, chunking, MCP setup, Railway, troubleshooting.

Quick links when deployed:


GitHub topics (for discoverability)

Suggested repository topics:

rag · document-processing · markdown · pdf · chunking · fastapi · mcp · model-context-protocol · vector-database · llm · ai-agents · document-ingestion · tiktoken · langchain · embeddings · railway · python


License

MIT — Copyright (c) 2026 DocForge contributors


DocForge — turn documents into RAG-ready Markdown and chunks.
Star on GitHub · Live demo · Agent guide · FAQ

About

docforge

Resources

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages