edu-markdown is a local-first CLI for turning mixed teaching materials into clean Markdown that is ready for search, chunking, and AI workflows.
It solves one narrow but useful problem:
- collect messy source material
- normalize it into readable Markdown
- keep lightweight metadata attached
- make the output portable for downstream tooling
Teaching and content workflows often start with a mess:
- public web pages
- copied notes
- exported HTML
- DOCX handouts
- text-based PDFs
Most converters stop at raw extraction. edu-markdown tries to do the next useful thing: produce Markdown that is readable enough for humans and structured enough for later pipelines.
- Convert a URL into article-style Markdown
- Convert local
html,txt, andmdfiles into normalized Markdown - Convert local
docxfiles into readable Markdown - Convert text-based
pdffiles into readable Markdown - Convert a whole directory in one pass
- Add YAML front matter so the output is ready for downstream tooling
- teachers building reusable lesson material
- content operators cleaning source material for knowledge bases
- AI workflow builders who need structured Markdown instead of raw files
- building a lesson-material knowledge base
- preparing source text for RAG or chunking pipelines
- cleaning exported documents before annotation or review
- normalizing mixed folders into one Markdown-first archive
Input folder:
reading-handout.pdflesson-plan.docxunit-notes.txtarticle.html
One command:
.\.venv\Scripts\edu-markdown.exe convert-dir ".\materials" -o ".\output" --recursiveOutput folder:
reading-handout.mdlesson-plan.mdunit-notes.mdarticle.md
Each generated file keeps readable Markdown plus YAML front matter.
cd C:\Users\admin\.openclaw\workspace\projects\edu-markdown
python -m venv .venv
.\.venv\Scripts\python.exe -m pip install -e .[dev].\.venv\Scripts\edu-markdown.exe convert "https://example.com/article" -o output\article.md.\.venv\Scripts\edu-markdown.exe convert ".\samples\lesson.html" -o output\lesson.md.\.venv\Scripts\edu-markdown.exe convert ".\notes\week1.txt".\.venv\Scripts\edu-markdown.exe convert ".\materials\lesson-plan.docx".\.venv\Scripts\edu-markdown.exe convert ".\materials\reading-handout.pdf".\.venv\Scripts\edu-markdown.exe convert-dir ".\materials" -o ".\output" --recursive---
title: Reading Handout
source: C:\materials\reading-handout.pdf
source_type: pdf_file
fetched_at: 2026-08-15T12:00:00Z
---# Reading Handout
Underline one sentence that reveals the setting.
Circle one word that shows the narrator's tone.If -o is omitted, the tool writes <input-name>.md next to the source file. For URLs, the default output name is derived from the page title or host.
Each output file keeps the original source attached:
titlesourcesource_typefetched_at
That makes the Markdown easier to audit, reprocess, index, or trace back later.
Generated files start with YAML front matter like this:
---
title: Example Article
source: https://example.com/article
source_type: url
fetched_at: 2026-08-15T12:00:00Z
---Then the cleaned Markdown body follows.
- Current release:
v0.1.0 - Status: usable MVP
- Focus: clean local conversion before OCR, chunk export, or richer pipeline features
Currently supported well:
httpandhttpsURLs- local
.html - local
.txt - local
.md - local
.docx - local
.pdfwith an extractable text layer - batch conversion for supported local files
Not yet supported:
- OCR from images
- chunk export
Those are sensible next steps, but not needed for a credible first release.
- Drop mixed teaching materials into one folder.
- Run
convert-dironce. - Feed the generated Markdown into search, chunking, annotation, or AI grading pipelines.
- See examples/README.md for sample source files and generated output.
- Regenerate them with:
.\.venv\Scripts\python.exe .\scripts\generate_examples.py- OCR fallback for image-only PDFs
- chunked JSON export
- front matter fields for subject / grade / unit
- plug-in converters for common education sources
- PDF support currently assumes the PDF has an extractable text layer.
- OCR is not implemented yet.
docxconversion currently favors clean paragraph extraction over rich formatting preservation.- The tool is local-first and CLI-first for now.
Run tests:
.\.venv\Scripts\python.exe -m pytest