Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions .github/workflows/snapshot.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -28,6 +28,7 @@ jobs:
sudo apt-get update
sudo apt-get install -y \
shellcheck \
ocrmypdf \
pdftk \
poppler-utils \
sane-utils \
Expand All @@ -37,6 +38,7 @@ jobs:
brew install \
bash \
shellcheck \
ocrmypdf \
pdftk-java \
poppler \
sane-backends \
Expand Down
61 changes: 56 additions & 5 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,7 @@

**pdfmt** is a command-line toolkit for the PDF tasks you keep doing again and again.

Merge PDFs. Split them. Extract pages. Reorder documents. Scan from your scanner. Add stamps. And when a task becomes repetitive, turn it into a reusable workflow.
Merge PDFs. Split them. Extract pages. Reorder documents. Scan from your scanner and OCR them. Add stamps. And when a task becomes repetitive, turn it into a reusable workflow.

No GUI required. No cloud upload required. Just your PDFs, your terminal, and the tools you already use.

Expand All @@ -27,11 +27,12 @@ Working with PDFs often means reaching for a different tool for every little tas
* "I scanned the front and back separately. Now I need to put them back together."
* "I need to merge these five documents."
* "I need to add a processed stamp."
* "I need to OCR my documents."
* "I have to do this every week."

`pdfmt` is designed for exactly these situations.

Instead of remembering a collection of different commands and tools, you can use one interface:
Instead of remembering a collection of different commands, options, pipes and tools, you can use one interface:

```bash
pdfmt <command> <subcommand> [arguments]
Expand Down Expand Up @@ -193,6 +194,20 @@ Here, page 1 becomes part of every resulting PDF.

---

### Add an OCR text layer to a PDF document

If you have scanned documents, you can simply add an OCR text layer and then search for text within those PDFs:

```bash
# add the OCR text layer
pdfmt ocr add file-without-ocr.pdf file-with-ocr.pdf

# show the OCR text layer (to use with grep)
pdfmt ocr show file-with-ocr.pdf
```

---

### Mark an invoice as processed

You received an invoice as a PDF and want to mark it as processed:
Expand Down Expand Up @@ -225,7 +240,7 @@ With a workflow, all of that can become:
pdfmt workflow run scan-duplex adf output.pdf
```

A workflow is simply an executable shell script:
A workflow is simply an executable shell script like this one, a simplified version of `scan-duplex` from the official [pdfmt-workflows](https://github.com/eifelcode/pdfmt-workflows) repository (see also pre-defined workflows below):

```bash
#!/usr/bin/env bash
Expand Down Expand Up @@ -302,9 +317,10 @@ chmod +x ~/.local/share/pdfmt/workflows/$WORKFLOW_NAME
`pdfmt` currently provides commands for:

| Task | Examples |
| ------------------ | ---------------------------------------------------------------- |
|--------------------|------------------------------------------------------------------|
| **Extract pages** | Extract even pages, odd pages, or arbitrary ranges |
| **Merge PDFs** | Merge documents, interleave duplex scans, insert documents |
| **OCR PDFs** | Execute OCR operations on PDF files |
| **Remove pages** | Remove even pages, odd pages, or arbitrary ranges |
| **Scan documents** | Scan using an ADF or flatbed scanner |
| **Sort pages** | Reverse, shuffle, swap, move and handle duplex page orders |
Expand Down Expand Up @@ -343,6 +359,10 @@ Below is a quick overview of all available commands. To get detailed usage instr
| | `merge duplex` | Interleave two PDFs (front/back scans) into one duplex document |
| | `merge help` | Help of category *Merge* |
| | `merge insert` | Insert a PDF into another document at a specific page |
| **OCR** | `ocr add` | Add an OCR text layer to a PDF document |
| | `ocr exists` | Check if a PDF document contains an OCR/text layer |
| | `ocr help` | Help of category *Ocr* |
| | `ocr show` | Display the extracted OCR text layer of a PDF document |
| **Remove** | `remove even` | Remove all even-numbered pages |
| | `remove help` | Help of category *Remove* |
| | `remove odd` | Remove all odd-numbered pages |
Expand Down Expand Up @@ -385,8 +405,9 @@ The required runtime tools are:
| [pdftk](https://www.pdflabs.com/tools/pdftk-the-pdf-toolkit/) | `pdftk` | `pdftk-java` | PDF manipulation |
| [ImageMagick](https://imagemagick.org/) | `imagemagick` | `imagemagick` | Image/PDF operations |
| [SANE](https://sane-project.org/) | `sane-utils` | `sane-backends` | Scanning |
| [OCRmyPDF](https://github.com/ocrmypdf/OCRmyPDF) | `ocrmypdf` | `ocrmypdf` | OCR operations |

Not every command requires every dependency. For example, SANE is only needed for the `scan` command.
Not every command requires every dependency. For example, SANE is only needed for the `scan` command and OCRmyPDF is only needed for `ocr add`.

#### Linux

Expand Down Expand Up @@ -502,6 +523,36 @@ Available settings include:
|------------------|----------------------------------|-----------------------------------------------------------------------------|
| `workflow.paths` | `~/.local/share/pdfmt/workflows` | List of directories separated by : where your workflow scripts are located |

### OCR configuration

The `ocr` command uses:

```text
~/.config/pdfmt/ocr.conf
```

Available settings include:

| Setting | Default | Description |
|----------------|---------|-------------------------------------------------------------------------|
| `ocr.language` | `eng` | Language codes used for OCR text recognition. |
| `ocr.plugin` | ` ` | The OCRmyPDF engine plugin to be used. If empty Tesseract will be used. |

**Note:** If a configured plugin like `ocrmypdf_rapidocr` is not installed on the system, `pdfmt ocr add` will exit with a clear error message rather than creating faulty PDFs.

#### Recommendation for OCR: RapidOCR or EasyOCR

`pdfmt` uses OCRmyPDF for text recognition, which also supports plugins. For optimal performance and accuracy (especially on CPUs and Intel Macs), the RapidOCR plugin is recommended. If you need maximum recognition quality on complex layouts and have a powerful GPU and RAM available, take a look at EasyOCR.

After installing your preferred plugin, you can enable it by updating the configuration file:

```bash
# Enable RapidOCR:
ocr.plugin=ocrmypdf_rapidocr
```

**Note:** If you get errors like `ValueError: Unsupported rec.lang_type='latin' for PP-OCRv6 small model.` ensure the correct model for RapidOCR is installed.

---

## A command-line tool that stays out of your way
Expand Down
3 changes: 2 additions & 1 deletion completions/pdfmt.bash
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@ function _pdfmt_completions()

# Level 1: Main commands
if [[ "$COMP_CWORD" -eq 1 ]]; then
local commands="extract merge remove scan sort split stamp version workflow help"
local commands="extract merge ocr remove scan sort split stamp version workflow help"
mapfile -t COMPREPLY < <(compgen -W "$commands" -- "$current")
return 0
fi
Expand All @@ -19,6 +19,7 @@ function _pdfmt_completions()
case "${COMP_WORDS[1]}" in
extract) mapfile -t COMPREPLY < <(compgen -W "even odd range help" -- "$current") ;;
merge) mapfile -t COMPREPLY < <(compgen -W "all duplex insert help" -- "$current") ;;
ocr) mapfile -t COMPREPLY < <(compgen -W "add exists remove show help" -- "$current") ;;
remove) mapfile -t COMPREPLY < <(compgen -W "even odd range help" -- "$current") ;;
scan) mapfile -t COMPREPLY < <(compgen -W "adf flatbed help" -- "$current") ;;
sort) mapfile -t COMPREPLY < <(compgen -W "duplex move random reverse swap help" -- "$current") ;;
Expand Down
1 change: 1 addition & 0 deletions sources/command/help.sh
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,7 @@ Usage:
Core Commands:
extract Save specific pages into a new PDF
merge Combine PDFs into a single document
ocr OCR operations for PDF files
remove Delete specific pages from a PDF
scan Create PDFs directly from a scanner
split Split a PDF into multiple files
Expand Down
14 changes: 14 additions & 0 deletions sources/command/ocr/_index.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
function command_ocr()
{
local -r command="${1:-}"
[[ -n "${command}" ]] && shift

case "${command}" in
add) command_ocr_add "$@" ;;
exists) command_ocr_exists "$@" ;;
show) command_ocr_show "$@" ;;

help|--help|-h|'') command_ocr_help "$@" ;;
*) _log_error "Unknown command '$command'"; return 1; ;;
esac
}
58 changes: 58 additions & 0 deletions sources/command/ocr/add.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,58 @@
function command_ocr_add()
{
if (( $# < 2 )); then
cat <<TEXT
Add an OCR text layer to a PDF document

Usage:
pdfmt ocr add <input file> <output file>

Examples:
pdfmt ocr add input.pdf output.pdf

TEXT
return 0
fi


# Load
# -----------------------------------------------------------------------------------------------------------------
_config_load_ocr

local -r input_file="${1:-}"
local -r output_file="${2:-}"
local -r language="${OCR_LANGUAGE:-eng}"
local -r plugin="${OCR_PLUGIN:-}"


# Validate
# -----------------------------------------------------------------------------------------------------------------
_assert_is_installed ocrmypdf || return $?
_assert_is_file "${input_file}" || return $?
_assert_is_not_file "${output_file}" || return $?
_assert_matches_regex '^[a-z]{3}$' "${language}" "Invalid language format: ${language}" || return $?

if [[ -n "${plugin}" ]]; then
if ! ocrmypdf --plugin "${plugin}" --help &>/dev/null; then
_log_error "OCRmyPDF plugin is not installed or invalid: '${plugin}'" >&2
return 1
fi
fi


# Handle
# -----------------------------------------------------------------------------------------------------------------
# Skip OCR if the PDF already contains extractable text.
# Extract text stream and test if at least one alphanumeric character exists,
if pdftotext -enc UTF-8 "${input_file}" - 2>/dev/null | grep -q '[[:alnum:]]'; then
return 0
fi

# build options
local opts=(-l "${language}")
if [[ -n "${plugin}" && "${plugin}" != "tesseract" ]]; then
opts+=(--plugin "${plugin}")
fi

ocrmypdf "${opts[@]}" "${input_file}" "${output_file}"
}
39 changes: 39 additions & 0 deletions sources/command/ocr/exists.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,39 @@
function command_ocr_exists()
{
if (( $# < 1 )); then
cat <<TEXT
Check if a PDF document contains an OCR/text layer

Usage:
pdfmt ocr exists <input file>

Examples:
pdfmt ocr exists document.pdf

TEXT
return 0
fi


# Load
# -----------------------------------------------------------------------------------------------------------------
local -r input_file="${1:-}"


# Validate
# -----------------------------------------------------------------------------------------------------------------
_assert_is_installed pdftotext || return $?
_assert_is_file "${input_file}" || return $?


# Handle
# -----------------------------------------------------------------------------------------------------------------
# Extract text stream and test if at least one alphanumeric character exists.
if pdftotext -enc UTF-8 "${input_file}" - 2>/dev/null | grep -q '[[:alnum:]]'; then
echo "true"
return 0
else
echo "false"
return 1
fi
}
18 changes: 18 additions & 0 deletions sources/command/ocr/help.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,18 @@
function command_ocr_help()
{
cat <<TEXT
OCR operations for PDF files

Usage:
pdfmt ocr <command> [arguments]

Commands:
add Add an OCR text layer to a PDF document
exists Check if a PDF document contains an OCR text layer
show Print the OCR text layer to stdout

General Commands:
help Display this help message

TEXT
}
32 changes: 32 additions & 0 deletions sources/command/ocr/show.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,32 @@
function command_ocr_show()
{
if (( $# < 1 )); then
cat <<TEXT
Display the extracted OCR text layer of a PDF document

Usage:
pdfmt ocr show <input file>

Examples:
pdfmt ocr show document.pdf

TEXT
return 0
fi


# Load
# -----------------------------------------------------------------------------------------------------------------
local -r input_file="${1:-}"


# Validate
# -----------------------------------------------------------------------------------------------------------------
_assert_is_installed pdftotext || return $?
_assert_is_file "${input_file}" || return $?


# Handle
# -----------------------------------------------------------------------------------------------------------------
pdftotext -enc UTF-8 "${input_file}" -
}
28 changes: 28 additions & 0 deletions sources/common/config.sh
Original file line number Diff line number Diff line change
Expand Up @@ -119,3 +119,31 @@ EOF
local value=
value="$(_config_get_value "$config_file" "workflow.paths")"; export WORKFLOW_PATHS="$value"
}

function _config_load_ocr()
{
local -r config_dir="${HOME}/.config/pdfmt"
local -r config_file="${config_dir}/ocr.conf"

if [[ ! -f "${config_file}" ]]; then
mkdir -p "$config_dir"
cat <<EOF > "$config_file"
# Config for command: ocr
# =====================================================================================================================
# Language codes used for OCR text recognition.
ocr.language=eng

# Used OCRmyPDF plugin.
# Default: -empty- (Tesseract will be used)
# Possible options if installed:
# - ocrmypdf_rapidocr
# - ocrmypdf_easyocr
ocr.plugin=

EOF
fi

local value=
value="$(_config_get_value "$config_file" "ocr.language")"; export OCR_LANGUAGE="$value"
value="$(_config_get_value "$config_file" "ocr.plugin")"; export OCR_PLUGIN="$value"
}
1 change: 1 addition & 0 deletions sources/main.sh
Original file line number Diff line number Diff line change
Expand Up @@ -26,6 +26,7 @@ function main()
case "${command}" in
extract) command_extract "$@" ;;
merge) command_merge "$@" ;;
ocr) command_ocr "$@" ;;
remove) command_remove "$@" ;;
scan) command_scan "$@" ;;
sort) command_sort "$@" ;;
Expand Down
Binary file added tests/_data/no-ocr.pdf
Binary file not shown.
Loading
Loading