This project provides a complete workflow for processing OCR output 📄 and other text-bearing files. It takes the output of OCR — ALTO XML, PAGE XML, hOCR, ABBYY FineReader XML, OCR JSON — and transforms it into structured statistics tables 📊, performs text classification, and filters low-quality OCR 🔍 results. ALTO XML, the format of the PSNC corpus, stays the worked example throughout this README.
Note
Formerly atrium-alto-postprocess. On 1 October 2026 the project continued in this repository under
its new name, with the pipeline, the scripts, the categories and the scoring unchanged (the ALTO-specific
scripts keep their names: they name steps that really are ALTO-specific). What changed: the repository and image
names (ghcr.io/ufal/atrium-ocr-postprocess, …-api), the service id (atrium-ocr-postprocess), and the program id
this tool stamps into the document record (ocr-postprocess; records written before keep alto-postprocess, which the
record contract accepts as the same writer). The images of the old repository stay published for consumers pinned to
them and receive no new tags; its issues and releases remain readable there.
The core of the quality filtering relies on language identification 🌐 and a composite quality score 📈 — combining structural detectors, perplexity 📉, and character-level metrics — to identify and categorize noisy or unreliable OCR 🔍 output.
Besides ALTO, the same categorization takes the other OCR formats (PAGE XML, hOCR, ABBYY FineReader XML, DjVuXML, Tesseract TSV, OCR/Doc-AI JSON) and PDF, office, e-mail and plain-text files — see Step 1 and the input formats reference 📚.
- ⚙️ Setup
- 🛤️ Workflow Stages
- Vendored code 📦
- Acknowledgements 🙏
Before you begin, set up your environment.
- Create and activate a new virtual environment 🖥️ in the project directory.
- Install the required Python 🐍 packages:
pip install -r setup/requirements.txt
- Download the FastText 🌐 model for language identification:
The file is the Hub's
wget "[https://huggingface.co/facebook/fasttext-language-identification/resolve/main/model.bin](https://huggingface.co/facebook/fasttext-language-identification/resolve/main/model.bin)" -O lid.176.binmodel.bin— the NLLB fastText language-ID model, 1.18 GB, CC BY-NC 4.0 — saved under the name the pipeline's config expects (FASTTEXT_MODEL = lid.176.bin); it is not the 126 MB fastTextlid.176.bin. The Docker image fetches the same file pinned to a Hub revision and checked against its SHA-256 (theFASTTEXT_REVISION/FASTTEXT_SHA256build args in the Dockerfile 📎). - Copy the
v3folder from the 📐layoutreader🔧 repository 1 to the project directory for the LR-based text extraction method:git clone [https://github.com/ppaanngggg/layoutreader.git](https://github.com/ppaanngggg/layoutreader.git) cp -r layoutreader/v3/ ./ rm -rf layoutreader/
Note
Docker on Linux: run as yourself. The published images are ghcr.io/ufal/atrium-ocr-postprocess:<version>
(the batch pipeline) and ghcr.io/ufal/atrium-ocr-postprocess-api:<version>, where <version> is the release
without its leading v. ./data is part of the clone and belongs to you, while the images run as uid 10001 by
default: docker-compose.yml runs the services as user: "${ATRIUM_UID:-10001}:0", so put your uid in .env
once — echo "ATRIUM_UID=$(id -u)" >> .env. With docker run, pass --user "$(id -u):0". Docker Desktop
(macOS, Windows) needs neither. (atrium-project#69)
Note
There is no alto-tools 🔧 install step any more. The statistics and text-extraction code paths this
pipeline used are vendored in alto_tools.py 📎 — see Vendored code 📦.
You are now ready to start the workflow.
The process is divided into sequential steps, starting from raw ALTO 📄 files — or generic OCR JSON, or (#31) any other text-bearing file: PDF, DOCX, ODT, XLSX, PPTX, EPUB, RTF, HTML/hOCR, PAGE XML, TEI, ABBYY/DjVu XML, Tesseract TSV, Markdown, CSV, subtitles, e-mail or plain text, also compressed or as a ZIP of page files — and ending with extracted linguistic and statistic data 📊.
You can run the entire pipeline end-to-end with a single command (see below), or run each stage individually as described in Steps 1–4.
The run_pipeline.py 🐍 orchestrator runs every stage sequentially (split → statistics → text extraction → classification → aggregation) and, at the end, merges all per-stage paradata 🗒️ logs into a single run summary describing every stage, the intermediate file formats produced, and the effective end-to-end output license ⚖️ (see Paradata logging).
python3 run_pipeline.py # all settings from config.txt
python3 run_pipeline.py --method glm # override just the extraction backend
python3 run_pipeline.py --method json-keys --input-dir data_samples/JSON # generic OCR JSON (#31)
python3 run_pipeline.py --method text-lines # PDF/DOCX/TXT/... from [PIPELINE].INPUT_DIR_TEXT (#31)
python3 run_pipeline.py --skip-split # PAGE_ALTO already populated
python3 run_pipeline.py --dry-run # print the resolved plan, run nothing- Configuration ⚙️: every setting is read from config.txt 📎
(section
[PIPELINE], withINPUT_CSVtaken from[EXTRACT]). Precedence is CLI flag > config value > built-in default. Point at a different config with--configor theLANGID_CONFIGenvironment variable. - Extraction method 🔀:
[PIPELINE] METHODselects the Step 3 backend —alto-tools,layoutreader(default), orglmfor ALTO XML;json-keysfor generic OCR JSON;text-linesfor every other text-bearing format (#31). The method also selects the Step 1/2 scripts for its input format. The choice flows through to the merged license: a LayoutReader 📐 run resolves to CC BY-NC-SA 4.0, an alto-tools 🧰 run to CC BY-NC 4.0. --input-csv: the page statistics CSV. Step 2 writes it, and Steps 3 and 4 read it — every extractor andclassify_TEXT.pytake--input-csv(#31 Phase 5; the flag used to reach Step 2 only). Without the flag the ALTO/JSON methods read[EXTRACT].INPUT_CSV/[CLASSIFY].INPUT_CSVas before, andtext-linesuses its own CSV,[EXTRACT].INPUT_CSV_TEXT(data_samples/text_stats.csv), which the orchestrator passes to Steps 3 and 4 — so a text-lines run no longer overwrites the ALTO methods'data_samples/test_alto_stats.csv.- text-lines extras (#31): without
--input-dirit reads[PIPELINE].INPUT_DIR_TEXT(data_samples/TEXT);--strict/--no-strictreachtext_split.pyandextract_TEXT_2_TXT.py(a strict failure stops the pipeline);--source-origin ocr:<engine>reaches the split stage of every method. - Output 📤: a merged
<YYMMDD-HHmmss>_pipeline-run.jsonin the paradata 📁 directory, alongside the individual per-stage logs.
Note
Every stage — Step 1 (page_split.py / text_split.py) included — writes its own paradata
log, so a full run merges five logged stages, and the dry-run tags all five [paradata].
The merged license is re-derived from the union of components used across all stages, so
the end-to-end most-restrictive rule holds.
Tip
Prefer to inspect or re-run a single stage? The individual scripts below remain fully usable on their own — the orchestrator simply calls them in order.
First, ensure you have a directory 📁 containing your document-level input files. This script
will split them into individual page-specific files — it supports both ALTO XML and generic
JSON, dispatched automatically by file extension. Every other text-bearing format goes through
text_split.py instead (--method text-lines).
python3 page_split.py <input_dir> <output_dir>
Each page-specific file retains the header from its original source document 📌.
- Input 📥:
../ALTO/(input directory with ALTO XML 📄 documents) - Output 📤:
../PAGE_ALTO/(output directory with ALTO XML 📄 files split into pages)
Note
page_split.py splits every ALTO version: the namespace is the root element's (v2, v3 — what
the ATRIUM ABBYY exports use — v4, BnF), and each page file keeps it. A .xml whose root is not a
namespaced <alto> (PAGE XML, TEI, namespace-less ALTO) is skipped with its root named; the
text-lines method reads those. A re-split
replaces a document's page files: pages the input no longer yields are removed, also when the
input now fails. Tesseract numbers its ALTO pages from 0 (PHYSICAL_IMG_NR="0"), so its page
files are <doc>-0.alto.xml, <doc>-1.alto.xml, …
Example of the output directory with divided per-page XML files: PAGE_ALTO 📁.
PAGE_ALTO/
├── <file1>
│ ├── <file1>-<page>.alto.xml
│ └── ...
├── <file2>
│ ├── <file2>-<page>.alto.xml
│ └── ...
└── ...
Real OCR/Doc-AI engines don't agree on how they represent multi-page documents, so the JSON
path is detected heuristically rather than assuming one vendor's schema: a nested page-list
container (e.g. Azure Document Intelligence's/docTR's pages array), a flat element list
tagged with a per-item page field (e.g. AWS Textract), or — when neither pattern is found —
today's single-page-per-file default (e.g. pero-ocr, OCR.space).
- Input 📥:
../JSON/(input directory with generic JSON OCR-engine 📄 documents) - Output 📤:
../PAGE_JSON/(output directory with JSON 📄 files split into pages)
PAGE_JSON/
├── <file1>
│ ├── <file1>-<page>.json
│ └── ...
├── <file2>
│ ├── <file2>-<page>.json
│ └── ...
└── ...
python3 text_split.py <input_dir> <output_dir> [--source-origin ocr:<engine>] [--strict | --no-strict]
The <format>-2-txt step for everything that is not ALTO or json-keys JSON: PDF (text
layer), DOCX, ODT/ODS/ODP, XLSX, PPTX, EPUB, RTF, HTML/XHTML and
hOCR, PAGE XML, TEI/TEITOK, the OCR exports ABBYY FineReader XML, DjVuXML and
Tesseract TSV, any other XML, JSON/JSONL, CSV/TSV, Markdown, SRT/WebVTT
subtitles, EML/MBOX e-mail and plain text in UTF-8/16/32 or a legacy code page (cp1250
first). A gzip/bzip2/xz-compressed file is read through (scan.txt.gz), and a ZIP of
per-page files — a Transkribus or eScriptorium export — is one multi-page document. Formats
are recognised by content, not by extension, so a misnamed file is still read correctly and
an unreadable one is refused with a reason.
Every file becomes an ordered list of pages, each an ordered list of lines. A page is a real page where the format has one (PDF, PAGE XML, hOCR, ALTO, DOCX/ODT page breaks). Otherwise it is the format's natural block: a sheet, a slide, a JSON child object, a JSONL record, an EPUB chapter, a form-feed section of a text file. A line is the format's own unit: a physical line, a paragraph, a table cell or a spreadsheet row. The standard behind each format, the tools that write it and what this repo keeps of it (Formats and their standards), the full per-format matrix, the normalisation rules, the limits and the reason codes are in docs/text_inputs.md 📚.
- Input 📥:
../TEXT/(any mix of the formats above — example: data_samples/TEXT 📁) - Output 📤:
../PAGE_TEXT/—<doc_id>/<doc_id>-<n>.txtper page, plus two reports:ingest_report.csv: one row per input file, with its status (ok/partial/error/ignored) and reason code, the detected kind and encoding, page/line counts and thesource.originrecorded;pages_report.csv: one row per page, with its original label (sheet name, PDF page label, JSON page number, bundle member) and, for PDFs, the text-layer class (none/garbled/ocr/digital), a needs-OCR reason and report-only flags (mojibake_cp1252,mirrored_text=N,rotated_text=N).
A file that cannot be read costs that file only: images (image_needs_ocr), legacy
.doc/.xls (legacy_office_unsupported), encrypted, corrupt, empty or oversized files are
listed with their reason and the run continues. A file read with a possible loss (recovered XML,
a skipped bundle member, a stray CSV quote, a bad JSONL record) is partial: written, processed
downstream, its reasons listed. --strict turns any failure or partial read into exit code 1.
PDFs are read in an isolated child process with a timeout, ZIP containers and compressed files are
checked against bomb caps before anything is unpacked, and XML with entity declarations is refused.
A re-run never leaves stale pages of a document that now fails, and a page directory is only ever
replaced when it holds nothing but that document's page files. DOCX/ODT footnotes end the page
that cites them ([TEXT_INGEST].NOTES). All caps live in [TEXT_INGEST] in
config.txt 📎; each can also be set as the environment variable
ATRIUM_TEXT_INGEST_<KEY>, which wins over the file (atrium-project#53; the full list is
service/README.md § Limits).
PAGE_TEXT/
├── ingest_report.csv
├── pages_report.csv
├── <file1>
│ ├── <file1>-1.txt
│ └── ...
└── ...
This is the pipeline's first stage to see the original file, so when
[DOCUMENT].JSON_DIR is configured it is also the first writer of the record's source
block (sha256, filename, media_type, page_count, origin). source is
immutable — first writer wins — and since the atrium_document §1a hardening its
origin is what authorises this repo to write the positional blocks
(pages/content/lines/tables), so that no document can end up with half an
OCR-derived plane and half a digital-born one. A value that matches no known originator
prefix silently switches that check off, which is why the resolution is verified and a
mismatch warns 📣.
| Input | Default origin |
Meaning |
|---|---|---|
| ALTO XML | ABBYY-ALTO |
ALTO from the ABBYY toolchain the extractors already assume (see extract_LytRdr_ALTO_2_TXT.py) |
| JSON | ocr:generic |
Generic OCR/Doc-AI export whose specific engine the file does not name |
For text_split.py (#31) the default is truthful per format. The OCR outputs are
ABBYY-ALTO (ALTO), ocr:page-xml, ocr:hocr — for an ALTO or hOCR file that names its engine
in its header, that engine instead (ocr:tesseract, ocr:pero, ocr:kraken, ocr:transkribus,
ocr:ocrd, …; #31 Phase 5) —, ocr:abbyy-finereader, ocr:djvu,
ocr:tesseract, ocr:pdf-text-layer (a PDF whose text is an invisible OCR layer under the page
image) and ocr:generic (TXT, Markdown, CSV/TSV, JSON/JSONL, TEI, XML, subtitles); a ZIP bundle
takes its members' origin. Born-digital documents are digital-born-<kind> (DOCX, ODT/ODS/ODP,
XLSX, PPTX, EPUB, RTF, plain HTML, e-mail, a PDF with visible text). A born-digital record is
originated by digital-convert, so this repo then writes only source into
it: document_hook holds back the pages/content/lines blocks of every stage
(atrium_document §1a), unless a page carries the needs_ocr hand-off. The CSV outputs of the
run are produced either way. --source-origin ocr:<engine> overrides the default when the
files are known OCR output.
digital-convert reads only PDF and DOCX, so born-digital records of the other kinds keep source
only. When files of such a kind really are OCR or transcription exports, say so per kind:
[DOCUMENT].SOURCE_ORIGIN_BY_KIND = xlsx = ocr:generic, pptx = ocr:generic. For text-lines the
precedence is CLI flag > env var > SOURCE_ORIGIN_BY_KIND > SOURCE_ORIGIN > per-format default
(docs/text_inputs.md §6).
Override it when the engine is known — the prefix must stay one this repo owns
(ABBYY-ALTO, ocr:<engine>, vlm:<engine>):
python3 page_split.py <input_dir> <output_dir> --source-origin ocr:pero
Also settable as [DOCUMENT].SOURCE_ORIGIN in setup/config.txt 📎
or via the DOCUMENT_SOURCE_ORIGIN env var, so orchestrated run_pipeline.py runs
(which invoke this script with no extra CLI flags) can set it too. Precedence: CLI flag >
env var > config value > per-format default.
Next, use the output directory from Step 1 as the input for this script to generate a foundational CSV 📊 statistics file.
python3 alto_stats_create.py <input_dir> -o output.csv # ALTO XML
python3 json_stats_create.py <input_dir> -o output.csv # generic JSON (json-keys)
python3 text_stats_create.py <input_dir> -o output.csv # text-lines (#31)
The three builders write the same columns. text_stats_create.py takes the document id from the
<doc_id>/ folder, so ids containing - stay whole. It counts non-blank lines as textlines,
words as strings and PDF image objects as illustrations.
This script writes a CSV 📊 file line-by-line, capturing metadata for each page:
file, page, textlines, illustrations, graphics, strings, path
CTX200205348, 1, 33, 1, 10, 163, /lnet/.../A-PAGE/CTX200205348/CTX200205348-1.alto.xml
CTX200205348, 2, 0, 1, 12, 0, /lnet/.../A-PAGE/CTX200205348/CTX200205348-2.alto.xml
...
The element counting is powered by the alto-tools 🔧 statistics code ^1, vendored in alto_tools.py 📎 (Apache-2.0) — see Vendored code 📦. No external binary is called and nothing is downloaded at run time.
- Input 📥:
../PAGE_ALTO/(input directory with ALTO XML 📄 files split into pages from Step 1) - Output 📤:
output.csv(table with page-level statistics and paths to ALTO files)
Important
This statistics table is the basis for subsequent processing steps. Example: test_alto_stats.csv 📎.
This script runs in parallel ⚡ (using multiple CPU 💻 cores) to extract text from ALTO XMLs 📄 into .txt 📝 files.
It reads the CSV 📊 from Step 2 — --input-csv <stats.csv>, else [EXTRACT].INPUT_CSV — and every
extractor keeps document ids such as 0001 as text. A page whose .txt is older than its input page
file is extracted again on the next run; one that is current is skipped (resume).
- Input 1 📥:
output.csv(from Step 2) - Input 2 📥:
../PAGE_ALTO/(input directory with ALTO XML 📄 files split into pages from Step 1) - Output 📤:
../PAGE_TXT/or../PAGE_TXT_LR/(directory containing raw text 📝 files)
Caution
The model responsible for spatial layout 📐 analysis requires a GPU 🚀 to run efficiently.
python3 extract_LytRdr_ALTO_2_TXT.py
Uses the LayoutReader 📐 framework ^9 to extract text and bounding boxes of XML 📄 elements
(specifically, <TextLine> elements containing Strings with CONTENT attribute),
process them to reconstruct the reading order of lines (columns-friendly), handle words split
between two lines (adding the full form of the word), and group page contents into paragraphs
based on the vertical spread of text lines.
Example of per-page text files: PAGE_TXT_LR 📁.
PAGE_TXT_LR/
├── <file1>
│ ├── <file1>-<page>.txt
│ └── ...
├── <file2>
│ ├── <file2>-<page>.txt
│ └── ...
└── ...
Note
The method is CPU 💻-bound and faster than the LayoutReader method, but the text lines may not be in the correct reading order, and full forms of hyphenated split words are not reconstructed.
python3 extract_ALTO_2_TXT.py
Uses the alto-tools 🔧 text extractor ^1 — vendored in
alto_tools.py 📎 (Apache-2.0), see Vendored code 📦 — to extract text lines from
XML 📄 elements directly, with no post-processing. Suitable for a quick overview of raw text content.
The output is byte-identical to what the upstream alto-tools -t command line produced.
Example of per-page text files: PAGE_TXT 📁.
PAGE_TXT/
├── <file1>
├── <file2>
│ ├── <file2>-<page>.txt
│ └── ...
└── ...
Warning
The method is GPU 🚀-bound, slower than the LayoutReader method, and requires a gpuram48G card.
python3 extract_LLM_ALTO_2_TXT.py
Uses the GLM-4v-9b 🤖 multimodal large language model ^10 to perform generative OCR 🔍 directly from
page images, prompted as Transcribe all text on this page exactly as it appears. The script
trims whitespace and resizes high-resolution images to fit model constraints.
Note
This method is significantly slower than parsing XML 📄 but often yields higher quality text for complex layouts 📐 or degraded scans. It patches the transformers configuration to run the GLM-4v architecture.
Example of per-page text files: PAGE_TXT_LLM 📁.
PAGE_TXT_LLM/
├── <file1>
├── <file2>
│ ├── <file2>-<page>.txt
│ └── ...
└── ...
Note
Use this method with the Generic JSON input split from Step 1 — the other three methods above all consume ALTO XML 📄 instead.
python3 extract_JSON_2_TXT.py
Reads each page's generic OCR/Doc-AI JSON 📄 and walks a whitelist of informative keys
(content, text, line, word, ... — no assumption about a particular vendor's schema
beyond "text lives under a key named roughly text/line/word"), yielding the page's text
in document order, each text once at line granularity (see the note below). This is the Extraction Layer: it makes no other change to the
JSON and does not know about doc.json at all.
Example input: data_samples/JSON 📁 (an Azure-style analyzeResult.pages document).
Note
json-keys re-serialises each split page with its document header, but reads only the page:
the page object a page file holds (Azure's whole-document analyzeResult.content no longer
opens every page), each text once at line granularity (a line's words[] are not repeated
after it; AWS Textract's WORD blocks are dropped when LINE blocks exist) — the rules the
text-lines method uses for the
same JSON (#31 Phase 5). On real AWS Textract responses and Microsoft's Azure example result both
methods give exactly the engine's LINEs per page; json-keys used to write 3–6× as many lines.
PAGE_TXT_JSON/
├── <file1>
│ ├── <file1>-<page>.txt
│ └── ...
├── <file2>
│ ├── <file2>-<page>.txt
│ └── ...
└── ...
When [DOCUMENT].JSON_DIR (or the DOCUMENT_JSON_DIR env var) is
configured, a second, separate Accretion Layer runs after extraction: for every document,
it reads back that document's already-written .txt pages and merges them into
<doc_id>.document.json as this repo's owned fields — pages[].ocr
(ocr-postprocess's share of the field-split pages block) and the whole content block —
leaving every other block (page_categories, entities, source, ...) untouched, per the
atrium_document.schema.json paired-hook contract. With no baseline doc.json yet on disk,
the record is created holding just this contribution (accretion rule 3). This mirrors what
the other three extraction methods above already do.
--force-single-page is a document-assembly policy, not an extraction-rule change — it
only decides how many pages[] rows describe the pages already extracted above:
python3 extract_JSON_2_TXT.py --force-single-page
pages[] shape |
content.text |
|
|---|---|---|
| default | one row per source page (page: "1", "2", ...) |
always the full document: every page's text concatenated in source order |
--force-single-page |
every source page collapses into one row (page: "1"), with ocr.source_pages listing the original page labels in concatenation order |
unchanged — same joined text either way |
--force-single-page can also be set as a config-file default —
[EXTRACT].FORCE_SINGLE_PAGE_JSON = true in setup/config.txt — so
run_pipeline.py orchestrated runs (which invoke this script with no extra CLI flags) can
still opt in. The CLI flag takes precedence over the config value when both are given.
Note
Use this method with text_split.py output from Step 1
and text_stats_create.py from Step 2.
python3 extract_TEXT_2_TXT.py [--input-csv CSV] [--output-dir DIR] [--lines-dir DIR] [--strict | --no-strict]
Prepares each page's text for categorization. Lines are normalized (NFC; control, zero-width
and bidi characters removed; soft hyphens resolved). Blank lines are dropped
([TEXT_INGEST].KEEP_BLANK_LINES). Any line longer than [TEXT_INGEST].MAX_LINE_CHARS
(default 1000) is wrapped at a word boundary, because one giant paragraph would otherwise pad a
whole perplexity batch to the model's full context. A paragraph, cell or physical line otherwise
stays one line.
- Output 📤:
../PAGE_TXT_TEXT/<file>/<file>-<page>.txt(what Step 4 reads) - Output 📤:
../DOC_LINES_TEXT/<file>.csv: the line table, i.e. the input's ordered text lines as CSV rows (file,page_num,line_num,text,page_label) before any categorization. Itspage_num/line_numare exactly the numbers Step 4 gives the same lines inDOC_LINE_CATEG. The script refuses to write it into theDOC_LINE_CATEGdirectory.
Page files in any encoding are accepted, so a hand-made directory of page TXT files works
with --skip-split. With [DOCUMENT].JSON_DIR set, pages[].ocr (engine: text-lines) and
content are accreted like the other methods (not into born-digital records, see Step 1). A
page, a line table or a document record that fails costs that page or document only; --strict
(or [TEXT_INGEST].STRICT) turns such a failure into exit code 1.
This is a key ⌛ time-consuming step that analyzes the text quality 📈 of each page line-by-line, assigning each line a quality category to filter out OCR 🔍 noise.
It uses the FastText language identification model 🌐 and perplexity 📉 scores from Qwen2.5-0.5B 🤖 to detect noise ^2 ^6.
More post-processing of TXT 📝 files can be found in the GitHub repository of the ATRIUM project, which covers NLP enrichment using Nametag for NER and UDPipe for CONLL-U files with lemmas & POS tags ^5.
As the script processes, it assigns each line one of five categories 🪧:
| Category | Action | Description |
|---|---|---|
| ✅ Clear | Ready to be processed by further NLP | Passes all structural checks; high composite quality score 📈. |
| Corrections of generally readable words are needed | Partially degraded: moderate quality score 📈 indicating isolated symbol issues, fused tokens, mid-word uppercase, or elevated perplexity 📉. | |
| 🗑️ Trash | Should be re-processed by another OCR 🔍 tool | Severely corrupted: composite quality score 📈 below the Trash threshold, or routed here by an override (unreadable all-caps line, inverted-scan page block). |
| 🔣 Non-text | May be checked for identifiers of finds/sites | Filtered by the CPU 💻 pre-filter: line is too short, has too few unique symbols, contains fewer than 30% alphabetic characters, or consists mostly of digits and punctuation. |
| 🫙 Empty | Can be ignored | Line contains only whitespace (paragraphs separator) |
Note
The table above describes what the program does with a line. It is not the definition an
annotator works to, and the two were drifting apart. For issue #30 the definition was settled by
@david-spacil on 2026-09-22, pending @DanaKriv: Trash means illegible. Anything legible is
Clear; anything easily decipherable through a mistake is Noisy — regardless of how useful
the line is. So a correctly scanned web address is Clear, not Trash. Non-text cannot be
distinguished from Trash without the page image and is out of scope for hand annotation.
The Trash row above is already consistent with this: re-process rather than delete.
See docs/issue30/issue30_review_request.md § 6.
Note
Hand-editable word lists (issue #30). The word lists the categoriser consults — units, reference
labels, section headings, short Czech function words, the rotation whitelists — live in
setup/word_lists.txt (WORD_LISTS_PATH), one [section] each, with the
matching setup/config.txt keys kept empty as overrides. Its [allowed] section is for real words
the archive uses that the program keeps mistaking for damage; since 2026-10-01 it ships the entries
the data providers reviewed (issue #30 Q5a), with the rest commented out. A listed word stops counting
against the quality score and is never evidence of damage to the shape witness (Q5b) — see
docs/categorization_logic.md.
Separately, DOMAIN_NOTATION_CATEG in [TEXT_UTILS] names the category every recognised web or
e-mail address gets; it ships Clear (issue #30 Q4; empty = off).
Note
This script generates two primary output directories:
DOC_LINE_LANG_CLASS/ and DOC_LINE_STATS/, while the
raw text 📝 files (primary input) are stored in ../PAGE_TXT/ generated from ../PAGE_ALTO/.
All input/output paths and tunable parameters are configured ⚙️ in config.txt 📎.
Parameters are organized into three sections: [CLASSIFY], [AGGREGATE], and [TEXT_UTILS].
[CLASSIFY]
BATCH_SIZE = 128 # Batch size for processing lines
WORKERS_MAX = 32 # Max CPU workers for parallel tasks
EXPECTED_LANGS = ces,deu,eng # Expected languages (ISO codes); first is default
TRUSTED_FOREIGN_LANGS = deu,eng,fra,pol,ita # Allowed foreign languages (ISO codes)
MODEL_NAME = Qwen/Qwen2.5-0.5B # Language model for perplexity scoring; English-only collections: distilgpt2
[TEXT_UTILS]
QS_WEIGHT_VALID_WORD = 0.35 # Weight for valid word ratio in QS
QS_WEIGHT_WEIRD = 0.18 # Weight for inverted word weirdness in QS
QS_WEIGHT_PERPLEXITY = 0.08 # Weight for inverted normalized perplexity in QS
QS_WEIGHT_LENGTH = 0.02 # Weight for length reward in QS
QS_WEIGHT_GARBAGE = 0.18 # Weight for inverted garbage density in QS
QS_WEIGHT_VOWEL = 0.07 # Weight for vowel quality in QS
QS_WEIGHT_LANG = 0.05 # Weight for language confidence in QS
QS_WEIGHT_GIBBERISH = 0.04 # Weight for inverted gibberish ratio in QS
QS_WEIGHT_FUSED = 0.03 # Weight for inverted fused word ratio in QS
QS_LENGTH_MAX = 100.0 # Max length for normalization
CATEG_TRASH_SCORE_MAX = 0.55 # Max QS for Trash category
CATEG_NOISY_SCORE_MAX = 0.80 # Max QS for Noisy category (#3 2026-07-02: lowered 0.85 -> 0.80)
REPEATED_DOUBLE_MIN = 2 # Minimum occurrence count for doubled-char penalty
SHORT_NOISY_QS_PENALTY = 0.20 # Opt-in QS penalty for short strings exhibiting OCR oddities
# --- New since last revision: Phase-2 categoriser overrides ---
LOWPPL_CLEAR_MAX = 50.0 # ppl ceiling for Override 3 (was hardcoded)
HARD_SWEEP_LANG_MAX = 0.45 # orig_lang_score ceiling for the hard-sweep route
HARD_SWEEP_PPL_MIN = 1000.0 # ppl floor for the hard-sweep route
GHOST_DOMINATED_MIN_RATIO = 0.5 # min ghost-token share to flag ghost_dominated
WORD_W_PENALTY = 0.20 # per-word weirdness penalty for tokens containing 'w'
ROT_HIGH_LANG_CONF = 0.90 # lang_score ceiling for the page-level rotation arm
Parameters that scale with the perplexity 📉 model:
These parameters must be re-tuned whenever you switch between multilingual Qwen2.5-0.5B🤖 and English-adapted distilgpt2🤖,
because the two models produce perplexity 📉 on very different numerical scales — Qwen2.5-0.5B🤖 assign scores roughly 3× lower
than distilgpt2🤖 on the same Czech 🇨🇿 text:
| Parameter | Qwen2.5-0.5B | distilgpt2 | What it controls |
|---|---|---|---|
PERPLEXITY_THRESHOLD_MAX |
1000.0 | 3000.0 | The ceiling used to normalise raw perplexity 📉 into [0, 1] for the quality score 📈. A value at or above this ceiling contributes 0 to the score (worst); a value of 0 contributes 1 (best). |
SHORT_PPL_CAP |
850.0 | 2500.0 | Maximum perplexity 📉 applied to 1–2 word lines before quality scoring. Short text fragments receive extreme perplexity 📉 scores from any LM because there is no context to condition on; this cap prevents legitimate short labels and codes from being unfairly penalised. |
PPL_INVERTED_MIN |
200.0 | 500.0 | Perplexity 📉 floor for the inverted-scan detection arm. A line is considered a candidate for the inverted-scan penalty only if the LM is also uncertain about it (perplexity 📉 above this value). |
CLEAN_PROSE_PPL_MAX |
400.0 | 1000.0 | Maximum perplexity 📉 a line may have to qualify for the near-boundary Clear promotion (Override 4). Lines with perplexity 📉 above this value are not promoted even if all other conditions are met. |
Parameters that are model-independent 🤖 and stable across different choices of perplexity 📉 model 🤖:
These parameters are expressed as ratios or quality-score fractions, not as perplexity 📉 values, so their meaning does not change between models and their defaults are stable across either choice:
| Parameter | Default | What it controls |
|---|---|---|
ROT_RATIO_INVERTED_MIN |
0.55 | Minimum fraction of structurally rotatable characters (pbqdnuwmoxszeyv) among alphabetic characters that must be present before a rotation penalty is even considered. A value of 0.55 means more than half of all letters in the line must belong to this ambiguous set. |
WEIRD_RATIO_INVERTED_MIN |
0.35 | Minimum mean per-word weirdness score required to confirm an inverted scan when rot_ratio is already above the threshold. This second condition prevents Czech 🇨🇿 sentences that happen to contain many p, d, b, q letters from being falsely penalised. |
CLEAN_PROSE_MIN_SCORE |
0.65 | Lower bound of the quality-score range within which the near-boundary promotion (Override 4) can fire. A line must score at least this well before it is a candidate for promotion from Noisy to Clear. |
CLEAN_PROSE_WEIRD_MAX |
0.08 | Maximum mean per-word weirdness a line may have to qualify for the near-boundary promotion. Even a single notably corrupted token disqualifies the line from being promoted. |
CLEAN_PROSE_WC_MIN |
4 | Minimum word count a line must have to qualify for near-boundary promotion. Very short lines (1–3 words) have unreliable perplexity 📉 scores and are therefore never promoted regardless of their quality score 📈. |
MOSTLY_READABLE_VALID_MIN |
0.85 | Minimum ratio of structurally valid words required. Semi-readable lines dipping below this ratio are capped at Noisy and prevented from achieving Clear. |
Note
The CLEAN_PROSE_* rows above (CLEAN_PROSE_PPL_MAX, CLEAN_PROSE_MIN_SCORE, CLEAN_PROSE_WEIRD_MAX,
CLEAN_PROSE_WC_MIN) parameterise the near-boundary "Override 4" clean-prose promotion, which has been
removed from determine_category() (see the callout in
Categorisation Logic). These
keys — together with the never-implemented CLEAR_BAND_WC_MIN guard — have now been removed from
config.txt as well (#7 Phase 0 of the config-coverage audit); they are not read by any current
scoring or categorisation path. The rows are kept here only as historical documentation of the removed override.
Language- and collection-specific data 🇨🇿 moved from hardcoded Python literals into the config (#7 Tier 1). Defaults are bit-identical to the previous in-code values, so the shipped config produces exactly the same categorisation:
| Parameter | Section | Default | What it controls |
|---|---|---|---|
DEU_DIACS |
[TEXT_UTILS] |
äöüßÄÖÜ |
German diacritic glyphs 🇩🇪; together with CZ_DIACS rebuilds the per-language diacritic map used by infer_lang_from_diacritics(). |
DIACRITIC_INFER_THRESHOLD |
[TEXT_UTILS] |
0.07 | Minimum diacritic share among alphabetic characters for diacritic-based language inference. |
WQX_CHARS |
[TEXT_UTILS] |
wqxWQX |
Letters rare in Czech 🇨🇿 — wqx-heavy tokens signal OCR noise in score_word, score_words_in_line and determine_category. |
ROT_WHITELIST |
[TEXT_UTILS] |
po,pod,do,od,on,ony,by,bez,ne,nebo,ven,den,zde,se,ve,mez,pouze,bude |
Czech 🇨🇿 function words recognisable upright; their mirror/rotation ghost images (ROT_GHOSTLIST) are derived at import time — changing this key requires re-import (override_constants() does not rebuild it). |
GHOST_WORD_COLLISIONS |
[TEXT_UTILS] |
no,bo |
Ghost images that collide with real words and must never count as ghost hits. |
TRAILING_FILL_CHARS |
[TEXT_UTILS] |
\x20._:-<\u2013\u2014 |
Trailing filler characters stripped before headline/short-line checks. Unicode-escape decoded — the leading space is written as \x20 because configparser strips leading whitespace from values. |
NONTEXT_MARKERS |
[TEXT_UTILS] |
IVerc |
Collection-specific literal markers (ARUP/B stamp) forcing the Non-text route in pre_filter_line(). |
FASTTEXT_MODEL |
[CLASSIFY] |
lid.176.bin |
Path to the FastText 🌐 language-ID weights loaded by each CPU worker. |
TRUST_TIER_TRUSTED |
[CLASSIFY] |
0.85 | Trust multiplier on the FastText 🌐 confidence for a known but unexpected language. The product (trust_lang_score) is what feeds both the quality score 📈 and the structural gates — not the stored lang_score. |
TRUST_TIER_UNKNOWN |
[CLASSIFY] |
0.50 | Trust multiplier for an unknown language. Because it caps trust_lang_score at 0.50 — below LANG_SCORE_REMAP (0.75) — gates of the form lang_score <= LANG_SCORE_REMAP are always true for unknown-language lines. |
REMAP_KEEP_SCORE_LANGS |
[CLASSIFY] |
slk |
Languages that keep their original FastText 🌐 confidence when remapped to the default language (Slovak ≈ Czech 🇨🇿, so the confidence stays meaningful after the label swap). |
This script reads the extracted text 📝 files, batches lines together 📦, and runs the FastText 🌐 and Qwen2.5-0.5B 🤖 models. It uses a CPU 💻/GPU 🚀 split architecture:
- A single dedicated GPU 🚀 worker holds the only Qwen2.5-0.5B 🤖 instance and processes perplexity 📉 batches to prevent VRAM OOM errors.
- Multiple CPU 💻 workers (up to
WORKERS_MAX, default 32) read files, run FastText 🌐 and structural detectors, and submit text batches to the GPU 🚀 worker via a shared queue. CPU 💻 workers poll the result dictionary while the GPU processes, running language identification 🌐 concurrently.
Warning
The first item of EXPECTED_LANGS list of languages 🌐 should be the most expected language in the processed
collection to work as a default replacement of ambiguous language recognition predictions.
python3 classify_TEXT.py [--input-csv <stats.csv>]
- Input 1 📥:
../PAGE_TXT/from Step 3 - Input 2 📥:
output.csvfrom Step 2 (--input-csv, else[CLASSIFY].INPUT_CSV) - Output 📤:
DOC_LINE_LANG_CLASS/containing per-document CSVs 📊 (e.g., DOC_LINE_CATEG 📁)
Tip
This script is resume-capable. If interrupted, run it again and documents whose output CSV is
at least as new as all of their page texts are skipped; a document with a page text newer than
its output (re-ingested, re-extracted) is classified again. Output CSVs of documents that are not
in the page CSV are listed in a warning, never deleted — aggregate_STAT.py aggregates every CSV
in the directory. The Step 3 extractors resume the same way, page by page, and Steps 1 and 3 rewrite a
page file only when its content changed, so a full re-run over unchanged inputs skips everything again.
<doc_name>.csv 📊: Detailed classification results for every single line within a document, columns:
file— document identifier 🆔page_num— page number 📄line_num— line number, starts from 1 for each page 🔢text— cleaned text of the line 📝original_text— original pre-repair text of the line 📝split_ws— hyphenated word prefix at the end of the line (split word start)split_we— hyphenated word suffix at the start of the line (split word end)word_count— count of whitespace-delimited tokens in the line (count of words)char_count— count of total character in the cleaned linegarbage_density— ratio of non-alphanumeric characters to total line length (calculated onoriginal_text)upper— count of words with unexpected mid-word uppercase lettersrepeated— count of words where a non-standard character makes up ≥ 30% of the word, or containing consecutive doubled garble charactersldl_fuses— count of words with a letter–digit–letter sandwich (e.g.,vyt1ačená), excluding valid measurements.fused_words— count of tokens that appear to be fused words (abnormal consonant/vowel runs or extreme length)gibberish— count of words flagged as gibberish (high vowel ratio)weird_wx— count of words with an abnormal density of 'w' or 'x' glyphsword_weird— mean per-word weirdness score in [0, 1]; combines strange-symbol (0.40), repeated-char (0.35), LDL-fusion (0.15), mid-uppercase (0.10), and aWORD_W_PENALTY-weighted (default 0.20) signal for tokens containing the letterw— rare in Czech and a strong inverted/mirror-OCR fingerprint — plus a separate caps-prefix penalty (0.20). The combined score is clamped to [0, 1]. Isolated single letters score 0.85 (OCR noise) or 0.25 (digit/measurement).vowel_ratio— ratio of vowel characters to total alphabetic and symbol characters in theoriginal_textrot_ratio— the ratio of structurally ambiguous/rotatable characters (pbqdnuwmoxszeyv) to the total number of alphabetic characters in the line.
<doc_name>.csv's key resulting output columns that depict the final classification and quality assessment:
quality_score— composite quality score 📈 in [0, 1] based on 9 combined signals; higher = cleanercateg— assigned category: Clear ✅, Noisy⚠️ , Trash 🗑️, Non-text 🔣, or Empty 🫙
<doc_name>.csv's columns useful for archive managers information apart from the quality score 📈 and category:
lang— predicted ISO language code from the FastText 🌐 model (remapped if unknown)lang_score— FastText 🌐 confidence score for the predicted language (capped if remapped, #3)original_lang— predicted language before remapping logicorig_lang_score— original FastText confidence before remappingperplex— Qwen2.5-0.5B 🤖 perplexity 📉 score of the line 📉caps_header— boolean flag indicating whether all alphabetic words in the line are uppercase (typical of section headers)
Diagnostic flags (#3):
Ten boolean audit columns follow caps_header. Six name the categoriser rule that decided the line —
allcaps_novowel, lowppl_clear, cleanprose_clear, trash_threshold, noisy_threshold, clear_threshold
(exactly one True, or none for Empty). Two further internal reason codes — trash_hard_sweep (route 1a)
and trash_inverted (route 1b) — also exist but are folded into the trash_threshold column rather than
getting their own column, so the per-line reason granularity is coarser in the CSV than inside the categoriser.
Four name the document-level post-pass that later changed it — pp_dedup (header/footer mode-harmonisation),
pp_surrounded_trash (rolling-window smoothing), pp_inverted_run (page-level inverted-scan sweep), and
pp_page_context (page-context Trash/Noisy adjustment, see below).
A categoriser flag and a pp_ flag may both be True on one line: that is the intended trail from the
original decision to the override.
The full decision logic — CPU pre-filter, language handling, structural detectors, the composite quality score, the categorisation gates/rescues, and the four document-level post-processing passes — lives in docs/categorization_logic.md.
Important
Known accuracy limitation since v1.5.0-beta (issue #30).
rule_short_garbage no longer convicts short diacritic-free lines on shape alone. Measured over
both collections (113,101 documents / 71.8M lines) the change promoted 367,208 lines and demoted
none, dropping Trash by 247,252 lines (−8.68%), and human grading of 484 promoted lines scored
76.5% right / 9.6% borderline / 13.9% wrong — a 5.5 : 1 improvement.
The cost is real and was accepted deliberately: roughly 26,000 genuine OCR-garbage lines now reach
Clear (95% CI [19,300; 34,600]). Clear is the top category, not a reviewable middle state, so
a Clear verdict on a 1–3 token line without Czech diacritics carries less confidence than the
label suggests. Downstream consumers filtering on categ == "Clear" should be aware of it.
The narrowing that pays this down (_has_shape_garbage_evidence(), a phonotactic second witness) is
implemented and wired, but ships disabled (SHORT_GARBAGE_WITNESS_ENABLE = false). It has been
measured against the 2,067-line human gold set in
tools/gold/sidecars/issue30_gold_2067.csv (see
tools/gold/GOLD.md): with the document-level smoothing it passes the adoption
gate — errors 513 → 503, readable lines lost 38 → 38, McNemar p = 0.013, and 503 → 502 / 38 → 37 with
the German/French vowel-run split — but it fails that gate per line, so the smoothing is a
precondition. It stays off until the data providers' annotation of the population it would newly
discard, and their agreement on what Trash means, are in. See
docs/issue30/README.md.
| Section | What it covers |
|---|---|
| CPU 💻 Pre-filter | pre_filter_line(): Empty/Non-text routing and the two OCR repairs, before any model runs |
| Language 🌐 Handling | FastText trust tiers, remap_lang(), and the LANG_REMAP_ALWAYS switch |
| Structural Detectors | the per-line signal detectors (rotation, gibberish, fused/vowel-less words, damage) |
| Composite Quality Score | the weighted compute_quality_score() formula and its dynamic adjustments |
| Categorisation Logic | determine_category(): the ordered gates, check_rescues(), and threshold routing |
| Post-Processing Smoothing | apply_document_postprocessing(): dedup, rolling window, page context, inverted-scan sweep |
To replay a config change over already-scored CSVs without a GPU, use
tools/recategorize_from_csv.py — it reads the same constants and
calls the very same scoring function as the pipeline (classify_TEXT.score_line()), so a re-scored
corpus differs from the shipped batch only by the change under test. The FastAPI /process endpoint
uses that same function too; see
Line Categorisation Logic, which opens with the three callers and
what each of them supplies.
Example of per-document CSV 📊 files: DOC_LINE_CATEG 📁 by Qwen2.5-0.5B 🤖 and DOC_LINE_CATEG_gpt 📁 by distilgpt2 🤖.
DOC_LINE_LANG_CLASS/
├── <docname1>.csv
├── <docname2>.csv
└── ...
This script processes the DOC_LINE_LANG_CLASS/ directory with CSV 📊 files in chunks 🧩 to produce
final page-level statistics. It is CPU 💻-bound and parallelized with ProcessPoolExecutor.
python3 aggregate_STAT.py
- Input 📥:
DOC_LINE_LANG_CLASS/(directory with CSV 📊 files from the previous step) - Output 1 📤:
final_page_stats.csv📊 (configurable viaOUTPUT_STATS) — global page-level summary across all documents - Output 2 📤:
DOC_LINE_STAT/(configurable viaOUTPUT_DOC_DIR) — per-document CSVs 📊 with the same schema
For each page, the aggregation computes features outputted in the following strict schema order:
Totals & Counts:
num_lines— the total number of valid lines processed on the pageClear,Noisy,Trash,Non-text,Empty— integer count of lines in each categorytotal_word_count— total number of words across scoreable linestotal_char_count— total number of characters across scoreable lines
Averages (mean over the same Clear ✅ and Noisy
avg_quality_score— mean composite quality score 📈 in [0, 1]; higher = cleaner OCR 🔍 outputavg_word_weird— mean per-word weirdness ratio in [0, 1]; 0 = fully clean, lower is better 📉avg_lang_score— mean FastText 🌐 confidence scoreavg_perplex— mean Qwen2.5-0.5B 🤖 perplexity 📉 scoreavg_vowel_ratio— mean vowel-to-alphabetic-character ratio per lineavg_rot_ratio— mean rotatable character ratio per linech_ratio— mean fraction of lines flagged as all-caps headers (caps_header = True)
Language profile:
main_lang— the statistical mode (most frequent) language 🌐 predicted for the page
Note
avg_* columns and main_lang will be NaN / None for pages whose only lines are
Empty or Non-text (i.e., pages with no scoreable text content).
Additional per-line diagnostic variables (e.g. weird_wx, original_lang, original_text) and flags added
and ignored for this page-level aggregation to ensure stability.
All numeric averages are rounded to 4 decimal places; totals are stored as integers.
- Examples: arup_page_stats_SHORT.csv 📊, arub_page_stats_SHORT.csv 📊
Example of per-document aggregate CSV 📊 files: DOC_LINE_STATS 📁 by Qwen2.5-0.5B 🤖 and DOC_LINE_STATS_gpt 📁 by distilgpt2 🤖:
DOC_LINE_STAT/
├── stats_<docname1>.csv
├── stats_<docname2>.csv
└── ...
This is the end of the text quality classification and filtering step. You can now use arup_page_stats_SHORT.csv 📎 to identify files that need another round of OCR 🔍 or manual correction based on the line type counts. Pages with the majority of Clear ✅ lines can be marked for further processing. The absence of clear lines combined with a high proportion of Trash 🗑️ lines may also indicate handwritten content, which can be excluded before Handwritten Text Recognition (HTR) is applied.
In addition to the batch pipeline, this repository ships with a FastAPI wrapper (service/text_api.py) that exposes
the core text_util_langID quality classification engine over HTTP. The /process endpoint accepts ALTO XML,
plain-text, and generic JSON uploads (task_type alto / text / json, or auto-detected from the file
extension), returning the same per-line classification fields as the batch pipeline for all three formats.
(#31) It also accepts every text-lines format (PDF, DOCX, ODT, XLSX, PPTX, EPUB, RTF, HTML/hOCR, PAGE XML,
TEI, ABBYY/DjVu XML, Tesseract TSV, Markdown, CSV/TSV, JSONL, SRT/VTT, EML/MBOX, compressed files and ZIP bundles)
as task_type document. auto keeps .txt/.json as before and decides .xml and every other upload
from its bytes: an uncompressed ALTO root stays on the ALTO path, a file of a kind the service does not read is
a 415 unsupported_media_type naming the reader's code (a 400 before atrium-project#32 round 2), a supported file
that cannot be read — a document without a single text line
included (no_text) — is a 422, and one over a [TEXT_INGEST] size or count cap is a 413 limit_exceeded.
Document results carry page/page_label per line and a pages summary, and are read and shaped by the same
text_formats.py code and the same [TEXT_INGEST] settings as the batch method; a .txt upload
is decoded the same way (any common encoding, cp1250 first).
The batch pipeline and the API service share the same text_util_langID categorization engine and config.txt
settings — including the default Qwen2.5-0.5B 🤖 perplexity model — to ensure zero drift between local processing
and web uploads.
For deployment instructions, endpoint specifications (/process, /info), and frontend integration details,
please see the dedicated Service Documentation.
This project incorporates a unified provenance and paradata 🗒️ logging system to seamlessly track the execution details of every pipeline stage. The logger automatically captures run-time metadata and saves it in a structured JSON 📄 format.
What gets logged?
- Provenance 🏛️: Captures the tool name, a tool version 🏷️ tag, the repository/runner reference, the running
container image (when set), the Python 🐍 version, and assigns a unique
run_idto each execution. The repository reference is resolved dynamically — environment overrides (ATRIUM_RUNNER_REPO,ATRIUM_RUNNER_REF,ATRIUM_RUNNER_IMAGE) take precedence over the static fallback in para_config.txt 📎 — so the log points at the image actually executing rather than a fixed fork. - Output license ⚖️: Computes the effective output license 📜 of the run from the licensed components it actually
exercised, and records it as
license/license_urlplus a detailedlicense_detailblock (per-component licenses, which component(s)determined_bythe result,is_non_commercial/is_share_alikeflags, and any unknown licenses). See Output licensing below. - Configuration ⚙️: Stores run-time configuration ⚙️, including script names, input/output paths, and specific model choices.
- Timing ⏱️: Records precise UTC start times, end times, and the total duration of the run in seconds.
- Statistics 📊: Tracks the total number of input files, successfully processed documents, and computes performance throughput (e.g., output files generated per minute).
- Error Tracking 🐛: Maintains a
skipped_files_detaillist that logs the exact filename and specific error reason if a file fails to process.
Log Location
By default, JSON 📄 logs are written to the paradata 📁 directory following the naming convention
<YYMMDD-HHmmss>_<program>.json. Paradata is intended to live alongside the outputs 📤 (not committed to the
repository); the paradata 🗒️ JSON files themselves are distributed under the CC BY-NC 4.0 license.
Important
The license of the files a run produces is not fixed — it is computed per run as the most restrictive license among the components (models, data, APIs) that the run actually used. The mechanism is data-driven via para_config.txt 📎 (component → license) and para_licenses.py 📎 (restrictiveness ranking + share-alike / non-commercial rules), so the licensing owner can adjust it without touching the logger.
Each repository ships a para_config.txt 📎 listing its components. Components flagged always count
toward every run (the worst-case baseline); components flagged conditional are only counted when the script that uses
them records it. For this repository the components and their effect on the effective output license 📜 are:
| Component | License | Counted | Used by |
|---|---|---|---|
| alto-tools 🔧 ^1 — vendored, see below | Apache-2.0 | always | page split, statistics, alto-tools text extraction |
| FastText 🌐 ^2 | CC BY-NC 4.0 | always | language identification (classify_TEXT.py) |
| Qwen2.5-0.5B 🤖 ^6 | Apache-2.0 | conditional | perplexity 📉 scoring (default, classify_TEXT.py) |
| distilgpt2 🤖 | Apache-2.0 | conditional | perplexity 📉 scoring (English-only alternative) |
| LayoutLMv3 📐 ^9 | CC BY-NC-SA 4.0 | conditional | LayoutReader text extraction (extract_LytRdr_ALTO_2_TXT.py) |
| GLM-4v-9b 🤖 ^10 | glm-4 | conditional | generative OCR 🔍 extraction (extract_LLM_ALTO_2_TXT.py) |
| pypdfium2 📄 (PDFium) — PDF text layers (#31) | Apache-2.0 2 | conditional | text-lines PDF reading (text_split.py) |
| charset-normalizer 🔤 — encoding detection (#31) | MIT | conditional | text-lines non-UTF-8 plain text (text_split.py) |
Because the always-on FastText 🌐 weights are CC BY-NC 4.0, the baseline effective output license for this repository is CC BY-NC 4.0 (non-commercial). Runs that additionally use the LayoutReader 📐 method escalate to CC BY-NC-SA 4.0 (non-commercial and share-alike), the most restrictive option here. A run that exercised only permissive components would resolve to Apache-2.0.
Note
The restrictiveness ordering encoded in para_licenses.py 📎 is a mechanical engineering approximation, not legal advice; unrecognised licenses are treated conservatively as maximally restrictive so a missing entry can never silently relax the recorded output license.
The per-document record this tool reads and writes — <doc_id>.document.json, the pair of the
paradata above — follows atrium_document.schema.json 📎, schema
version 1.0. That version is frozen as the hub tag
doc-schema-v1
(ufal/atrium-project@544298b), and it is the baseline this tool implements:
- tests/test_schema_freeze.py 📎 checks this repository's schema against the frozen copy beside it, atrium_document.schema.doc-schema-v1.json 📎: the copy is exactly the tagged file, nothing declared at the freeze has been removed or renamed, and every change made since is registered;
- all three files, like
atrium_document.py, are vendored from the hub and kept byte-identical to itsv1tag by thepara-driftCI check, so they are never edited in this repository. The hub additionally checks that every record shape the tools write validates under both the frozen and the current schema.
What may change after the freeze, and what a new major version takes, is in the hub's
Freeze & conformance.
CITATION.cff 📎 carries the same reference under references.
Some third-party code is copied into this repository rather than installed as a dependency. Vendored code keeps its original licence; the repository's own LICENSE 📎 (MIT) does not apply to it. Licence texts live in LICENSES/ 📁.
| Upstream | cneud/alto-tools 3 by Clemens Neudecker |
| Commit | 1f4f01e5f6ac3562740e39948442973b9cb94be4 (__version__ = "0.1.0") |
| Source file | src/alto_tools/alto_tools.py |
| Licence | Apache-2.0 — full text in LICENSES/alto-tools-Apache-2.0.txt 📎 |
| Local copy | alto_tools.py 📎 |
| Used by | alto_stats_create.py 📎 (Step 2, statistics) and extract_ALTO_2_TXT.py 📎 (Step 3, alto-tools method) |
| Parity tests | tests/test_alto_tools.py 📎 |
Only the used functionality was copied — the code reachable from the two command-line flags this pipeline ever called:
| Vendored | Upstream flag | Used by |
|---|---|---|
alto_parse() |
(shared) | both |
alto_text() / text_from_file() |
alto-tools -t |
extract_ALTO_2_TXT.py |
alto_statistics() / statistics_from_file() |
alto-tools -s |
alto_stats_create.py |
Left behind, because nothing here calls it: -c mean word confidence, -i illustration boxes, -g graphic
boxes, the --dehyphenate / --detect-hyphens branches (this repo does its own de-hyphenation in
extract_ALTO_2_TXT._dehyphenate), the argparse command line, stdin input, directory walking and the
--xml-encoding sniffing path. Every deviation from the original is marked with a VENDORED: comment in the
file, as Apache-2.0 §4(b) requires for a derived work.
Why 📌 (issue #50) — alto-tools has no PyPI
release carrying the -s statistics flag, so it had to be declared as
alto-tools @ git+https://github.com/cneud/alto-tools.git@<commit>. That made the published Docker image
resolve a GitHub URL in order to build itself, and the end-to-end lane then re-installed a moving master
inside the already-released image, at run time — so what CI exercised was not the artefact being shipped.
Copying the two used code paths retires the git+ requirement, the shutil.which("alto-tools") guard, the
PATH export and the run-time pip install, all at once, and replaces two subprocess calls per page with two
function calls.
Important
The vendored code produces byte-identical output to the upstream CLI it replaced: the same page text
(block/line separators, spacing, <HYP> handling, reading order) and the same element counts.
tests/test_alto_tools.py 📎 pins this against output captured from the upstream
command line at the commit above, and re-derives it independently for every ALTO file in
data_samples 📁.
Updating it 🔄 — re-copy the two code paths from a newer upstream commit, update UPSTREAM_COMMIT in
alto_tools.py 📎 and the table above, then run pytest tests/test_alto_tools.py. Golden values
in that file are upstream CLI output and must only change when upstream's behaviour genuinely does.
For support write to: lutsai.k@gmail.com — responsible for this GitHub repository ^8 🔗
- Developed by UFAL ^7 👥
- Funded by ATRIUM ^4 💰
- Shared by ATRIUM ^4 & UFAL ^7 🔗
- Vendored code:
- alto-tools 🔧 ^1 by Clemens Neudecker, Apache-2.0 — ALTO text extraction and element statistics (alto_tools.py 📎)
- Models used:
©️ 2026 UFAL & ATRIUM
Footnotes
-
pypdfium2 is dual-licensed Apache-2.0 or BSD-3-Clause;
para_licensesdoes not parse SPDXOR, so it is recorded as Apache-2.0 (either is permissive). The text-lines readers use no other licensed component: DOCX, XLSX, PPTX, ODF, EPUB, HTML and XML are read with the standard library and lxml, so a text-lines run stays at the CC BY-NC 4.0 baseline. ↩