wikisource-catalog extracts a reusable book catalog from any Wikimedia
Wikisource Special:IndexPages URL. It uses the MediaWiki action API, discovers
localized ProofreadPage namespace names at runtime, follows API continuation,
and writes Unicode-safe CSV and structured JSON.
Python 3.10 or newer is required.
python -m venv .venv
. .venv/bin/activate
pip install -e .For development tools, use pip install -e '.[dev]'.
wikisource-catalog \
'https://ml.wikisource.org/wiki/പ്രത്യേകം:IndexPages' \
--output-dir output --format csv,jsonInspect a wiki without downloading every Index revision:
wikisource-catalog 'https://fr.wikisource.org/wiki/Spécial:IndexPages' --inspectRun a small extraction:
wikisource-catalog URL --limit 10 --verboseProgress is printed after every completed Index page. CSV output is appended
immediately, JSON is refreshed periodically, and a sibling
*.checkpoint.jsonl file records each completed item. If a run is interrupted,
continue it with the same arguments plus --resume:
wikisource-catalog URL --output-dir output --format csv,json --resumeThe resume operation rebuilds the selected outputs from the checkpoint and
skips Index pages already completed. Starting without --resume intentionally
starts a fresh checkpoint and replaces prior result files with the same name.
Useful controls include --delay, --timeout, --max-retries, --no-cache,
--clear-cache, --dry-run, --resume, and --debug. Cached API responses
live in .cache/wikisource-catalog; output defaults to output/.
Wikidata access is disabled by default. Normal catalog extraction makes no requests to the Wikidata API.
Use --wikidata only when explicit Wikidata IDs found in Index metadata should
be verified and supplemented with labels and descriptions:
wikisource-catalog URL --wikidataCandidate search is a separate opt-in operation because most runs do not need it. Enable it explicitly with:
wikisource-catalog URL --wikidata-search--wikidata verifies explicit item IDs and retrieves labels/descriptions.
--wikidata-search records title-search candidates separately and never claims
that a candidate is an established match. The options may be combined when
both behaviors are required.
The extractor supports ProofreadPage JSON content and older template wikitext.
It records both the <pagelist> scan count and the number of corresponding
pages found in the Page namespace. Failures enriching one book produce a
partial record instead of aborting the complete catalog.
See architecture, API behavior, and the
data model. Small illustrative outputs are in examples/.
pytest
ruff check .
mypy cli.py wikisource_catalog.py api exporters models parsing services storage utilsTests mock API responses and do not require network access.