Skip to content

Repository files navigation

Wikisource Catalog

wikisource-catalog extracts a reusable book catalog from any Wikimedia Wikisource Special:IndexPages URL. It uses the MediaWiki action API, discovers localized ProofreadPage namespace names at runtime, follows API continuation, and writes Unicode-safe CSV and structured JSON.

Install

Python 3.10 or newer is required.

python -m venv .venv
. .venv/bin/activate
pip install -e .

For development tools, use pip install -e '.[dev]'.

Usage

wikisource-catalog \
  'https://ml.wikisource.org/wiki/പ്രത്യേകം:IndexPages' \
  --output-dir output --format csv,json

Inspect a wiki without downloading every Index revision:

wikisource-catalog 'https://fr.wikisource.org/wiki/Spécial:IndexPages' --inspect

Run a small extraction:

wikisource-catalog URL --limit 10 --verbose

Progress is printed after every completed Index page. CSV output is appended immediately, JSON is refreshed periodically, and a sibling *.checkpoint.jsonl file records each completed item. If a run is interrupted, continue it with the same arguments plus --resume:

wikisource-catalog URL --output-dir output --format csv,json --resume

The resume operation rebuilds the selected outputs from the checkpoint and skips Index pages already completed. Starting without --resume intentionally starts a fresh checkpoint and replaces prior result files with the same name.

Useful controls include --delay, --timeout, --max-retries, --no-cache, --clear-cache, --dry-run, --resume, and --debug. Cached API responses live in .cache/wikisource-catalog; output defaults to output/.

Optional Wikidata access

Wikidata access is disabled by default. Normal catalog extraction makes no requests to the Wikidata API.

Use --wikidata only when explicit Wikidata IDs found in Index metadata should be verified and supplemented with labels and descriptions:

wikisource-catalog URL --wikidata

Candidate search is a separate opt-in operation because most runs do not need it. Enable it explicitly with:

wikisource-catalog URL --wikidata-search

--wikidata verifies explicit item IDs and retrieves labels/descriptions. --wikidata-search records title-search candidates separately and never claims that a candidate is an established match. The options may be combined when both behaviors are required.

Data and behavior

The extractor supports ProofreadPage JSON content and older template wikitext. It records both the <pagelist> scan count and the number of corresponding pages found in the Page namespace. Failures enriching one book produce a partial record instead of aborting the complete catalog.

See architecture, API behavior, and the data model. Small illustrative outputs are in examples/.

Development

pytest
ruff check .
mypy cli.py wikisource_catalog.py api exporters models parsing services storage utils

Tests mock API responses and do not require network access.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages