Skip to content

Repository files navigation

Docker Dataset

CI

Pre-populated sample databases as Docker images — ready-to-run PostgreSQL, MySQL, CockroachDB, SQLite, DuckDB, ClickHouse, Apache Druid, and Apache Pinot containers loaded with real, valid sample data (Chinook, Northwind, Sakila/Pagila, World, AdventureWorks, Stack Exchange, and more). Ever needed a database already populated with valid data — to practice SQL, run tests, demo an app, or benchmark — without hand-crafting rows or hunting for a usable dump? Every image ships exactly one dataset, so you just docker run and connect.

Quick start

Run the PostgreSQL world image, wait for it to initialize, run a real query, and clean up:

docker run -d --name pg-ds-world aa8y/postgres-dataset:world
# The first start loads the dataset; give it a few seconds. Watch for the
# second "database system is ready to accept connections":
docker logs -f pg-ds-world
docker exec -it pg-ds-world psql -d world -c 'SELECT name, population FROM city ORDER BY population DESC LIMIT 5'
docker rm -f pg-ds-world

That is the whole shape of it: pull a tag from the matrix below, docker run, connect. Each engine connects a little differently (server engines take a client over a port; SQLite and DuckDB open a file), and each has per-dataset notes — see its guide: PostgreSQL · MySQL · CockroachDB · SQLite · DuckDB · ClickHouse · Apache Druid · Apache Pinot.

Dataset support matrix

Each cell is the image tag to pull for that dataset on that engine; — means it isn't shipped there (yet). The dataset name links to its upstream source when every engine pulls from the same one; where engines use different upstreams, the source link is on the individual tag instead. All images are published for linux/amd64 and linux/arm64.

Dataset PostgreSQL MySQL CockroachDB SQLite DuckDB ClickHouse Apache Druid Apache Pinot
AdventureWorks adventureworks — — — — — — —
Airlines airlines — — airlines airlines — — —
Chinook chinook chinook chinook chinook chinook chinook — —
Dell DVD Store dellstore dellstore dellstore dellstore dellstore dellstore dellstore dellstore
Employees employees employees employees employees employees — — —
French Towns frenchtowns frenchtowns frenchtowns frenchtowns frenchtowns frenchtowns frenchtowns frenchtowns
GeoNames geonames geonames geonames geonames geonames geonames geonames geonames
ISO 3166 iso3166 iso3166 iso3166 iso3166 iso3166 iso3166 iso3166 iso3166
MoMA moma moma moma moma moma moma moma moma
Northwind northwind northwind northwind northwind northwind — — —
NYC Taxi Trip Records — — — — nyc-taxi nyc-taxi — —
OMDb omdb — — — — — — —
OpenFlights openflights openflights openflights openflights openflights openflights openflights openflights
PGExercises pgexercises pgexercises pgexercises pgexercises pgexercises pgexercises pgexercises pgexercises
Sakila / Pagila pagila sakila sakila sakila sakila — — —
SportsDB sportsdb sportsdb sportsdb sportsdb sportsdb — — —
Stack Exchange¹ stackexchange-<site> stackexchange-<site> stackexchange-<site> stackexchange-<site> stackexchange-<site> stackexchange-<site> — —
USDA usda usda usda usda usda usda usda usda
World world world world world world world world world

¹ <site> is one of beer, coffee, poker, woodworking, chess, cooking, outdoors, boardgames (e.g. stackexchange-chess).

Every engine also publishes a latest tag. It is an alias, not a dataset of its own: it opens the same dataset one of the named tags does — world on PostgreSQL, MySQL, Apache Druid, and Apache Pinot; chinook on CockroachDB, SQLite, and DuckDB; and nyc-taxi on ClickHouse.

Tag naming

A tag names the dataset an image carries. How that dataset is addressed once the container is up depends on the engine:

  • PostgreSQL, MySQL, CockroachDB, SQLite, DuckDB, ClickHouse put the dataset in a named database (or, for SQLite/DuckDB, a named file). The name is the tag minus any stackexchange- prefix (e.g. stackexchange-beer → beer). Two things are not databases named after the tag: the latest alias opens its target dataset's database (e.g. world, not latest), and on ClickHouse the nyc-taxi tag's database is nyc_taxi, because the base image interpolates the name into SQL unquoted and a hyphen does not survive that.
  • Apache Druid and Apache Pinot have no per-dataset database to name — Druid exposes the dataset's tables as datasources in its single druid schema, and Pinot as a flat namespace of tables — so you query the tables directly. Each engine's guide shows how.

Documentation

Docker Hub repositories: aa8y/postgres-dataset · aa8y/mysql-dataset · aa8y/cockroach-dataset · aa8y/sqlite-dataset · aa8y/duckdb-dataset · aa8y/clickhouse-dataset · aa8y/druid-dataset · aa8y/pinot-dataset

Dataset licenses

This repository's own software and packaging are MIT licensed. Each bundled dataset keeps its upstream license — see docs/ATTRIBUTION.md for the per-dataset sources, licenses, and required attributions (notably the Stack Exchange dumps, which are CC BY-SA 4.0).

Future Work

The remaining matrix gaps, and what each one is waiting on:

  • The last matrix gaps. adventureworks and omdb stay PostgreSQL-only by design (their upstreams lean on PostgreSQL-specific machinery — see Datasets not ported to MySQL). airlines on MySQL and CockroachDB is a volume problem, not a dialect one: 10.7M rows replayed as init-time INSERTs would blow the smoke test's readiness budget, so it needs a bulk-load path first (MariaDB's LOAD DATA INFILE, CockroachDB's IMPORT INTO as the employees tag already does) and a ~500 MB data payload in the image.
  • More OLAP datasets. ClickHouse carries the pgFoundry family, chinook, nyc-taxi, the CSV datasets and the Stack Exchange sites; its remaining gaps need per-dataset work (northwind's bytea columns, sportsdb's 107 mostly-empty tables — see Datasets not ported to ClickHouse). Druid needs an input format that tolerates embedded newlines (the Stack Exchange sites) or per-tag JVM sizing (nyc-taxi) — see Datasets not ported to Druid. Pinot's obvious next tag is nyc-taxi, which needs its Parquet ingestion job rather than the controller's synchronous CSV endpoint — see Datasets not ported to Pinot.
  • More Parquet-native datasets. DuckDB and ClickHouse both ship nyc-taxi from Parquet already (DuckDB reads it at build time; ClickHouse loads it at container start), and the open-data world publishes plenty more sources in that shape.
  • More free data sources across every engine.

About

Docker database images with pre-populated data for testing and/or practice.

Topics

Resources

Stars

39 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages