Pre-populated sample databases as Docker images — ready-to-run PostgreSQL, MySQL, CockroachDB, SQLite, DuckDB, ClickHouse, Apache Druid, and Apache Pinot containers loaded with real, valid sample data (Chinook, Northwind, Sakila/Pagila, World, AdventureWorks, Stack Exchange, and more). Ever needed a database already populated with valid data — to practice SQL, run tests, demo an app, or benchmark — without hand-crafting rows or hunting for a usable dump? Every image ships exactly one dataset, so you just docker run and connect.
Run the PostgreSQL world image, wait for it to initialize, run a real query, and clean up:
docker run -d --name pg-ds-world aa8y/postgres-dataset:world
# The first start loads the dataset; give it a few seconds. Watch for the
# second "database system is ready to accept connections":
docker logs -f pg-ds-world
docker exec -it pg-ds-world psql -d world -c 'SELECT name, population FROM city ORDER BY population DESC LIMIT 5'
docker rm -f pg-ds-world
That is the whole shape of it: pull a tag from the matrix below, docker run, connect. Each engine connects a little differently (server engines take a client over a port; SQLite and DuckDB open a file), and each has per-dataset notes — see its guide: PostgreSQL · MySQL · CockroachDB · SQLite · DuckDB · ClickHouse · Apache Druid · Apache Pinot.
Each cell is the image tag to pull for that dataset on that engine; — means it isn't shipped there (yet). The dataset name links to its upstream source when every engine pulls from the same one; where engines use different upstreams, the source link is on the individual tag instead. All images are published for linux/amd64 and linux/arm64.
| Dataset | PostgreSQL | MySQL | CockroachDB | SQLite | DuckDB | ClickHouse | Apache Druid | Apache Pinot |
|---|---|---|---|---|---|---|---|---|
| AdventureWorks | adventureworks |
— | — | — | — | — | — | — |
| Airlines | airlines |
— | — | airlines |
airlines |
— | — | — |
| Chinook | chinook |
chinook |
chinook |
chinook |
chinook |
chinook |
— | — |
| Dell DVD Store | dellstore |
dellstore |
dellstore |
dellstore |
dellstore |
dellstore |
dellstore |
dellstore |
| Employees | employees |
employees |
employees |
employees |
employees |
— | — | — |
| French Towns | frenchtowns |
frenchtowns |
frenchtowns |
frenchtowns |
frenchtowns |
frenchtowns |
frenchtowns |
frenchtowns |
| GeoNames | geonames |
geonames |
geonames |
geonames |
geonames |
geonames |
geonames |
geonames |
| ISO 3166 | iso3166 |
iso3166 |
iso3166 |
iso3166 |
iso3166 |
iso3166 |
iso3166 |
iso3166 |
| MoMA | moma |
moma |
moma |
moma |
moma |
moma |
moma |
moma |
| Northwind | northwind |
northwind |
northwind |
northwind |
northwind |
— | — | — |
| NYC Taxi Trip Records | — | — | — | — | nyc-taxi |
nyc-taxi |
— | — |
| OMDb | omdb |
— | — | — | — | — | — | — |
| OpenFlights | openflights |
openflights |
openflights |
openflights |
openflights |
openflights |
openflights |
openflights |
| PGExercises | pgexercises |
pgexercises |
pgexercises |
pgexercises |
pgexercises |
pgexercises |
pgexercises |
pgexercises |
| Sakila / Pagila | pagila |
sakila |
sakila |
sakila |
sakila |
— | — | — |
| SportsDB | sportsdb |
sportsdb |
sportsdb |
sportsdb |
sportsdb |
— | — | — |
| Stack Exchange¹ | stackexchange-<site> |
stackexchange-<site> |
stackexchange-<site> |
stackexchange-<site> |
stackexchange-<site> |
stackexchange-<site> |
— | — |
| USDA | usda |
usda |
usda |
usda |
usda |
usda |
usda |
usda |
| World | world |
world |
world |
world |
world |
world |
world |
world |
¹ <site> is one of beer, coffee, poker, woodworking, chess, cooking, outdoors, boardgames (e.g. stackexchange-chess).
Every engine also publishes a latest tag. It is an alias, not a dataset of its own: it opens the same dataset one of the named tags does — world on PostgreSQL, MySQL, Apache Druid, and Apache Pinot; chinook on CockroachDB, SQLite, and DuckDB; and nyc-taxi on ClickHouse.
A tag names the dataset an image carries. How that dataset is addressed once the container is up depends on the engine:
- PostgreSQL, MySQL, CockroachDB, SQLite, DuckDB, ClickHouse put the dataset in a named database (or, for SQLite/DuckDB, a named file). The name is the tag minus any
stackexchange-prefix (e.g.stackexchange-beer→beer). Two things are not databases named after the tag: thelatestalias opens its target dataset's database (e.g.world, notlatest), and on ClickHouse thenyc-taxitag's database isnyc_taxi, because the base image interpolates the name into SQL unquoted and a hyphen does not survive that. - Apache Druid and Apache Pinot have no per-dataset database to name — Druid exposes the dataset's tables as datasources in its single
druidschema, and Pinot as a flat namespace of tables — so you query the tables directly. Each engine's guide shows how.
- Engine guides: PostgreSQL · MySQL · CockroachDB · SQLite · DuckDB · ClickHouse · Apache Druid · Apache Pinot — each has the client invocation, host-connection details, and per-dataset notes for its engine.
- Building images —
dave, custom images, and build caching - Testing — structure tests and integration smoke tests
- Dataset attribution and licenses
Docker Hub repositories: aa8y/postgres-dataset · aa8y/mysql-dataset · aa8y/cockroach-dataset · aa8y/sqlite-dataset · aa8y/duckdb-dataset · aa8y/clickhouse-dataset · aa8y/druid-dataset · aa8y/pinot-dataset
This repository's own software and packaging are MIT licensed. Each bundled dataset keeps its upstream license — see docs/ATTRIBUTION.md for the per-dataset sources, licenses, and required attributions (notably the Stack Exchange dumps, which are CC BY-SA 4.0).
The remaining matrix gaps, and what each one is waiting on:
- The last matrix gaps.
adventureworksandomdbstay PostgreSQL-only by design (their upstreams lean on PostgreSQL-specific machinery — see Datasets not ported to MySQL).airlineson MySQL and CockroachDB is a volume problem, not a dialect one: 10.7M rows replayed as init-timeINSERTs would blow the smoke test's readiness budget, so it needs a bulk-load path first (MariaDB'sLOAD DATA INFILE, CockroachDB'sIMPORT INTOas theemployeestag already does) and a ~500 MB data payload in the image. - More OLAP datasets. ClickHouse carries the pgFoundry family,
chinook,nyc-taxi, the CSV datasets and the Stack Exchange sites; its remaining gaps need per-dataset work (northwind'sbyteacolumns,sportsdb's 107 mostly-empty tables — see Datasets not ported to ClickHouse). Druid needs an input format that tolerates embedded newlines (the Stack Exchange sites) or per-tag JVM sizing (nyc-taxi) — see Datasets not ported to Druid. Pinot's obvious next tag isnyc-taxi, which needs its Parquet ingestion job rather than the controller's synchronous CSV endpoint — see Datasets not ported to Pinot. - More Parquet-native datasets. DuckDB and ClickHouse both ship
nyc-taxifrom Parquet already (DuckDB reads it at build time; ClickHouse loads it at container start), and the open-data world publishes plenty more sources in that shape. - More free data sources across every engine.