A small, dependency-free Linux process supervisor and event-driven job scheduler with an HTTP control plane and Prometheus metrics.
watchdog forks two workers — a scheduler and an HTTP server — and supervises them. The scheduler runs a configurable job on a fixed interval (or on demand), and the HTTP server exposes health, status, and control endpoints. Everything logs as structured JSON and exposes Prometheus metrics, making it a drop-in job supervisor for containers, sidecars, and lightweight daemons.
- Process supervision —
mainforks and monitors the scheduler and HTTP workers, aggregates their logs, and shuts down gracefully onSIGINT/SIGTERM. - Scheduled execution — runs a job every
SCHEDULER_EVERY_MINUTESminutes, with optionalKILL_ON_TIMEOUTto kill a still-running job when the next interval fires. - On-demand control —
POST /triggerandPOST /stopspawn or terminate a job via a Unix-domain-socket IPC channel. - Health & status —
/livez,/readyz,/startupz, and/status. - Prometheus metrics —
/metricsin text exposition format. - Optional auth — constant-time SHA-256 verification of a Bearer token or
X-API-Keyon mutating endpoints. - Structured JSON logging — one JSON object per line on stdout.
flowchart LR
M[main supervisor] -->|fork| S[scheduler worker]
M -->|fork| H[http_server worker]
S <-->|Unix socket| H
S -->|exec| J[job process]
H -->|TCP| C[HTTP clients]
S -.->|log pipe| M
H -.->|log pipe| M
The scheduler and HTTP server communicate over a Unix domain socket (SOCKET_PATH); worker/journal output is captured over pipes and forwarded to the supervisor's stdout.
Requires Linux and GCC. The code uses epoll, timerfd, signalfd, Unix sockets, and close_range, so _GNU_SOURCE is required (already defined in src/main.c).
make # build ./watchdogNo third-party libraries are needed. The test suite additionally requires Python 3.
Ready-made manifests live under deploy/:
deploy/kubernetes/— genericConfigMap,Secret(example),Service,Deployment,StatefulSet, and PrometheusServiceMonitor, with probes on/livez,/readyz, and/startupz.
The directory has its own README with prerequisites and step-by-step instructions.
Configuration is via environment variables:
| Variable | Required | Default | Description |
|---|---|---|---|
SCHEDULER_EVERY_MINUTES |
yes | — | Job interval in minutes (positive integer). |
SOCKET_PATH |
yes | — | Path of the Unix domain socket used for scheduler ↔ HTTP IPC. |
HTTP_PORT |
no | 8080 |
TCP port for the HTTP control plane. |
KILL_ON_TIMEOUT |
no | unset | If 1, kill a still-running job when the next interval fires (default: skip). |
JOB_EXECUTABLE |
no | /bin/true |
Path of the program to run. |
JOB_ARGS |
no | empty | Space/tab/newline-separated arguments passed to the job. |
STOP_GRACE_PERIOD_SECONDS |
no | 5 |
Seconds a job gets to exit on SIGTERM before it is SIGKILLed (integer 0–300; 0 escalates immediately). Invalid values abort startup. |
HISTORY_MAX |
no | 20 |
Number of recent runs kept in the in-memory history buffer (capped at 64). |
AUTH_TOKEN_HASH_SHA256 |
no | unset | 64-char hex SHA-256 of the token; when set, mutating endpoints require auth. |
SCHEDULER_EVERY_MINUTES=5 \
SOCKET_PATH=/tmp/watchdog.sock \
HTTP_PORT=8080 \
JOB_EXECUTABLE=/usr/local/bin/backup \
JOB_ARGS="--daily" \
./watchdog| Method | Endpoint | Description |
|---|---|---|
| GET | /livez /healthz /startupz |
Liveness, health, and startup probes → 200. |
| GET | /readyz |
Readiness; 200 when the scheduler is reachable, else 503. |
| GET | /status |
JSON status (scheduler state, active pid, next run, interval). |
| GET | /history |
JSON array of recent runs (id, start/end, duration, exit code, trigger source). |
| GET | /metrics |
Prometheus metrics in text format. |
| POST | /trigger /run |
Spawn a job now → 200 triggered, 409 busy, 401 unauthorized. |
| POST | /stop /kill |
Stop the running job → 200, 401 unauthorized. |
| — | anything else | 404. |
GET /history returns the last HISTORY_MAX completed runs as a JSON array in
chronological (oldest-first) order:
[
{
"id": 7,
"start": "2026-09-25T00:47:10Z",
"end": "2026-09-25T00:47:10Z",
"duration_seconds": 0.412,
"exit_code": 0,
"trigger_source": "manual"
}
]trigger_source is "scheduled" for an interval tick or "manual" for
POST /trigger. A run terminated by /stop, a timeout, or a signal reports a
negative exit_code (e.g. -15 for SIGTERM).
History is held in memory only and is intentionally not persisted: it is
empty after a restart. This keeps the daemon dependency-free; tools around it
(Prometheus via /metrics, a log shipper, or a curl on an interval) are the
right place to retain history if you need it. JOB_ARGS is deliberately
not recorded, so job arguments cannot leak through this endpoint. When
AUTH_TOKEN_HASH_SHA256 is set, /history requires auth like the other
sensitive endpoints.
When AUTH_TOKEN_HASH_SHA256 is set, POST /trigger and POST /stop require either:
Authorization: Bearer <token>or:
X-API-Key: <token>where the SHA-256 hash of <token> must match the configured value.
When a running job must stop — POST /stop, a KILL_ON_TIMEOUT timeout, or
supervisor shutdown — the scheduler performs a SIGTERM → grace period →
SIGKILL handshake:
SIGTERMis sent immediately and the grace timer is armed forSTOP_GRACE_PERIOD_SECONDS.- If the job exits within the grace period (the normal case for cooperative jobs), it is reaped asynchronously and no signal escalation occurs.
- If it is still running when the grace period expires, it is
SIGKILLed.
Escalation is driven by a dedicated non-blocking timerfd in the scheduler's
epoll set — not a blocking sleep. The event loop therefore keeps serving
IPC, /status, /metrics, and reaping children while a job is stopping, so a
stop request never stalls the control plane: POST /stop returns 200 as soon
as SIGTERM is sent, rather than waiting for the job to die. A
STOP_GRACE_PERIOD_SECONDS of 0 skips the grace period and escalates to
SIGKILL at once.
On SIGINT/SIGTERM to the daemon, main forwards SIGTERM to both workers.
The scheduler then stops its interval loop, terminates any in-flight job with
the same handshake, and waits for it inline (bounded by
STOP_GRACE_PERIOD_SECONDS) before cleaning up the socket and exiting.
Exposed on /metrics:
| Metric | Type |
|---|---|
watchdog_uptime_seconds |
gauge |
watchdog_job_runs_total |
counter |
watchdog_job_success_total |
counter |
watchdog_job_failures_total |
counter |
watchdog_job_timeouts_total |
counter |
watchdog_job_status |
gauge |
watchdog_next_run_in_seconds |
gauge |
watchdog_job_last_exit_code |
gauge |
watchdog_job_last_duration_seconds |
gauge |
watchdog_process_resident_memory_bytes |
gauge |
watchdog_process_cpu_user_seconds_total |
counter |
watchdog_process_cpu_system_seconds_total |
counter |
watchdog_job_last_exit_code is -1 until the first run completes and is
negative for a signal-terminated run (e.g. -15 for SIGTERM). The
watchdog_process_* metrics describe the scheduler worker, read from
/proc/self/stat[m].
make test # build + run the unittest suite
make check # static analysis (requires cppcheck)
make asan # build with AddressSanitizer + UBSan
make test-asan # build with sanitizers + run testsThe normal build can be checked for leaks and invalid memory access with
Valgrind. The daemon reads required configuration from the environment, so pass
it on the command line — a bare ./watchdog would exit immediately with code 1
because SCHEDULER_EVERY_MINUTES and SOCKET_PATH are missing:
SCHEDULER_EVERY_MINUTES=60 \
SOCKET_PATH=/tmp/watchdog-vg.sock \
HTTP_PORT=8080 \
JOB_EXECUTABLE=/bin/true \
valgrind --leak-check=full --errors-for-leak-kinds=definite --error-exitcode=1 ./watchdogThen exercise the running daemon from another shell (this is what makes the run meaningful — an idle process that never handles a request exercises little):
for i in 1 2 3; do curl -s -o /dev/null -X POST localhost:8080/trigger; done
curl -s localhost:8080/history
curl -s localhost:8080/metrics
curl -s localhost:8080/statusStop it with Ctrl-C (or SIGTERM) so Valgrind can run the leak check on
exit. A clean run ends with:
ERROR SUMMARY: 0 errors from 0 contexts (suppressed: 0 from 0)
Useful flags:
--leak-check=full— report each leak site and its backtrace.--errors-for-leak-kinds=definite --error-exitcode=1— exit non-zero only on definite leaks or errors, which is convenient in a script or CI step.--track-fds=yes— report file descriptors still open at exit (useful when touching the socket/pipe plumbing).--log-file=vg.log— write the report to a file instead of mixing it with the daemon's own stdout/JSON logs.
The daemon forks a scheduler and an HTTP worker, so Valgrind traces all three
processes; each prints its own ERROR SUMMARY line.
Planned work — retry with backoff, rate limiting, cron-expression scheduling, and Helm packaging — is tracked in TODO.md.
- The HTTP worker is a blocking, single-connection accept loop. A client
that never sends a full header block is bounded by a 5-second
SO_RCVTIMEO, but that is a per-read timeout: a client that trickles bytes can still hold the loop. A per-connection deadline on anepoll/timerfdreactor is the planned fix (see TODO.md). - The scheduler already runs an
epoll+timerfd+signalfdevent loop;main(the supervisor) usespollover two log pipes plussignalfd.