Skip to content

Latest commit

 

History

15 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

watchdog

A small, dependency-free Linux process supervisor and event-driven job scheduler with an HTTP control plane and Prometheus metrics.

watchdog forks two workers — a scheduler and an HTTP server — and supervises them. The scheduler runs a configurable job on a fixed interval (or on demand), and the HTTP server exposes health, status, and control endpoints. Everything logs as structured JSON and exposes Prometheus metrics, making it a drop-in job supervisor for containers, sidecars, and lightweight daemons.

Features

  • Process supervision — main forks and monitors the scheduler and HTTP workers, aggregates their logs, and shuts down gracefully on SIGINT/SIGTERM.
  • Scheduled execution — runs a job every SCHEDULER_EVERY_MINUTES minutes, with optional KILL_ON_TIMEOUT to kill a still-running job when the next interval fires.
  • On-demand control — POST /trigger and POST /stop spawn or terminate a job via a Unix-domain-socket IPC channel.
  • Health & status — /livez, /readyz, /startupz, and /status.
  • Prometheus metrics — /metrics in text exposition format.
  • Optional auth — constant-time SHA-256 verification of a Bearer token or X-API-Key on mutating endpoints.
  • Structured JSON logging — one JSON object per line on stdout.

Architecture

flowchart LR
    M[main supervisor] -->|fork| S[scheduler worker]
    M -->|fork| H[http_server worker]
    S <-->|Unix socket| H
    S -->|exec| J[job process]
    H -->|TCP| C[HTTP clients]
    S -.->|log pipe| M
    H -.->|log pipe| M
Loading

The scheduler and HTTP server communicate over a Unix domain socket (SOCKET_PATH); worker/journal output is captured over pipes and forwarded to the supervisor's stdout.

Building

Requires Linux and GCC. The code uses epoll, timerfd, signalfd, Unix sockets, and close_range, so _GNU_SOURCE is required (already defined in src/main.c).

make        # build ./watchdog

No third-party libraries are needed. The test suite additionally requires Python 3.

Deployment

Ready-made manifests live under deploy/:

  • deploy/kubernetes/ — generic ConfigMap, Secret (example), Service, Deployment, StatefulSet, and Prometheus ServiceMonitor, with probes on /livez, /readyz, and /startupz.

The directory has its own README with prerequisites and step-by-step instructions.

Configuration

Configuration is via environment variables:

Variable Required Default Description
SCHEDULER_EVERY_MINUTES yes — Job interval in minutes (positive integer).
SOCKET_PATH yes — Path of the Unix domain socket used for scheduler ↔ HTTP IPC.
HTTP_PORT no 8080 TCP port for the HTTP control plane.
KILL_ON_TIMEOUT no unset If 1, kill a still-running job when the next interval fires (default: skip).
JOB_EXECUTABLE no /bin/true Path of the program to run.
JOB_ARGS no empty Space/tab/newline-separated arguments passed to the job.
STOP_GRACE_PERIOD_SECONDS no 5 Seconds a job gets to exit on SIGTERM before it is SIGKILLed (integer 0–300; 0 escalates immediately). Invalid values abort startup.
HISTORY_MAX no 20 Number of recent runs kept in the in-memory history buffer (capped at 64).
AUTH_TOKEN_HASH_SHA256 no unset 64-char hex SHA-256 of the token; when set, mutating endpoints require auth.

Example

SCHEDULER_EVERY_MINUTES=5 \
SOCKET_PATH=/tmp/watchdog.sock \
HTTP_PORT=8080 \
JOB_EXECUTABLE=/usr/local/bin/backup \
JOB_ARGS="--daily" \
./watchdog

HTTP API

Method Endpoint Description
GET /livez /healthz /startupz Liveness, health, and startup probes → 200.
GET /readyz Readiness; 200 when the scheduler is reachable, else 503.
GET /status JSON status (scheduler state, active pid, next run, interval).
GET /history JSON array of recent runs (id, start/end, duration, exit code, trigger source).
GET /metrics Prometheus metrics in text format.
POST /trigger /run Spawn a job now → 200 triggered, 409 busy, 401 unauthorized.
POST /stop /kill Stop the running job → 200, 401 unauthorized.
— anything else 404.

Run history

GET /history returns the last HISTORY_MAX completed runs as a JSON array in chronological (oldest-first) order:

[
  {
    "id": 7,
    "start": "2026-09-25T00:47:10Z",
    "end": "2026-09-25T00:47:10Z",
    "duration_seconds": 0.412,
    "exit_code": 0,
    "trigger_source": "manual"
  }
]

trigger_source is "scheduled" for an interval tick or "manual" for POST /trigger. A run terminated by /stop, a timeout, or a signal reports a negative exit_code (e.g. -15 for SIGTERM).

History is held in memory only and is intentionally not persisted: it is empty after a restart. This keeps the daemon dependency-free; tools around it (Prometheus via /metrics, a log shipper, or a curl on an interval) are the right place to retain history if you need it. JOB_ARGS is deliberately not recorded, so job arguments cannot leak through this endpoint. When AUTH_TOKEN_HASH_SHA256 is set, /history requires auth like the other sensitive endpoints.

When AUTH_TOKEN_HASH_SHA256 is set, POST /trigger and POST /stop require either:

Authorization: Bearer <token>

or:

X-API-Key: <token>

where the SHA-256 hash of <token> must match the configured value.

Graceful termination

When a running job must stop — POST /stop, a KILL_ON_TIMEOUT timeout, or supervisor shutdown — the scheduler performs a SIGTERM → grace period → SIGKILL handshake:

  1. SIGTERM is sent immediately and the grace timer is armed for STOP_GRACE_PERIOD_SECONDS.
  2. If the job exits within the grace period (the normal case for cooperative jobs), it is reaped asynchronously and no signal escalation occurs.
  3. If it is still running when the grace period expires, it is SIGKILLed.

Escalation is driven by a dedicated non-blocking timerfd in the scheduler's epoll set — not a blocking sleep. The event loop therefore keeps serving IPC, /status, /metrics, and reaping children while a job is stopping, so a stop request never stalls the control plane: POST /stop returns 200 as soon as SIGTERM is sent, rather than waiting for the job to die. A STOP_GRACE_PERIOD_SECONDS of 0 skips the grace period and escalates to SIGKILL at once.

On SIGINT/SIGTERM to the daemon, main forwards SIGTERM to both workers. The scheduler then stops its interval loop, terminates any in-flight job with the same handshake, and waits for it inline (bounded by STOP_GRACE_PERIOD_SECONDS) before cleaning up the socket and exiting.

Metrics

Exposed on /metrics:

Metric Type
watchdog_uptime_seconds gauge
watchdog_job_runs_total counter
watchdog_job_success_total counter
watchdog_job_failures_total counter
watchdog_job_timeouts_total counter
watchdog_job_status gauge
watchdog_next_run_in_seconds gauge
watchdog_job_last_exit_code gauge
watchdog_job_last_duration_seconds gauge
watchdog_process_resident_memory_bytes gauge
watchdog_process_cpu_user_seconds_total counter
watchdog_process_cpu_system_seconds_total counter

watchdog_job_last_exit_code is -1 until the first run completes and is negative for a signal-terminated run (e.g. -15 for SIGTERM). The watchdog_process_* metrics describe the scheduler worker, read from /proc/self/stat[m].

Testing

make test        # build + run the unittest suite
make check       # static analysis (requires cppcheck)
make asan        # build with AddressSanitizer + UBSan
make test-asan   # build with sanitizers + run tests

Valgrind (dynamic analysis)

The normal build can be checked for leaks and invalid memory access with Valgrind. The daemon reads required configuration from the environment, so pass it on the command line — a bare ./watchdog would exit immediately with code 1 because SCHEDULER_EVERY_MINUTES and SOCKET_PATH are missing:

SCHEDULER_EVERY_MINUTES=60 \
SOCKET_PATH=/tmp/watchdog-vg.sock \
HTTP_PORT=8080 \
JOB_EXECUTABLE=/bin/true \
valgrind --leak-check=full --errors-for-leak-kinds=definite --error-exitcode=1 ./watchdog

Then exercise the running daemon from another shell (this is what makes the run meaningful — an idle process that never handles a request exercises little):

for i in 1 2 3; do curl -s -o /dev/null -X POST localhost:8080/trigger; done
curl -s localhost:8080/history
curl -s localhost:8080/metrics
curl -s localhost:8080/status

Stop it with Ctrl-C (or SIGTERM) so Valgrind can run the leak check on exit. A clean run ends with:

ERROR SUMMARY: 0 errors from 0 contexts (suppressed: 0 from 0)

Useful flags:

  • --leak-check=full — report each leak site and its backtrace.
  • --errors-for-leak-kinds=definite --error-exitcode=1 — exit non-zero only on definite leaks or errors, which is convenient in a script or CI step.
  • --track-fds=yes — report file descriptors still open at exit (useful when touching the socket/pipe plumbing).
  • --log-file=vg.log — write the report to a file instead of mixing it with the daemon's own stdout/JSON logs.

The daemon forks a scheduler and an HTTP worker, so Valgrind traces all three processes; each prints its own ERROR SUMMARY line.

Roadmap

Planned work — retry with backoff, rate limiting, cron-expression scheduling, and Helm packaging — is tracked in TODO.md.

Known limitations

  • The HTTP worker is a blocking, single-connection accept loop. A client that never sends a full header block is bounded by a 5-second SO_RCVTIMEO, but that is a per-read timeout: a client that trickles bytes can still hold the loop. A per-connection deadline on an epoll/timerfd reactor is the planned fix (see TODO.md).
  • The scheduler already runs an epoll + timerfd + signalfd event loop; main (the supervisor) uses poll over two log pipes plus signalfd.

About

Dependency-free Linux process supervisor and job scheduler in C — epoll/timerfd/signalfd, HTTP control plane, Prometheus metrics

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages