Demo results (mock vLLM): TTFT p50 84 ms / p99 244 ms · TBT p50 17 ms · E2E p50 384 ms / p99 925 ms · KV-cache ~98% · 32/32 tests passing.
Your APM says CPU is fine, memory is fine, and every request returns 200 — while users wait seconds for the first token.
Generic monitoring cannot see how LLM serving fails. InferSight can.
InferSight is a lightweight sidecar for vLLM (and any OpenAI-compatible server) that measures the signals that matter for inference: TTFT, TBT, E2E latency, KV-cache pressure, scheduler queue depth, and exact token throughput.
| Gap in generic APM | What InferSight captures |
|---|---|
| Request latency only | TTFT — time until the first token (what users feel) |
| No streaming insight | TBT — inter-token gaps (decode / smoothness health) |
| No GPU KV visibility | KV-cache usage — leading indicator of preemption & OOM |
| Opaque backlogs | Queue depth — running / waiting / swapped |
| Approximate token counts | Exact usage via injected stream_options.include_usage |
flowchart LR
C[Clients / SDKs] -->|OpenAI API| S[InferSight Sidecar :8020]
S -->|passthrough stream| V[vLLM / OpenAI server :8000]
V -->|engine /metrics| S
S -->|Prometheus scrape| P[(Prometheus)]
P --> G[Grafana Dashboard]
S -.->|optional batched ship| H[Hosted Tier :9000]
H --> A[Slack / PagerDuty Alerts]
Hot-path design: streamed bytes pass through untouched. Timing uses two perf_counter() reads per network chunk — out-of-band of the client response. Overhead stays in the microsecond range.
sequenceDiagram
participant Client
participant Sidecar as InferSight
participant Engine as vLLM
Client->>Sidecar: POST /v1/chat/completions (stream)
Sidecar->>Engine: forward (+ inject include_usage)
loop SSE chunks
Engine-->>Sidecar: token chunk
Note over Sidecar: timestamp out-of-band
Sidecar-->>Client: same bytes, unmodified
end
Sidecar->>Sidecar: observe TTFT / TBT / E2E / tokens
Sidecar-->>Client: final usage + [DONE]
Full design notes: docs/architecture.md · Project report: docs/REPORT.md
pip install -e .
infersight run --upstream http://localhost:8000 --port 8020Point clients at :8020. Scrape http://localhost:8020/metrics. Import dashboards/infersight-vllm.json into Grafana.
infersight discover # probe for OpenAI-compatible / vLLM serversgit clone https://github.com/ArchanaChetan07/InferSight && cd InferSight
docker compose up --build| Service | Default URL |
|---|---|
| Mock vLLM | http://localhost:8000 |
| InferSight sidecar | http://localhost:8020 |
| Prometheus | http://localhost:9090 |
| Grafana (anon Admin) | http://localhost:3000 |
| Hosted tier (optional) | http://localhost:9000 |
If host ports are taken:
MOCK_PORT=18080 SIDECAR_PORT=18020 PROM_PORT=19091 GRAFANA_PORT=13001 HOSTED_PORT=19000 \
docker compose up --build| Metric | Type | Purpose |
|---|---|---|
infersight_ttft_seconds |
Histogram | Time to first token |
infersight_tbt_seconds |
Histogram | Time between tokens |
infersight_e2e_latency_seconds |
Histogram | End-to-end request latency |
infersight_prompt_tokens_total |
Counter | Prompt tokens |
infersight_completion_tokens_total |
Counter | Completion tokens |
infersight_requests_total |
Counter | Requests by status |
infersight_requests_in_flight |
Gauge | Concurrent requests |
infersight_kv_cache_usage_ratio |
Gauge | Engine KV-cache utilization |
infersight_queue_depth |
Gauge | Scheduler running / waiting / swapped |
Grafana panels (pre-built): TTFT · TBT · E2E · Throughput · KV-cache · Queue · In-flight & error rate.
Measured against the included mock vLLM + InferSight sidecar + Prometheus stack (model: meta-llama/Llama-3.1-8B-Instruct).
| Signal | P50 | P99 |
|---|---|---|
| TTFT | 84 ms | 244 ms |
| TBT | 17 ms | 39 ms |
| E2E latency | 384 ms | 925 ms |
| Volume / health | Value |
|---|---|
| Requests observed | 18 (17 × HTTP 200) |
| Prompt tokens | 4,291 |
| Completion tokens | 300 |
| KV-cache usage | ~98% |
| Queue (running / waiting) | 7 / 3 |
Values reflect a local demo workload. Re-run
docker compose upand send traffic through the sidecar to regenerate live series in Grafana.
xychart-beta
title "Demo latency percentiles (ms)"
x-axis [TTFT, TBT, E2E]
y-axis "Milliseconds" 0 --> 1000
bar [84, 17, 384]
bar [244, 39, 925]
Ship metrics without running Prometheus/Grafana yourself:
infersight run --upstream http://localhost:8000 \
--hosted-api-key isk_your_keyIncludes:
- Zero-ops dashboard (TTFT / TBT / E2E percentiles, throughput, errors)
- LLM-aware alerts to Slack / PagerDuty (P99 regressions, KV pressure, error spikes)
- Multi-model / multi-cluster comparison tables
Self-host the ingest API from this repo:
uvicorn hosted.ingest:app --port 9000
curl -X POST localhost:9000/v1/admin/tenants \
-H "Authorization: Bearer $INFERSIGHT_ADMIN_TOKEN" \
-H "Content-Type: application/json" \
-d '{"name":"my-team"}'Precedence: CLI flags → INFERSIGHT_* env → --config file → defaults
INFERSIGHT_UPSTREAM_URL=http://vllm:8000
INFERSIGHT_LISTEN_PORT=8020
INFERSIGHT_HOSTED__API_KEY=isk_... # nested keys use __| Engine | Request timing | KV cache / queue |
|---|---|---|
| vLLM | Yes | Yes |
| Any OpenAI-compatible server | Yes | — |
| TGI / SGLang gauges | Planned | Planned |
InferSight/
├── infersight/ # Sidecar proxy, metrics, discovery, CLI
├── hosted/ # Optional multi-tenant ingest + alerts + UI
├── dashboards/ # Import-ready Grafana JSON
├── deploy/ # Prometheus + Grafana provisioning
├── examples/mock_vllm.py
├── tests/ # 32 unit / hosted / e2e / regression tests
├── docs/ # Architecture, config, limitations, report, assets
└── docker-compose.yml # One-command demo stack
pip install -e ".[dev]"
pytest # 32 testsflowchart TB
subgraph tests [Test suite]
U[Unit — timer, metrics, config]
F[Forwarder — batching / retry]
H[Hosted — tenancy, alerts, API]
E[E2E — mock vLLM + live proxy]
R[Regressions — SSE buffer, cardinality]
end
U --> F --> H --> E --> R
| Document | Contents |
|---|---|
| docs/REPORT.md | Full project report — goals, design, results, roadmap |
| docs/architecture.md | Design decisions & data model |
| docs/configuration.md | Config reference |
| docs/limitations.md | Known limitations (v0.1) |
llm · vllm · observability · prometheus · grafana · inference · ttft · streaming · openai-compatible · fastapi · sidecar · kv-cache · sre · mlops
Apache-2.0. Built by Archana Suresh Patil.
Feedback and design partners welcome — open an issue on GitHub.



