A systems-grade SDK that makes Solana RPC and transaction submission reliable — built on web3.js v2.0.
Public RPCs rate-limit, drop transactions, lag, and occasionally fall over. This SDK wraps web3.js v2.0 with a resilience layer so your dApp keeps working anyway: health-aware load balancing, automatic failover, circuit breaking, MEV (Jito) routing with RPC fallback and atomic bundles, dropped-transaction rebroadcast, dynamic priority-fee estimation (native/Helius/Triton/QuickNode), WebSocket subscription failover, OpenTelemetry/Datadog/Prometheus observability, and a live diagnostics CLI.
The whole layer is implemented as a custom web3.js v2 RpcTransport, so it composes behind the standard createSolanaRpc API. You get a normal v2 RPC object — your existing code doesn't change.
import { createResilientRpc } from "solana-resilience-sdk";
const rpc = createResilientRpc({
endpoints: [
{ url: "https://your-primary-rpc" },
{ url: "https://your-fallback-rpc", weight: 2 },
],
});
// Exactly the web3.js v2 API — now load-balanced, failed-over and observable.
const slot = await rpc.getSlot().send();web3.js v2.0 builds its RPC client from a swappable transport function:
createSolanaRpcFromTransport(transport). We provide a transport that, on every
call, picks the healthiest endpoint, fails over on error, records metrics, and
trips circuit breakers — all transparently.
createResilientRpc(config)
└─ createSolanaRpcFromTransport( resilientTransport )
resilientTransport:
NodePool.pick() → CircuitBreaker gate → upstream transport
│ │
└──── on failure: backoff + failover to next healthy node
└──── always: record latency/outcome → MetricsCollector → OTel/Datadog
Because it's "just a transport," every concern is isolated and unit-testable, and you keep 100% of the v2 API surface and types.
npm install solana-resilience-sdk @solana/web3.jsRequires @solana/web3.js@^2 (a.k.a. @solana/kit) and Node ≥ 18.
createResilientRpc returns a standard Rpc<SolanaRpcApi>. createResilientClient
additionally exposes the pool, metrics and health checker.
import { createResilientClient } from "solana-resilience-sdk";
const client = createResilientClient({
endpoints: [{ url: primary }, { url: fallback }],
strategy: "least-latency", // or round-robin | least-inflight | weighted-random
healthCheck: { intervalMs: 10_000 },// background getHealth probes
breaker: { failureThreshold: 5, openDurationMs: 10_000 },
});
const slot = await client.rpc.getSlot().send();
console.log(client.metrics.snapshot()); // p50/p95, per-endpoint, per-method
await client.close();- Load balancing across only the endpoints whose circuit is closed/half-open and that pass health checks.
- Automatic failover: a transport error fails the call over to the next-best endpoint with exponential backoff + jitter.
- Circuit breaker per endpoint (closed → open → half-open) keeps traffic off degraded nodes and probes recovery.
- Health checker steers traffic before a user request hits a bad node.
Route through the Jito Block Engine to avoid public-mempool frontrunning; fall back to a normal RPC automatically if the relay is unavailable, and re-broadcast dropped transactions until they confirm or the blockhash expires.
import { createJitoRelay, createResilientSender } from "solana-resilience-sdk";
const relay = createJitoRelay({ blockEngineUrl: "https://mainnet.block-engine.jito.wtf" });
const sender = createResilientSender({ rpc: client.rpc, relay });
const result = await sender.send(
{ base64Transaction, signature, lastValidBlockHeight },
{ commitment: "confirmed", maxRebroadcasts: 5 },
);
// → { signature, route: "jito" | "rpc", rebroadcasts, confirmed }A Jito-routed transaction must tip. relay.tipInstruction({ from }) builds the
required SystemProgram transfer (to a random canonical tip account) as a ready-to-use
web3.js v2 instruction — just append it to your transaction message before signing:
import { appendTransactionMessageInstruction } from "@solana/web3.js";
message = appendTransactionMessageInstruction(
relay.tipInstruction({ from: payer.address }), // default tip = relay.tipLamports
message,
);Atomic bundles. Assemble an ordered, all-or-nothing multi-transaction bundle
(setup → swap → tip) with relay.bundle() — it enforces the 5-transaction Block
Engine limit — then submit it in one shot. One transaction in the bundle carries
the tip (append relay.tipInstruction(...) before signing it):
const bundle = relay
.bundle()
.add(setupTxBase64)
.add(swapTxBase64) // one of these carries the tip instruction
.add(tipTxBase64);
const bundleId = await relay.sendBundle(bundle); // atomic: lands together or not at allAggregate priority-fee estimates from multiple sources (native on-chain fees, Helius, Triton, QuickNode, or any custom source) with caching and safety clamps.
import {
createFeeEstimator,
nativeRecentFeesSource,
heliusPriorityFeeSource,
tritonPriorityFeeSource,
quickNodePriorityFeeSource,
} from "solana-resilience-sdk";
const fees = createFeeEstimator({
sources: [
nativeRecentFeesSource(client.rpc, { percentile: 75 }),
heliusPriorityFeeSource({ url: HELIUS_URL, priorityLevel: "High" }),
tritonPriorityFeeSource({ url: TRITON_URL, percentile: 75 }),
quickNodePriorityFeeSource({ url: QUICKNODE_URL, level: "high" }),
],
aggregate: "max", // safest for landing
multiplier: 1.25,
cacheTtlMs: 2_000,
});
const { microLamportsPerCu } = await fees.estimate({ accounts: writableAccounts });Sources are queried in parallel and individual failures are tolerated — as long
as one source responds, you still get an estimate. Drop in your own by
implementing the one-method FeeSource interface.
Plug any standard wallet (Phantom, Solflare, Backpack, …) into the resilience layer. The wallet signs; broadcast + confirmation go through the resilient sender, so every wallet send gets MEV routing, RPC fallback and rebroadcast.
import { createResilientWalletAdapter, fromWalletStandard } from "solana-resilience-sdk";
const adapter = createResilientWalletAdapter({
wallet: fromWalletStandard(wallet.features["solana:signTransaction"]),
address: account.address,
sender,
});
const { signature, confirmed } = await adapter.signAndSend({ transaction, lastValidBlockHeight });import {
createOpenTelemetryExporter,
createDatadogExporter,
createPrometheusExporter,
} from "solana-resilience-sdk";
client.metrics.addExporter(createOpenTelemetryExporter()); // → OTLP → Datadog/Grafana/Honeycomb
client.metrics.addExporter(createDatadogExporter({ apiKey: DD_KEY })); // direct Datadog metrics intake
const prom = createPrometheusExporter();
client.metrics.addExporter(prom);
const server = await prom.serve({ port: 9_464 }); // GET /metrics for Prometheus to scrapeAll three export request counts, failures, failovers, and a latency histogram,
tagged/labelled by endpoint and method. The Prometheus exporter emits the
standard text exposition format (cumulative counters + _bucket/_sum/_count
histogram series); mount prom.requestListener() on a server you already run, or
use prom.serve(...) to stand one up.
npx srpc doctor --rpc url1,url2 # per-endpoint health, slot, version, latency
npx srpc bench --rpc url1,url2 -c 50 # latency distribution across the pool
npx srpc monitor --rpc url1,url2 # live reliability dashboardSolana RPC Resilience — Live Monitor
requests=42 failures=1 failRate=2.4% p50=48ms p95=179ms
ENDPOINT HEALTH CIRCUIT AVG P95 REQ FAIL INFLT
-----------------------------------------------------------------------------------------
rpc-primary up closed 46ms 120ms 28 0 1
rpc-fallback up half-open 210ms 640ms 14 1 0
createMonitor({ collector, pool }) also exposes snapshot() / onUpdate() for
building your own dashboard.
RPC calls are only half the story — accountSubscribe / slotSubscribe /
signatureSubscribe run over a WebSocket that silently drops when a node
restarts or a connection goes stale, and your stream just stops. createResilientSubscriptions
wraps web3.js v2 subscriptions so a dropped socket transparently reconnects and
fails over to another endpoint — exposed as one continuous async iterable.
import { createResilientSubscriptions } from "solana-resilience-sdk";
const subs = createResilientSubscriptions({
endpoints: ["wss://primary-rpc", "wss://fallback-rpc"], // rotated on reconnect
backoff: { baseDelayMs: 250, maxDelayMs: 10_000 },
});
const stream = subs.subscribe((rpc, signal) =>
rpc.slotNotifications().subscribe({ abortSignal: signal }),
);
for await (const slot of stream) {
console.log("slot", slot); // keeps flowing across drops + node failovers
}- Auto-reconnect with exponential backoff + jitter; backs off only while the stream keeps failing, and reconnects promptly once a connection has delivered.
- Endpoint rotation on every reconnect, so one flaky node can't kill the stream.
- Clean teardown — abort the caller's
AbortSignal(orbreakthe loop) and the underlying socket is closed. ExhaustingmaxReconnectsthrowsSubscriptionClosedError.
The core (resilientSubscription) is transport-agnostic, so the whole
reconnect/failover state machine is unit-tested offline with no sockets.
The suite runs fully offline against a deterministic network simulator
(test/mocks/networkSimulator.ts) that injects latency, dropped/timed-out
calls, HTTP 429 rate-limit bursts, and intermittent errors — with a manual clock
so backoff/expiry/health timing is exact and fast.
npm test # 178 tests, fully offline
npm run test:cov # coverage with enforced thresholds (CI-gated)Statements : 100%
Branches : 100%
Functions : 100%
Lines : 100%
Failure modes covered: failover to a healthy node, circuit open/half-open/close, retry/backoff limits, rate-limit handling, Jito→RPC fallback, fee aggregation with failing providers, dropped-tx rebroadcast, and blockhash-expiry fast-fail.
| Requirement | Where | Test |
|---|---|---|
| web3.js v2.0 compatibility | createResilientRpc → createSolanaRpcFromTransport |
resilientRpc.test.ts |
| Wallet adapter (1+ major wallet) | wallet/adapter.ts (fromWalletStandard) |
wallet.test.ts |
| MEV routing + RPC fallback | relay/jitoRelay.ts, relay/sender.ts |
jitoRelay.test.ts, sender.test.ts |
| Dynamic fee estimates | fees/feeEstimator.ts, fees/providers.ts (native/Helius/Triton/QuickNode) |
feeEstimator.test.ts, providers.test.ts |
| Healthy-node distribution | core/nodePool.ts, core/healthChecker.ts |
nodePool.test.ts, healthChecker.test.ts |
| Observability (OTel/Datadog/Prometheus) | observability/{otel,datadog,prometheus}.ts |
otel.test.ts, datadog.test.ts, prometheus.test.ts |
| Real-time monitor | monitor/monitor.ts |
monitor.test.ts |
| Diagnostics CLI | cli/index.ts (srpc) |
live doctor/bench/monitor |
| 90%+ coverage w/ network sim | test/mocks/networkSimulator.ts + suite |
100% lines / 100% branches, 178 tests |
| Script | Purpose |
|---|---|
npm run build |
Bundle ESM + CJS + types (tsup) |
npm test / npm run test:cov |
Run tests / with coverage |
npm run typecheck |
tsc --noEmit |
npm run example |
Run examples/basic.ts against devnet |
npm run example:subs |
Stream live slots via resilient subscriptions (devnet) |
npm run cli -- doctor |
Run the CLI from source |
- Triton / QuickNode dedicated fee sources (
tritonPriorityFeeSource,quickNodePriorityFeeSource) - Jito tip-transfer instruction helper (
relay.tipInstruction) - Full Jito bundle helper (multi-tx + tip assembly) (
relay.bundle/JitoBundle) - WebSocket subscription failover (
createResilientSubscriptions) - Prometheus
/metricsexporter alongside OpenTelemetry & Datadog (createPrometheusExporter) - Adaptive strategy that auto-switches between load-balancing modes under load
Issues and PRs are welcome. To get set up:
git clone https://github.com/deviverr/solana-resilience-sdk
cd solana-resilience-sdk
npm install
npm run typecheck && npm run test:cov && npm run buildThe whole suite runs offline against the deterministic network simulator, so
npm test needs no RPC endpoint. Please keep coverage above the enforced
thresholds (99% lines / 98% branches) and add a test for any new failure mode.
Built and maintained by deviverr.
MIT © deviverr