A local-first, multi-agent voice assistant. A browser mic feeds a FastAPI/WebSocket backend that runs speech detection, transcription, LLM reasoning, and text-to-speech synthesis, streaming audio back while the LLM is still generating — no cloud STT/TTS APIs in the loop. The focus is inference-path latency and correctness under real conversational conditions (interruptions, background noise, multiple speakers), not just a chat UI with a mic bolted on.
Stack: FastAPI + WebSocket · LangGraph + Claude Haiku · faster-whisper / Parakeet MLX · Kokoro / Piper · Silero VAD · Resemblyzer · Next.js.
See ARCHITECTURE.md for the full technical write-up: sentence-chunked streaming, barge-in, speaker verification, multi-agent orchestration, and the WebSocket protocol.
- Streamed everywhere. The LLM response is split into sentences as it's generated; each sentence is synthesized and streamed to the browser as soon as it's ready, so perceived latency is time-to-first-sentence, not time-to-full-response.
- Real barge-in. A second, always-on mic tap runs Silero VAD continuously. On confirmed speech, playback pauses immediately and the turn is interrupted once the speaker is verified — not before.
- Speaker verification. Resemblyzer enrolls a voiceprint from the first genuine utterance and checks every later hands-free utterance against it, so a TV or another person in the room can't hijack the session.
- Multi-agent. Sessions can run several agents with distinct personas, routed by intent, sharing a running memory and awareness of each other's answers.
- Local inference. STT (faster-whisper / Parakeet MLX), TTS (Kokoro / Piper), and VAD (Silero) all run on-device — only the LLM call goes over the network.
make install # venv + backend deps + frontend deps
make dev # backend (:8000) + frontend (:3000) togetherRequires ANTHROPIC_API_KEY (and optionally EXA_API_KEY for web search) in a
repo-root .env.