An agentic browser automation system that watches a webpage, decides what to do next, and acts — built from scratch over a 14-day incremental build, no LangChain/LangGraph, just raw asyncio + Playwright + an LLM.
You type something like "go to flipkart and search for earbuds under ₹500", and BrowserPilot opens a real Chromium browser, reads the page, plans the next action with an LLM, executes it, takes a screenshot, and repeats — streaming every step live to a React frontend over WebSocket.
Google search → results
Amazon product search (stealth mode bypassing bot detection)
┌─────────────┐ WebSocket ┌──────────────┐
│ React UI │ ◄──────────────────► │ FastAPI │
│ (screenshot │ task / step /done │ WS endpoint │
│ + step log)│ └──────┬───────┘
└─────────────┘ │
┌──────────▼──────────┐
│ Agent Loop │
│ Observe → Plan → Act │
└──────────┬────────────┘
┌─────────────────────┼─────────────────────┐
▼ ▼ ▼
┌────────────────┐ ┌─────────────────┐ ┌──────────────────┐
│ dom_utils.py │ │ planner.py │ │ executor.py │
│ BeautifulSoup │ │ Gemini 2.0 Flash │ │ Playwright │
│ DOM → ~3k tok │ │ (Groq fallback) │ │ click/type/nav │
└────────────────┘ └─────────────────┘ └──────────────────┘
│
┌────────▼────────┐
│ Chromium (stealth)│
└──────────────────┘
Every loop iteration: take a pruned DOM snapshot of the live page → send it (with the instruction, URL, and action history) to the LLM → get back a JSON array of actions → execute them with Playwright → screenshot → stream to the frontend → repeat until the LLM returns [].
Phase 1 — Backend foundation ✅
- FastAPI + WebSocket server, async Playwright launch/teardown
- Action executor:
navigate,click,type,press,scroll,wait,extract
Phase 2 — AI planner ✅
- DOM snapshot extraction with BeautifulSoup — strips scripts/styles/hidden elements, whitelists selector-relevant attributes, capped at ~12k chars
- Gemini 2.0 Flash planner with defensive JSON parsing (handles markdown fences, bare dicts, newline-delimited JSON — Gemini violates its own output format constantly)
- Groq (
llama-3.3-70b-versatile) as an automatic fallback when Gemini quota runs out, switched via a one-time session flag instead of retrying on every call - Per-call LLM logging to
llm_calls.jsonlfor debugging what the model actually saw vs. returned
Phase 3 — React frontend ✅
- Split-pane UI: live screenshot view + sidebar with task input and step log
useWebSockethook, base64 screenshot rendering, connection status badge
Phase 4 — Full agentic loop 🔧 (in progress)
- Pause/resume via
asyncio.Event, Stop button, task queue pattern - ✅
MAX_STEPSguard +MAX_CONSECUTIVE_ERRORSescape hatch (re-observes the page instead of retrying a failing selector blindly) - ✅Self-healing re-plan: failed actions get appended to history with their error so the LLM sees what went wrong and changes approach
- Stealth Chromium config to get past bot detection on real e-commerce sites (Amazon was serving a JS challenge page to every headless request)
| Layer | Tech |
|---|---|
| Backend | FastAPI, asyncio, Playwright (async, headed Chromium) |
| AI planner | Gemini 2.0 Flash, Groq Llama 3.3 70B (fallback) |
| Frontend | React 18, Vite, Tailwind CSS v4 |
| Realtime | WebSocket |
| DOM parsing | BeautifulSoup4 + lxml |
| Package mgmt | uv (backend), npm (frontend) |
# Backend
cd backend
uv sync
uv run python main.py # NOT `uvicorn main:app` — see note below
# Frontend (separate terminal)
cd frontend
npm install
npm run devAdd a .env in the project root:
GEMINI_API_KEY=your_key_here
GROQ_API_KEY=your_key_here # optional fallback
LLM_PROVIDER=auto # auto | gemini | groq
Windows note: the server must be launched via
python main.pywith uvicorn embedded, not theuvicornCLI. Playwright needsWindowsProactorEventLoopPolicyset before uvicorn initializes its own loop, and the CLI sets up the loop too late for that to take effect.
- Gemini and Llama both occasionally prepend prose before the JSON array or return a bare object instead of a list — handled defensively in
_parse_actions(), but worth knowing if you fork this - Bot-protected sites (Amazon especially) need the stealth Chromium config in
browser.py— headless mode alone gets blocked - DOM-text-only planning means the agent can miss purely visual cues (e.g. a button that's only distinguishable by color/icon) — see roadmap below
- Finish Day 13: Amazon-specific DOM attribute preservation (
data-asin,data-component-type) + 60s timeout guard - Day 14: real-world task testing across Amazon/Flipkart/Wikipedia, error-state UI polish, README setup walkthrough
- Swap from pure DOM-text planning to vision-based planning (screenshot as model input) — flagged as the single biggest capability upgrade
- Test Groq fallback under sustained load, not just quota-exhaustion edge cases
- Headless mode option for CI / non-visual runs once stealth config is solid enough to not need a visible window
- Basic auth / multi-session support if this goes beyond a single-user dev tool
Wanted to actually understand the observe-plan-act loop, JSON parsing failure modes, and selector strategy from the ground up rather than inherit someone else's abstractions. Turns out "just write the loop" is maybe 150 lines of code — the hard parts were never the framework, they were bot detection, flaky LLM JSON output, and Playwright selector strictness.
MIT — do whatever you want with it.

