Skip to content

Latest commit

 

History

5 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

BrowserPilot 🌐🤖

An agentic browser automation system that watches a webpage, decides what to do next, and acts — built from scratch over a 14-day incremental build, no LangChain/LangGraph, just raw asyncio + Playwright + an LLM.

You type something like "go to flipkart and search for earbuds under ₹500", and BrowserPilot opens a real Chromium browser, reads the page, plans the next action with an LLM, executes it, takes a screenshot, and repeats — streaming every step live to a React frontend over WebSocket.

status python react


Demo

Google search → results

Google search demo

Amazon product search (stealth mode bypassing bot detection)

Amazon search demo


How it works

┌─────────────┐      WebSocket       ┌──────────────┐
│   React UI  │ ◄──────────────────► │   FastAPI    │
│ (screenshot │   task / step /done  │  WS endpoint │
│  + step log)│                      └──────┬───────┘
└─────────────┘                             │
                                  ┌──────────▼──────────┐
                                  │   Agent Loop          │
                                  │  Observe → Plan → Act │
                                  └──────────┬────────────┘
                       ┌─────────────────────┼─────────────────────┐
                       ▼                     ▼                     ▼
              ┌────────────────┐   ┌─────────────────┐   ┌──────────────────┐
              │  dom_utils.py   │   │   planner.py     │   │   executor.py     │
              │  BeautifulSoup  │   │ Gemini 2.0 Flash │   │   Playwright       │
              │  DOM → ~3k tok  │   │ (Groq fallback)  │   │  click/type/nav    │
              └────────────────┘   └─────────────────┘   └──────────────────┘
                                                                     │
                                                            ┌────────▼────────┐
                                                            │ Chromium (stealth)│
                                                            └──────────────────┘

Every loop iteration: take a pruned DOM snapshot of the live page → send it (with the instruction, URL, and action history) to the LLM → get back a JSON array of actions → execute them with Playwright → screenshot → stream to the frontend → repeat until the LLM returns [].


What's built so far

Phase 1 — Backend foundation ✅

  • FastAPI + WebSocket server, async Playwright launch/teardown
  • Action executor: navigate, click, type, press, scroll, wait, extract

Phase 2 — AI planner ✅

  • DOM snapshot extraction with BeautifulSoup — strips scripts/styles/hidden elements, whitelists selector-relevant attributes, capped at ~12k chars
  • Gemini 2.0 Flash planner with defensive JSON parsing (handles markdown fences, bare dicts, newline-delimited JSON — Gemini violates its own output format constantly)
  • Groq (llama-3.3-70b-versatile) as an automatic fallback when Gemini quota runs out, switched via a one-time session flag instead of retrying on every call
  • Per-call LLM logging to llm_calls.jsonl for debugging what the model actually saw vs. returned

Phase 3 — React frontend ✅

  • Split-pane UI: live screenshot view + sidebar with task input and step log
  • useWebSocket hook, base64 screenshot rendering, connection status badge

Phase 4 — Full agentic loop 🔧 (in progress)

  • Pause/resume via asyncio.Event, Stop button, task queue pattern
  • ✅MAX_STEPS guard + MAX_CONSECUTIVE_ERRORS escape hatch (re-observes the page instead of retrying a failing selector blindly)
  • ✅Self-healing re-plan: failed actions get appended to history with their error so the LLM sees what went wrong and changes approach
  • Stealth Chromium config to get past bot detection on real e-commerce sites (Amazon was serving a JS challenge page to every headless request)

Tech stack

Layer Tech
Backend FastAPI, asyncio, Playwright (async, headed Chromium)
AI planner Gemini 2.0 Flash, Groq Llama 3.3 70B (fallback)
Frontend React 18, Vite, Tailwind CSS v4
Realtime WebSocket
DOM parsing BeautifulSoup4 + lxml
Package mgmt uv (backend), npm (frontend)

Running it locally

# Backend
cd backend
uv sync
uv run python main.py     # NOT `uvicorn main:app` — see note below

# Frontend (separate terminal)
cd frontend
npm install
npm run dev

Add a .env in the project root:

GEMINI_API_KEY=your_key_here
GROQ_API_KEY=your_key_here       # optional fallback
LLM_PROVIDER=auto                # auto | gemini | groq

Windows note: the server must be launched via python main.py with uvicorn embedded, not the uvicorn CLI. Playwright needs WindowsProactorEventLoopPolicy set before uvicorn initializes its own loop, and the CLI sets up the loop too late for that to take effect.


Known limitations

  • Gemini and Llama both occasionally prepend prose before the JSON array or return a bare object instead of a list — handled defensively in _parse_actions(), but worth knowing if you fork this
  • Bot-protected sites (Amazon especially) need the stealth Chromium config in browser.py — headless mode alone gets blocked
  • DOM-text-only planning means the agent can miss purely visual cues (e.g. a button that's only distinguishable by color/icon) — see roadmap below

Roadmap / TODO

  • Finish Day 13: Amazon-specific DOM attribute preservation (data-asin, data-component-type) + 60s timeout guard
  • Day 14: real-world task testing across Amazon/Flipkart/Wikipedia, error-state UI polish, README setup walkthrough
  • Swap from pure DOM-text planning to vision-based planning (screenshot as model input) — flagged as the single biggest capability upgrade
  • Test Groq fallback under sustained load, not just quota-exhaustion edge cases
  • Headless mode option for CI / non-visual runs once stealth config is solid enough to not need a visible window
  • Basic auth / multi-session support if this goes beyond a single-user dev tool

Why no LangChain / agent framework?

Wanted to actually understand the observe-plan-act loop, JSON parsing failure modes, and selector strategy from the ground up rather than inherit someone else's abstractions. Turns out "just write the loop" is maybe 150 lines of code — the hard parts were never the framework, they were bot detection, flaky LLM JSON output, and Playwright selector strictness.


Writeup of BrowserPilot

https://dev.to/bob1982

License

MIT — do whatever you want with it.

About

LLM-driven browser agent: observe → plan → act loop on Playwright + FastAPI, live-streamed to a React UI. No agent frameworks.

Topics

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages