TapHound provides two integration surfaces for Agents:
- Use
taphound verify --jsonto deterministically verify an existing Journey. - Use a Project Context and
taphound generation ... --jsonto generate a new Journey, which TapHound then fully Replays from the initial state and publishes after verification.
An external Agent may analyze source code, judge whether a goal is complete, and propose the next step, but TapHound Core never invokes a model. Project Context validation, device-state binding, proposal validation, risk confirmation, ADB execution, final Replay, and assertions are all handled by deterministic code.
TapHound intentionally stops at the Journey boundary. External Workflow Skills own requirement analysis, planning, coding, build/install, multi-Case orchestration, completion gates, and diagnosis. An orchestrator can invoke TapHound once per independent Case and adapt the public JSON, Report, and evidence paths into its own Task/Result protocol. Workflow correlation and Requirement/Plan identities remain outside TapHound.
For one Case, an external orchestrator may supply the optional Skill-level
journeyBrief binding:
{
"path": ".android-agent-workflow/req-search-001/cases/CASE-002/taphound-journey-brief.md",
"sha256": "<SHA-256 of the exact file bytes>"
}The Markdown format is defined by
taphound-journey-brief.example.md.
This is not a Core CLI argument. It provides static Case hints to the Journey
Skill; Project Context, live Snapshot binding, risk policy, and final Replay
remain authoritative.
The taphound-journey-brief-author Skill
is the recommended producer of the Brief. It combines Android source analysis
with read-only taphound observe to author one Brief per Case, then
returns {path, sha256} for the Journey Skill to consume. It uses only
read-only commands and never modifies device state.
A multi-Case orchestrator dispatches one brief-author subagent per Case.
Configure the subagent with a name and a PROMPT field whose content
is copied verbatim from
brief-author-role.md.
That file is a lean bootstrap: it establishes the role, capability boundary,
inputs, output format, and key rules, then directs the subagent to read
SKILL.md from the installed skill directory for the full execution
procedure. This keeps the PROMPT short enough for agent runtimes that
impose a length limit on the PROMPT configuration field. The orchestrator then
dispatches a dynamic task message per Case with these explicit inputs:
| Field | Required | Description |
|---|---|---|
project |
yes | Android project root path |
caseGoal |
yes | One Case's test scenario (natural language) |
caseId |
no | Case identifier for frontmatter |
contextPaths |
no | Explicit path array; the subagent reads ONLY these files for surrounding context |
observeSnapshot |
no | Pre-captured taphound observe --json result |
output |
no | Brief output path (defaults to .taphound/journeys/taphound-journey-brief.md) |
The subagent returns a single JSON summary:
{
"status": "authored",
"caseId": "CASE-002",
"path": ".taphound/journeys/taphound-journey-brief.md",
"sha256": "<64-char hex hash>",
"edgesVerified": 2,
"edgesNeedsObservation": 1
}The orchestrator never re-parses raw exploration content from the subagent; it consumes only this structured summary.
The brief-author subagent MUST NEVER search for or assume files named
plan.md, requirement.md, or any convention. It reads ONLY files the
orchestrator explicitly passes via contextPaths. If no contextPaths are
supplied, it works from caseGoal alone plus source code and Project Context.
This rule is encoded in both SKILL.md and brief-author-role.md.
To enable parallel brief authoring across multiple Cases without device
contention, the orchestrator pre-captures one taphound observe --json
snapshot and passes it to all parallel brief-author subagents via the
observeSnapshot input. When observeSnapshot is provided, the subagent
uses it directly and MUST NOT call taphound observe itself.
# Orchestrator captures once, before dispatching parallel subagents:
taphound observe --project <project> --device <serial> --logcat-lines 200 --json
# Then passes the result as observeSnapshot to each parallel subagent.The Brief is an inspectable artifact between the planning phase and Journey
generation. After a brief-author subagent returns status: "authored", the
orchestrator should present the Brief path and summary to the user for
Review before dispatching the downstream Journey generation subagent with
the journeyBrief: {path, sha256} binding. Review is a Skill convention,
not a Core CLI gate; the Journey Skill's own Brief validation (SHA-256
check, frontmatter, required sections, Goal match) remains enforced.
Orchestrator
|-- dispatch brief-author subagent (per Case, parallel)
| output: {path, sha256, caseId, edgesVerified, edgesNeedsObservation}
|
|-- human Review (brief is an inspectable artifact)
|
|-- dispatch journey subagent
input: {project, goal, journeyBrief: {path, sha256}}
output: {journeyPath, reportPath, verified}
A typical flow: a developer uses Claude Code or another Agent CLI to implement a requirement, then has the Agent call TapHound Journey to verify whether the code meets expectations.
taphound verify \
--project /workspace/android-app \
--config /workspace/android-app/.taphound/config.json \
--journey /workspace/android-app/.taphound/journeys/search.json \
--device emulator-5554 \
--json- The stdout of
verify --jsoncontains exactly one JSON value and a trailing newline, with no progress text. - stderr receives pre-checks, progress, and diagnostics, which the Agent may save separately.
- The process exit code matches the JSON
exitCode. 0means passed;1is a product verification failure;2is invalid input;3is an unavailable environment;4is a TapHound internal error.- When a report exists, read
reportPath,report.primaryFailure,report.secondaryErrors, and the layered results. - When no report exists, read the top-level
failure.codeandfailure.message.
Do not merely search stdout text for "passed"; first check the process status and exitCode, then read the structured fields.
import { spawn } from "node:child_process";
const child = spawn("taphound", [
"verify",
"--project", projectRoot,
"--journey", journeyPath,
"--json"
], { shell: false });
let stdout = "";
let stderr = "";
child.stdout.setEncoding("utf8").on("data", chunk => { stdout += chunk; });
child.stderr.setEncoding("utf8").on("data", chunk => { stderr += chunk; });
child.on("close", code => {
const result = JSON.parse(stdout);
if (code !== result.exitCode) throw new Error("TapHound exit contract mismatch");
// Feed result.report.primaryFailure back to the development Agent.
});The caller must also use an argument array and keep shell: false, to avoid turning project paths or user input into a Shell command.
Generation sessions bind one exact UI backend descriptor (id, adapter and
engine version, and configuration hash). Runtime Snapshot v2 carries that
descriptor, the physical-display viewport, observation ID, and capture timing.
An Agent must submit the exact referenced snapshot and must not substitute a
snapshot captured by another backend. A legacy session without evidence may
bind once on its first snapshot; a legacy session with evidence must be
restarted rather than migrated implicitly.
ui.backend=auto chooses only at provider open and does not silently fail over
during an authoritative operation. appium-uiautomator2 is opt-in and supplies
only the page-source tree; application lifecycle, actions, waiting, Logcat, and
verification remain TapHound/ADB responsibilities.
The generation flow uses the in-repo taphound-journey-generator Skill:
project describe --jsonoutputs stable Package and Activity information.- The
taphound-journey-brief-authorSkill analyzes each Gradle module independently and produces the Project Context root index plus module shards. It owns the full Context lifecycle: initial generation,context refreshre-hashing, and full regeneration. context listexposes the compact module catalog.context validate/context statuscheck the index, shard hashes, source evidence, and per-module file inventory.context refreshrecomputes evidence hashes for an existing Context without re-analyzing source.journey list-flows --jsonvalidates reusable local prefixes. The Agent selects the deepest applicable valid Flow, never by filename alone. The first resolved step must begin at a stable Activity deterministically reached after cold launch. A transient Splash must not be required to remain foreground;core/launch-homeshould be await: Home -> Homereadiness anchor with an expectation for a unique Home element.generation start --module ... [--base-flow ...]binds the project, config, selected module dependency closure, device, and optional cleanly replayed Flow prefix. Without a Base Flow, Core force-stops, launches the configured Activity, and waits for the App process before creating the session.run.activityis only the cold-launch entry; the subsequent observation's idle/layout checks establish the stable post-redirect state.- The Agent uses
generation observe --compact --json, reads the project- relative authoritativesnapshotRef, and submits a proposal strictly bound to that full snapshot. Compact successful steps returnnextBindingandnextSnapshotRef; the Agent reads the reference before the next proposal. Active references point into the Store-owned.<generationId>.workbundle; publication atomically moves the same evidence into the final<generationId>bundle. It usesgeneration confirmwhen human approval is required, andgeneration manualfor local TTY overrides. Confirmation defaults to a local prompt. In a non-TTY sandbox, the Agent may pass--decision approve|declineonly after the user explicitly reviews that exact challenge; the decision remains bound to Core-owned evidence. generation statusexposes durable state. Interrupted work is retried only after explicitgeneration recover --decision retryacknowledgement.generation finalize --detachsurvives caller interruption and fully Replays from the initial state. The Journey and immutable evidence are published only after exact verification passes.generation list --jsonenumerates all sessions in the workspace (active, archived, and published).generation archive --session <id>marks an idle active session as archived so it no longer clutters active listings. Archive is only permitted on sessions with no in-flight step or pending confirmation; recoveryRequired sessions must be recovered first.
taphound project describe --project /workspace/android-app --json
taphound context validate \
--project /workspace/android-app \
--context .taphound/context/project-context.json \
--json
taphound journey list-flows \
--project /workspace/android-app \
--json
taphound generation start \
--project /workspace/android-app \
--context .taphound/context/project-context.json \
--module :feature:search \
--base-flow search/open \
--device emulator-5554 \
--jsonThe application module is always selected; dependencies declared by selected modules are expanded automatically. Omitting --module selects all modules. The exact root-index hash and selected shard IDs/hashes are returned as contextSelection and bound to the session. Modules cannot be added later. The device is bound at generation start. generation observe, step, confirm, manual, status, and recover use that binding via --session and do not accept --device; generation finalize reloads exactly the bound module set and may explicitly provide --device, but must not change the session identity binding.
context refresh recomputes evidence hashes for an existing Context without re-analyzing source. It backfills the optional semanticSha256 for every evidence file, rehashes files whose change was formatting or comments only, repairs index entries whose shard hash drifted, and rewrites only the shards and index that actually changed.
taphound context refresh \
--project /workspace/android-app \
--context .taphound/context/project-context.json \
--jsonRefresh never invents semantic knowledge. It stops with exitCode: 1 and status: "blocked" when evidence changed semantically, when a module's file inventory changed, or when an evidence file is missing or unreadable, because those cases need module re-analysis. --module <id...> limits refresh to selected modules. --accept-source-changes additionally rehashes semantically changed evidence and inventory drift; use it only when the recorded module summary is still accurate, since the summary itself is not updated. Missing or unreadable evidence always blocks.
By default, changed source evidence stops generation with CONTEXT_STALE. For frequent implementation-only edits, an agent may explicitly pass --allow-evidence-drift to both generation start and generation finalize. This does not bypass Context shard integrity, project/config/session bindings, locator safety, or final replay verification. It only allows the validator's evidence-file drift result to proceed; the final replay remains authoritative. JSON output reports evidenceDriftAllowed: true when this opt-in is active.
Generation's --json commands likewise write only one machine-readable JSON
value to stdout and indicate the result with exitCode. observe always
returns snapshotRef; --compact omits the duplicate inline snapshot.
step, confirm, and manual similarly replace nextSnapshot with
nextSnapshotRef in compact mode. The referenced file is still the full
RuntimeSnapshot required by the proposal envelope. The Agent must retain
generationId, baseRevision, snapshotHash, and that exact snapshot, and
must not fabricate or reuse expired bindings. Step results include phase timing
for freshness, evidence setup, observation, action, idle wait, expectations,
Logcat, and optional next observation. Detached finalize progress and stdout
live under .taphound/build/jobs/<generationId>/, outside the authoritative
bundle. For the full protocol, retry rules, and Context update strategy, see
the Skill's GUIDE.md.
The generation step --input envelope is a strict object with exactly three
top-level fields. Unknown, missing, or flat (unwrapped) fields are rejected as
CONTEXT_INVALID; the JSON failure includes a hint describing the required
shape:
proposal.binding must match the nextBinding/binding returned by the most
recent observe (or prior step), and snapshot must be that exact
RuntimeSnapshot. activity.after and expect are optional on a proposal;
Core records the observed post-action Activity and evaluates any supplied
expectation.
When Base Flow verification fails, generation start --json returns
FLOW_REPLAY_FAILED details containing the Flow name, Verify report path,
primary failure, the failed step's Activity/locator/expectation summary, and
recovery guidance. The Agent must not silently skip reuse or treat a device
already showing Home as an exact replay. It should repair or re-record the
Flow, and restart without --base-flow only after the user explicitly chooses
that bypass.
When an indexed Locator is resolvable from the bound snapshot, Core persists
versioned, non-geometric semantic evidence for the selected element. Replay
recomputes that evidence before mutation and fails instead of using annotated
fallback when the indexed element's represented content changed. Existing
Journeys without this optional evidence retain their previous ordinal
behavior.
generation status includes pendingConfirmation and its computed expired
flag. While a challenge is pending, observe returns
RISK_CONFIRMATION_REQUIRED with challenge details rather than an internal
error. An expired challenge cannot be approved; resolve it with the exact
challenge ID (for example confirm --decision decline) before observing and
submitting a fresh proposal.
For approved actions, the challenge ID and approvalMode (localTty or
delegated) are stored in the in-flight attempt before device mutation and in
the immutable step result, so interrupted-action recovery retains the approval
audit.
taphound init copies the TapHound Journey Generator Skill from the npm package into
each Agent's Skill directory. The interactive multi-select requires choosing
at least one Agent; you can also specify non-interactively with --agent:
taphound init --agent claude,codex,cursor,droid --jsonGlobal install (user-level directory):
taphound init --agent claude --globalSupported Agents and paths:
| Agent | Project-level path | User-level path |
|---|---|---|
| Claude Code | .claude/skills/ |
~/.claude/skills/ |
| Codex | .agents/skills/ |
~/.agents/skills/ |
| Cursor | .cursor/skills/ |
~/.cursor/skills/ |
| Droid | .factory/skills/ |
~/.factory/skills/ |
| Other | .agents/skills/ |
~/.agents/skills/ |
The Skill ships with the npm package (assets/skills/), and taphound init
copies it into the target Skill root. Re-running init overwrites existing
files in the installed Skill directory.
After implementation is complete, run:
taphound verify --project . --journey .taphound/journeys/search.json --json
Parse the JSON; acceptance passes only when exitCode=0.
If it fails, report report.primaryFailure first, and include reportPath.
Do not modify the Journey to mask implementation defects.
TapHound Replay never invokes AI. The Agent may select an existing Journey or propose new steps in a generation session, but the final judgment of Locator, Activity, Layout Diff, risk policy, and Expect is performed by deterministic code. The Agent must not automatically loosen assertions, swap the Package, delete steps, or bypass confirmation after a failure.
Generation binds the normalized config for the lifetime of the session. Agents
must choose idle.strategy before generation start and must start a new
session after any config change. hybrid falls back from active frame counters
to Core-owned UIAutomator layout hashes; layoutDiff selects structural
stability directly. An action attempt that returns
status: "recoveryRequired" may already have executed. Inspect
generation status, obtain explicit user approval before recover, and never
assume that recovery committed the interrupted action or returned a snapshot.
TapHound Core enforces single-package determinism. When a step causes the
foreground to leave the configured target package — for example tapping a
button that opens the system camera, a file picker, or a third-party share
sheet — generation step and generation manual reject the post-action
snapshot with PACKAGE_ESCAPE. This is by design: TapHound cannot
deterministically bind or replay actions that execute in a process it does not
own.
The bridge action lets Core own the trigger click and the return detection
for cross-app flows within a single generation step. Core clicks the trigger,
detects the package escape, optionally executes a bound External Flow's steps
inside the escaped package, waits for the foreground to return, and captures
the post-return snapshot.
Bridge steps run in two replay modes:
- Auto (
replayMode: "auto"):--flow <name>resolves a bound External Flow. Core stamps the flow's steps asexternalStepsand replays them deterministically duringfinalizewith no human operator. - Manual (
replayMode: "manual"): no--flow. A human operator completes the external action during replay.
Bind External Flows at session start:
taphound generation start \
--project . \
--external-flow camera/photo-capture \
...List available flows (built-in and project):
taphound journey list-flows --project . --include-external --jsonUse generation bridge with one of the built-in scenarios:
photoCapture— system camera (validates escaped package)pickImage— system image picker (validates escaped package)pickFile— system file picker (validates escaped package)custom— any other cross-app flow (skips package validation, requires--description)
# Auto bridge (deterministic, no operator)
taphound generation bridge \
--project . \
--session <id> \
--scenario photoCapture \
--trigger-locator '{"resourceId":"com.example.app:id/camera_button"}' \
--flow camera/photo-capture \
--return-timeout-ms 60000 \
--escape-timeout-ms 3000 \
--json
# Manual bridge (human operator completes external action)
taphound generation bridge \
--project . \
--session <id> \
--scenario photoCapture \
--trigger-locator '{"resourceId":"com.example.app:id/camera_button"}' \
--return-timeout-ms 60000 \
--jsonLike generation manual, bridge goes through risk confirmation. If the
action is not auto-approved, the response carries status: "confirmationRequired" and the Agent calls generation confirm with the human
decision.
BRIDGE_NO_ESCAPE— the foreground did not leave the target package withinescapeTimeoutMs(default 3 seconds) of the trigger click.SCENARIO_PACKAGE_MISMATCH— the escaped package is not in the known system package list for the selected scenario. Usecustomto bypass.BRIDGE_NOT_RETURNED— the foreground did not return to the target package withinreturnTimeoutMs.EXTERNAL_FLOW_NOT_FOUND—--flownames a flow not bound to this session.EXTERNAL_FLOW_STALE— the bound flow file changed sincegeneration start.EXTERNAL_PACKAGE_MISMATCH— an external step ran in a different package thanescapedPackageName.EXTERNAL_ACTIVITY_MISMATCH— an external step'sexpectedActivitydid not match.EXTERNAL_STEP_FAILED— an external step's action failed (e.g., locator not found).EXTERNAL_LOCATOR_STRICTNESS— an external step locator lacks aresourceId(v1 requires XML-only resource IDs for external steps).MANUAL_STEP_REQUIRED— a non-interactivefinalizeencountered areplayMode: "manual"step. Bind an External Flow or run finalize in a TTY.
See docs/journey-schema.md for the full bridge schema
and examples/bridge-camera.journey.json
for a complete Journey.
{ "version": 1, "proposal": { "action": "click", "locator": { "resourceId": "btn_more" }, "binding": { "generationId": "<session-id>", "baseRevision": 2, "snapshotHash": "<sha256-of-snapshot>" }, "activity": { "before": "com.example.app.AIChatActivity" } }, "snapshot": { /* full RuntimeSnapshot returned by generation observe */ } }