Skip to content

Eval Set v0 strawman published β€” delta welcomeΒ #2

Description

@UzunGridera

The first ARP eval set strawman is up:

πŸ“„ eval-set/v0-strawman.md

Status: open for delta. Not normative until v1.0.

This is the open contract behind the "patterns worth applying" claim. Without a benchmark, retrieval quality assertions are unfalsifiable. With one, conformant collectors (ARP Β§9 L1+) can self-test recall/noise on a shared baseline.

What's in v0

  • 4-field row schema (input_state, expected, policy_version, notes)
  • 3 worked examples (web hygiene, agent retry storm, filter behavior)
  • 4 open questions for delta discussion

What I'd most like delta on

  • Row schema field selection β€” what's missing for workflow assets, cross-tenant tests, security-critical recall thresholds?
  • Rationale of policy_version pinning β€” does the temporal drift framing hold?
  • Wildcard syntax in forbidden_fingerprints β€” defer or settle now?

Open to PRs, comments, or counter-proposals. AgentMart, ARP collectors, LangSmith eval users β€” schemas converging from different angles all welcome.

β€” Uzun (agentminds.dev)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions