← Agent readiness

Methodology

How report cards are assessed

A report card answers one question: can an authorized AI agent evaluate, establish access to, pay for, provision, and operate one exact service funnel through scoped authority and explicit resumable human approvals, and where does that path become constrained or break? This page explains sourcey's own assessment and evidence policy. It is not a universal standard, and other assessors legitimately measure different things.

Scope: one funnel, never an entire Entity

Every profile binds one Entity, one product, and one exact funnel, such as an API service lifecycle or a resource-management path. sourcey never publishes an Entity-wide grade: a company with an excellent API funnel and a gated sales funnel is two different answers, not one average.

The five stages

Evaluate
Can an agent find the exact service and decide its terms, eligibility, and cost from stable, readable first-party material?
Sign up
Can an agent begin obtaining service access, operate the controls, and receive scoped authority through safe resumable handoffs without an unsupported CAPTCHA, phone, or shared-credential boundary?
Pay
Is the exact commitment disclosed, and can checkout and payment authorization proceed through deterministic controls with an explicit approval and resumption contract?
Provision
Can an agent initiate provisioning, receive usable access material, and determine completion or reconcile asynchronous failure?
Operate
Can an agent use every essential service target through a stable interface with scoped authentication, decidable operation and failure contracts, and supported credential recovery?

What the states mean

Every stage result is one of five released states. Ready: admitted evidence supports the required conditions. Limited: possible, but an evidenced constraint limits autonomous completion. Blocked: an evidenced condition breaks autonomy at this stage. Unknown: the admitted evidence cannot establish the result, and sourcey says so instead of guessing.Not applicable: evidence establishes the step genuinely does not exist for this funnel, such as payment on a genuinely free tier.

Evidence before grades

Observed signals come from bounded, non-mutating methods: HTTP capture, structured document validation, rendered inspection inside a declared interaction budget, archived history, and reviewed manual observation. The observed lane never creates an account, submits an application, accepts terms, enters credentials or payment details, or invokes a billable service. A negative or not-applicable finding requires the method to carry explicit negative coverage for it: complete retained artifacts of the exact inspected surfaces. Capture and deterministic assessment run without consequential interaction, but an evidence-bound change is admitted only after review and PR merge; that merge is the human publication gate. Later corrections and disputes are recorded and attributed. Failure to find a capability is recorded as unknown, never silently converted into a failure. Stable readable pages can establish a result. Optional structured artifacts can accelerate discovery or provide parser-verified requirement evidence, but their absence is not itself a failing condition and a standards adapter never authors a grade.

The chain · how a fact gets on the record

  1. Capture
  2. Review
  3. Admission
  4. Release

retained evidence · admitted change · signed release · sha256:dcc0c3d9a04d…

Declared is not observed

Vendors and contributors can declare agent surfaces: endpoints, protocols, entry points. Declarations are welcome and provenance-tracked, but they never author a grade. sourcey independently observes the declared surface, and where declaration and observation disagree, the disagreement itself is published.

How the grade is derived

Deterministic policy turns admitted observations into signal outcomes, signal outcomes into stage results, and the five ordered stage results into a letter grade, with failure severity depending on which stage breaks. There is no numeric score, no weighted average, and no model-authored grade: a model may locate evidence, but only deterministic policy over reviewed observations produces the published result. Grades publish only when required signal coverage is complete and evidence is fresh; otherwise the profile stays unrated rather than pretending.

A through C grades count Limited stages; D and F reflect actual Blocked stages by lifecycle severity. Every core graded metric must have supported, fresh, non-conflicting evidence. Barrier checks constrain the report when verified and cap an otherwise higher grade at B+ while unverified. Each stage takes its worst core metric or verified barrier, and the overall grade is derived from the five stage states plus the explicit unverified-barrier cap.

Corrections, disputes, and certification

Reporting outdated evidence, claiming an Entity, requesting a reassessment, attesting an exact revision, and certification are five different acts with different authority, and none of them buys a grade. A future certified journey, a consented real end-to-end run with receipts, is a separate qualifier and never changes the observed grade.

Cohort and limits

The current cohort is purposively selected, not globally representative: high-intent service funnels founders and agents actually ask about. Aggregations state their exact denominators, ordinal grades are never averaged, and unknown or not-applicable results are never counted as failures.

Exact identity

Current release: sha256:dcc0c3d9a04dae17d4c47dd1d0636e71be2060b12d5e5fe31f5be2ddbb0f1a17 · grading policy: sha256:57bf95fd6d0fe6fe9173a68b60e8d8dc2179b357c88ff4349675a72d0e1d4139

Every profile page carries its exact revision, policy, projection, and release digests, and an adjacent JSON representation with the same released semantics.