Methodology

How report cards are assessed

A report card answers one question: can an authorized AI agent independently evaluate, establish access to, pay for, provision, and operate one exact service funnel, and where does autonomy first break? This page explains sourcey's own assessment and evidence policy. It is not a universal standard, and other assessors legitimately measure different things.

Scope: one funnel, never a vendor

Every profile binds one vendor, one product, and one exact funnel, such as an API service lifecycle or a resource-management path. sourcey never publishes a vendor-wide grade: a company with an excellent API funnel and a gated sales funnel is two different answers, not one average.

The five stages

Evaluate
Can an agent discover the funnel and determine terms, eligibility, and price from machine-readable material?
Sign up
Can an agent enter and complete onboarding without an undelegated human-only gate such as a CAPTCHA or phone verification?
Pay
Can an agent understand and complete any required economic step on a supported rail, without sales intervention?
Provision
Does successful onboarding produce usable access: automatic activation, credential delivery, machine-readable status, bounded delay?
Operate
Can an agent use and recover the service afterward: API access, usable authentication, documented limits and errors, delegated recovery?

What the states mean

Every stage result is one of five released states. Ready: admitted evidence supports the required conditions. Limited: possible, but an evidenced constraint limits autonomous completion. Blocked: an evidenced condition breaks autonomy at this stage. Unknown: the admitted evidence cannot establish the result, and sourcey says so instead of guessing.Not applicable: evidence establishes the step genuinely does not exist for this funnel, such as payment on a genuinely free tier.

Evidence before grades

Observed signals come from bounded, non-mutating methods: HTTP capture, structured document validation, rendered inspection inside a declared interaction budget, archived history, and reviewed manual observation. The observed lane never creates an account, submits an application, accepts terms, enters credentials or payment details, or invokes a billable service. A negative or not-applicable finding requires the method to carry explicit negative coverage for it, and every negative finding receives human review. Failure to find a capability is recorded as unknown, never silently converted into a failure.

The chain · how a fact gets on the record

  1. Authoring
  2. Review
  3. Admission
  4. Release

one atomic effect · what every footer proves ·sha256:fbebb3b52183…

Declared is not observed

Vendors and contributors can declare agent surfaces: endpoints, protocols, entry points. Declarations are welcome and provenance-tracked, but they never author a grade. sourcey independently observes the declared surface, and where declaration and observation disagree, the disagreement itself is published.

How the grade is derived

Deterministic policy turns admitted observations into signal outcomes, signal outcomes into stage results, and the five ordered stage results into a letter grade, with failure severity depending on which stage breaks. There is no numeric score, no weighted average, and no model-authored grade: a model may locate evidence, but only deterministic policy over reviewed observations produces the published result. Grades publish only when required signal coverage is complete and evidence is fresh; otherwise the profile stays unrated rather than pretending.

Complete, fresh evidence for the core service lifecycle first applies hard stage severity, then grades the share of applicable evidenced signals that are constrained. Optional blocker probes affect the grade when established and remain visible as unknown when public evidence cannot establish them. Any incomplete, contradictory, unsupported, or non-fresh required evidence produces an unrated profile. A failed Evaluate, Sign up, or Pay stage derives D; a failed Provision or Operate stage derives F; otherwise constrained-signal ratio derives A+ through C.

Corrections, disputes, and certification

Reporting outdated evidence, claiming a vendor, requesting a reassessment, attesting an exact revision, and certification are five different acts with different authority, and none of them buys a grade. A future certified journey, a consented real end-to-end run with receipts, is a separate qualifier and never changes the observed grade.

Cohort and limits

The current cohort is purposively selected, not globally representative: high-intent service funnels founders and agents actually ask about. Aggregations state their exact denominators, ordinal grades are never averaged, and unknown or not-applicable results are never counted as failures.

Exact identity

Current release: sha256:fbebb3b521836bb3594967c8473c3934901240d4c0d73ceace004453591a5808 · grading policy: sha256:ffb2babe9adfc714aa365dfdd8c4185dfd5d4fdb7157052091a795a6066dbe0e

Every profile page carries its exact revision, policy, projection, and release digests, and an adjacent JSON representation with the same released semantics.