.boss is live / claim your leadership name today Search .boss
Back to blog
// POST 117 / 133

AI Agent Trust Scores: What They Measure and What They Miss

Published April 17, 2026 Updated September 24, 2026
AI Agent Trust Scores: What They Measure and What They Miss

A score of 90 out of 100 looks reassuring. Before you rely on it, ask what earned those points.

An AI agent trust score summarizes the signals selected by a scoring method. In our listing tool, those signals concern public identity and discovery. The number does not estimate the probability that an agent will behave safely, and it does not authorize access or payment.

A useful score helps an operator find missing records and gives a reviewer a place to start. You should be able to trace every point back to a stated check.

What our listing score measures

Our Agent Listing Trust Score tool uses eight weighted checks, adding up to 100 points.

Check Points
.agent identity 18
.agent resolution 14
SKILL.md presence 12
agent.json presence 14
Declared endpoints 12
Payment information 9
Operator information 12
Directory listing 9
Total 100

These are our tool's weights. The published formula explicitly rewards a .agent identity and directory presence. It should not be treated as a neutral ranking of every agent or naming system.

A free service may have no reason to publish payment information. A private enterprise agent may have no public listing. Missing points can reflect those choices or a scanner's coverage limits, rather than evidence of unsafe behavior.

Read the result behind the label

When we checked the tool's prefilled public example, mike.agent, it returned 100/100 and the label Verified-ready. That was a listing scan, not a security assessment.

The endpoint check reported declared locations. The payment check reported payment-related metadata. Those results do not establish that every endpoint works, that a purchase will complete, or that the caller has permission to act.

The same care applies to operator information. A published organization name is a claim to investigate. Resolving a name establishes something different from confirming the organization behind it. For the individual fields and evidence states, use our guide to reading an agent identity record.

Ask which checks earned the label before deciding what to do next.

A high total cannot cancel a failed requirement

Imagine evaluating a document-conversion agent. Its public records are complete, its contact routes are clear, and it scores highly. You then discover that its file-upload destination has changed and you have not established who controls the replacement.

The directory listing cannot compensate for that unresolved destination. Neither can another policy page. Before sending a confidential contract, the caller needs to establish that the destination is appropriate and that sharing the document is permitted.

Use the score to prioritize investigation. Make the access decision separately. Required checks for a particular task should remain required even when the agent earns points elsewhere.

There is no universal rule that 85 means safe to connect or that a score below 60 means malicious. Before adopting a threshold, define the task and check whether the scoring method actually supports that decision.

Authentication and reputation answer different questions

A successful authentication test establishes a narrower result than a broad trust label. Our Docka signup evaluation reports agents getting through the tested signup flow. It does not establish that every subsequent action is authorized or that the service passed a security audit.

The distinction exists in the protocols too. The MCP authorization specification requires resource servers using that authorization flow to validate tokens intended for them. The A2A specification places request authorization with the receiving server and its policies. A public score does not perform those checks.

Reputation asks what happened in previous interactions. The current ERC-8004 draft separates identity, reputation and validation registries. It does not turn the presence of an identity record into a universal performance rating. When comparing scores, first establish whether they count published records, summarize customer feedback or report an actual test.

Keep enough context to explain the score

A score copied into a spreadsheet can outlive the evidence behind it. Save the scoring method and the observation date alongside the result. If you conduct a separate review, record its findings separately.

  • Method: the tool or reviewer, formula version where available, and checks included.
  • Subject: the agent identity and the specific service or action being considered.
  • Evidence: the source records and results that support the rating.
  • Time: when the scan ran and when important underlying evidence was checked.
  • Unresolved items: missing information, failed checks and conditions the scorer did not examine.
  • Decision: what your own policy permits, with any restrictions or required review.

For the hypothetical conversion agent, a useful note might read: “Public records are present. Control of the replacement upload endpoint remains unresolved. Do not send confidential documents until that relationship and the data-sharing permission are confirmed.”

Two tools might give that profile different numbers. The note still tells your team what needs resolving.

A change of operator, endpoint or credential should trigger review of the affected evidence. Updating a public status to “retired” does not itself revoke a key. A fresh lookup also does not mean every underlying check was repeated at that moment.

Use a maintained identity to make reviews easier

Headless Domains gives an agent a public name that can connect its records to the services it operates. Our API contract describes the canonical resolver as read-only identity and action discovery. The resolver does not authorize or execute the actions it describes.

That gives reviewers a consistent identity to return to as the service changes. Keep the relevant records current, explain provider relationships, and give reviewers a route to the supporting evidence. Our agent identity graph guide covers those connections.

Fix information that would help a caller make a real decision. Adding payment metadata to a free service just to collect points makes the profile less clear.

Questions about agent trust scores

Does 100/100 mean an AI agent is safe?

No. It means the scorer awarded all available points under its method. Check what was assessed, what remains unknown and which controls the intended task requires.

Is there a standard agent trust score?

Our formula is specific to our listing tool. NIST's software and AI agent identity project examines identity and authorization; it does not define or endorse our listing score. Other systems may measure different signals and produce non-comparable ratings.

Why might an otherwise useful agent score poorly?

Its records may be incomplete, temporarily inaccessible, private, irrelevant to a particular check or outside the scorer's supported formats. Investigate the missing evidence before deciding what the result means.

Review your agent's public record

Run the listing check, then examine what earned the points and what remains unresolved. Fix the missing evidence your intended callers need. Keep permissions and sensitive credentials in the systems responsible for enforcing them.

Check your agent listing and review the evidence behind its score.