Trust & methodology · reviewed Aug 2, 2026

How the system works

The scoring, the refresh rules, and where the line runs between what we measured and what the model wrote.

How is the security score broken down?

Every profile carries five 1-10 sub-scores under Security breakdown. Four measure protection, so higher is safer. The fifth measures risk, so higher is worse.
  • Sandboxing — isolation from the host OS. 10 is fully virtualised or containerised; 1 is direct local execution.
  • API security — how external integrations are handled. 10 is scoped, encrypted and multi-user safe; 1 is plaintext keys.
  • Network isolation — outbound traffic control. 10 is air-gapped or strictly whitelisted; 1 is unrestricted internet access.
  • Telemetry safety — what the tool reports back. 10 is no telemetry; 1 is extensive logging to external servers.
  • Shell access risk (higher is riskier) — 10 is raw unmonitored shell access; 1 is no unsupervised execution.
The 0-100 security score shown on profiles and compare pages is separate: it weighs these axes together with the wider evidence we hold.

What is measured, and what is written by AI?

Measured: GitHub metadata, releases, stars, commit history, language and runtime clues — anything directly observable. AI-written: clone summaries, "why choose this over OpenClaw" blocks, compare verdicts, tradeoffs and confidence notes, generated from that evidence. Community: Reddit and public search coverage, used as supporting evidence. The ecosystem report keeps the same split — its tables are computed, its commentary is labelled.

What does evidence confidence mean?

How strong the backing evidence looks behind an AI-heavy claim. High usually means multiple direct signals agree. Low often means the repo is young, the docs are thin, or sources conflict. Read low-confidence content as a directional nudge, not a verdict.

Why do some fields say unknown or mixed?

Because unknown is more honest than filler. If a project does not clearly publish something like deployment posture or privacy behaviour, we show unknown or mixed rather than polished speculation.

How often does everything refresh?

Targets: repo data nightly, clone profiles reviewed weekly, compare verdicts checked more aggressively, methodology pages monthly. These are targets, not guarantees — the About page shows the dates that actually landed, so you can check the gap yourself.

What triggers an update outside the normal cycle?

Security incidents, major releases, architecture changes, momentum spikes, confidence drops, or a meaningful change in OpenClaw itself. Credible community reports of a factual miss also send a profile back for review.

Which external sources feed the analysis?

GitHub for repository facts, Reddit for community discussion, and public search and news coverage for wider signals. They are complementary evidence, not a replacement for reading the repo.

Should I trust the compare verdict or the raw metrics?

The verdict is a shortcut; the metrics are the evidence. If the two disagree, the evidence wins.

Can I suggest a correction or a new clone?

Yes. For a missing project use the Add clone button in the header; for a wrong or stale figure open a correction issue. Both land as public GitHub issues, and a link to the source beats a description. We would rather queue a review than leave a misleading claim live.

Something looks misread or stale? Send the evidence and we will review it.

Nominate a clone

Add a new Claw

Paste a GitHub repository and tell us why it belongs on the tracker.

Opens a prefilled issue on GitHub — every nomination is public. Comfortable with a PR? Adding the repo to projects.json is faster.