claw-score
Audit or refresh OpenClaw maturity scorecard docs from root taxonomy, maturity scores, and QA evidence artifacts without using maintainer discrawl data or committed inventory reports.
Savant verdict: Analyzed
Scanned or evaluated by Savant; not yet both passing.
Live evaluation
Not evaluated yet. Workspaces can request a live evaluation.
Safety (NVIDIA SkillSpector)
Risk score
0/100
Recommendation
CAUTION
Severity
LOW
Savant decision
Passed (low risk)
SkillSpector rated it CAUTION, but its risk score is below 20/100, so Savant's policy passes it. Its findings are still listed for review. SkillSpector 2.12.0, static analysis.
0 patterns found
Structure
- No license declaredConfirm you may reuse this skill before importing it into your repository.
SKILL.md
---
name: claw-score
description: Audit or refresh OpenClaw maturity scorecard docs from root taxonomy, maturity scores, and QA evidence artifacts without using maintainer discrawl data or committed inventory reports.
---
# claw-score
Use this skill when working on the OpenClaw maturity scorecard in this repo.
This is the openclaw-local version of the maintainer `claw-score` workflow:
it keeps the taxonomy and scorecard concepts, but excludes discrawl and the old
committed `inventory/` report tree.
## Authority
This skill owns the operational workflow for:
- `taxonomy.yaml`
- `qa/maturity-scores.yaml`
- `docs/concepts/qa-e2e-automation.md`
- `qa/scenarios/index.yaml`
Keep person-specific, maintainer-private, Discord archive, and discrawl facts
out of this repo. If a score needs private evidence, use the redacted
`qa-evidence.json` artifact shape generated by OpenClaw QA workflows.
## Source Model
- `taxonomy.yaml` is the hand-edited source of truth for surfaces, levels,
QA profiles, categories, feature coverage IDs, docs refs, LTS overrides, and
completeness-instruction paths.
- Each feature has exactly one `coverageIds` entry. Keep that evidence ID
unique to the feature; broader many-to-many evidence mapping is not part of
the current taxonomy schema.
- Coverage IDs use dotted `namespace.behavior` form, with lowercase
alphanumeric/dash segments. Profile, surface, and category IDs may remain
dashed or dotted.
- Keep categories and feature names unique, product-shaped, and broader than raw
coverage IDs. Do not promote generic IDs into standalone feature names.
- Avoid duplicate coverage-ID bundles under different feature names in one
category.
- `qa/maturity-scores.yaml` is the committed aggregate source for Quality,
Completeness, and LTS review state.
- `extensions/qa-lab/src/scorecard-taxonomy.ts` exports
`readValidatedQaMaturityScoreSources`; use it to validate score output.
- Generated public docs are `docs/maturity/scorecard.md` and
`docs/maturity/taxonomy.md`; both come from `pnpm maturity:render`. Do not
hand-edit generated Markdown to change score results.
- `qa-evidence.json` artifacts provide per-run QA scorecard evidence. Release
profile artifacts are the source of truth for Coverage. They can enrich
generated artifact docs, but they are not committed as inventory.
## Commands
Run from the openclaw repo root.
Validate taxonomy YAML structure and the maturity score schema after source
edits:
```bash
node --import tsx --input-type=module <<'NODE'
import fs from "node:fs";
import YAML from "yaml";
import { readValidatedQaMaturityScoreSources } from "./extensions/qa-lab/src/scorecard-taxonomy.ts";
for (const file of ["taxonomy.yaml", "qa/scenarios/index.yaml"]) {
YAML.parse(fs.readFileSync(file, "utf8"));
}
readValidatedQaMaturityScoreSources();
NODE
```
Check docs when touching docs prose:
```bash
pnpm check:docs
```
Run focused QA/profile checks when changing coverage IDs or profile membership:
```bash
pnpm openclaw qa coverage --json
```
## Full Generation Runs
For a direct full scorecard run that publishes the generated-doc pull request,
use floating `main` resolution by default:
```bash
gh workflow run maturity-scorecard.yml \
--repo openclaw/openclaw \
--ref main \
-f ref=main \
-f expected_sha='' \
-f publish_pull_request=true \
-f allow_failures=true
```
Do not resolve `main` locally and pass that commit as both `ref` and
`expected_sha` for an ordinary manual generation run. OpenClaw's `main` moves
quickly, so the caller-selected commit can become stale before validation. The
workflow then correctly rejects publication when the pull request base contains
newer maturity inputs, and QA never starts.
With `ref=main` and a blank `expected_sha`, the workflow's
`floating_default_branch` path fetches and freezes the current remote default
branch inside validation before handing an immutable revision to downstream
jobs. Use an explicit SHA only when the requested evidence must remain bound to
that exact revision, such as a release-candidate workflow call or an
artifact-only historical reproduction. If that exact-revision run also requests
publication and `main` has changed relevant inputs, expect validation to fail and
dispatch again from floating `main` instead.
## Scoring Workflow
When asked to score or refresh a surface:
1. Read the surface in `taxonomy.yaml`.
2. Read the surface completeness rubric under
`.agents/skills/claw-score/references/completeness/`.
3. Gather public repo evidence from docs, source, tests, and QA scenario
metadata.
4. Prefer existing release profile `qa-evidence.json` artifacts for executed
proof.
5. Update `qa/maturity-scores.yaml` only for Quality, Completeness, and LTS
review state backed by public or redacted artifact evidence.
6. Run the schema validation command from this skill.
7. Run `pnpm check:docs` if docs prose changed, and focused QA coverage checks
if coverage IDs or profile membership changed.
For subjective score changes, make the smallest defensible edit and leave the
evidence path in the PR or task summary. Keep manual prose in current docs and
keep score data in `qa/maturity-scores.yaml`.
## Default Completeness Process
Completeness is scored against the intended operator-visible workflow for each
category, not against test breadth or implementation quality. The completeness
reference files under `references/completeness/` define the category scope and
any surface-specific variation from this default process.
By default, Completeness measures how fully OpenClaw exposes the intended
surface capability set to the user, operator, author, or maintainer persona for
that surface. Score whether each category delivers the full expected workflow,
including setup, normal use, status or inspection, recovery, and important
platform, provider, channel, security, or lifecycle variants where they apply.
Treat `Surface-Specific Scoring Questions` and `Surface-Specific Guidance` as
higher-priority instructions for that surface. The surface instructions may
flesh out, narrow, or intentionally conflict with the default ideas here; when
they do, follow the surface instructions and make the score rationale reflect
that surface-specific instruction. If a reference file does not include
surface-specific questions or guidance, apply this default process to the
surface's `Category Scope`.
For each category, ask:
- Can the intended user or operator complete the category workflow end to end?
- Are the taxonomy features present as supported capabilities rather than
isolated implementation fragments?
- Are the important lifecycle stages represented: setup, normal operation,
status/inspection, recovery, and upgrade or removal where relevant?
- Are the important environment, provider, platform, channel, or security
branches present for this surface?
- Do the known gaps leave major user-visible capability branches missing?
Default guidance:
- Favor higher Completeness when the category supports the full
operator-visible workflow described by taxonomy and category evidence.
- Lower Completeness when only the happy path exists, when important variants
are undocumented or unimplemented, or when recovery/status paths are missing.
- Do not lower Completeness because tests are thin; that is Coverage.
- Do not lower Completeness because implementation quality is fragile; that is
Quality.
Default Completeness bands:
- `Clawesome` (95-100): complete across expected workflows, variants, and
recovery branches, with only minor polish gaps.
- `Stable` (80-95): the expected workflow set is broadly present, with only
bounded missing branches.
- `Beta` (70-80): the main workflow exists, but meaningful branches or recovery
paths are still absent.
- `Alpha` (50-70): only a partial capability set is present; users can complete
some core tasks but not the full expected workflow.
- `Experimental` (0-50): the category exposes only fragments of the intended
capability.
## Decision Context
Record an optional `decision` beside `score` and `label` for surface and category
Quality/Completeness, or beside `supported` for category LTS. In `taxonomy.yaml`,
use optional `level_decision` beside the canonical surface `level`.
Each record contains `value`, `rationale`, `reviewer`, `evidence_refs`, and
`revalidate_when`. Use an integer from 0–100 for Quality/Completeness, a boolean
for LTS, and a declared taxonomy level ID for `level_decision`. Supply nonempty
text fields and at least one evidence reference. Name the actual reviewer and
the condition that should trigger another review.
Leave unavailable history absent: it is unknown, not an invitation to invent
reviewers, rationale, or evidence. A record does not overwrite the current score,
support flag, or canonical level. If its value differs, retain both; generated
docs show a non-gating mismatch, including under strict input validation.
Do not attach decisions to Coverage, computed rollups, surface LTS summaries, or
the copied level in score aggregates. Decision context does not change coverage
identity, score calculations, support commitments, or release gates.
## Score Semantics
- Coverage: deterministic release validation coverage derived from the release
profile `qa-evidence.json.scorecard` feature fulfillment data.
- Quality: reliability, maintainability, operator safety, and regression
confidence for the category.
- Completeness: how much of the intended operator-visible workflow exists for
the category. Use the default completeness process plus any surface-specific
variation before changing this score.
- LTS: derived from Quality, release-evidence Coverage, and
`human_lts_override`; do not hand-edit generated Markdown to change LTS
status.
Bands:
- `Clawesome`: 95-100
- `Stable`: 80-95
- `Beta`: 70-80
- `Alpha`: 50-70
- `Experimental`: 0-50
## Artifacts
Do not add the maintainer repo's `docs/kevinslin/maturity-scorecard/inventory/`
tree to openclaw. Evidence-enriched scorecard outputs belong in short-lived
artifacts, not committed generated docs, unless this repo adds an explicit
renderer/check workflow first.