CatalogClawHub · rustyorb

Agent Evaluation

Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks Use when: agent testing, agent evaluation, benchmark agents, agent reliability, test agent.

Pending analysisCommunity189 installs

Savant verdict: Pending analysis

Imported from the source hub; Savant's analysis is still running.

Live evaluation

Not evaluated yet. Workspaces can request a live evaluation.

Safety (NVIDIA SkillSpector)

Scan pending.

SKILL.md

This listing is enumerated but its package hasn't been fetched yet; it's queued by popularity.