← Retour à /research
Classement · WAB
Top 100 · workspaces agentiques audités · 12 piliers · L0–L4
| # | Workspace | Type | Grade | Score | cluster | ELO | Piliers matures | Point faible | Stack | Audité | Evidence |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Madani Workspace• B2B services portfolio · iter-39 | ref | A | 87.08 | 87 A95 B89 C73 D | 1900 | — | — | Claude Code · Python · n8n · launchd · auto-promote-engine | 2026-05-25 | audit ↗ |
| 2 | Hermes Agent · NousResearch• skill-curator + RL self-evolution | ext | C | 50.83 | 30 A63 B53 C60 D | 1650 | — | — | Python · agent/curator.py · skill_manage · GRPO | 2026-05-24 | audit ↗ |
| 3 | OpenClaw• agentic platform · plugin ecosystem | ext | D | 47.50 | 23 A57 B58 C50 D | 1580 | — | — | TypeScript · Node.js · plugin system | 2026-05-24 | audit ↗ |
| 4 | OpenAI Agents SDK · Python• agent SDK library | ext | D | 40.83 | 23 A42 B49 C50 D | 1450 | — | — | Python · agents framework | 2026-05-20 | audit ↗ |
| 5 | Cline · IDE Agent• VS Code agentic IDE | ext | D | 32.50 | 13 A47 B35 C35 D | 1480 | — | — | TypeScript · VS Code extension | 2026-05-24 | audit ↗ |
| 6 | Anthropic Cookbook• code-sample repository | ext | F | 27.50 | 7 A33 B35 C35 D | 1380 | — | — | Python · Jupyter · Claude Agent SDK | 2026-05-20 | audit ↗ |
Affichage 6 sur 6
Légende
✓ verified · audit vérifié par les maintainers du benchmark.
• self-reported · audit exécuté par le submitter · re-audit serveur prévu en v0.5.
Piliers matures · nombre de piliers au niveau maximum de maturité (L4 Optimizing) sur 12 au total. Ex. 9/12 = 9 piliers à L4.
Point faible · le pilier avec la maturité la plus basse · où le workspace a le plus grand écart.
Cluster A·B·C·D · moyennes des 4 clusters (Cognition, Action, Trust, Operations).
ELO · dérivé du composite (1200 + composite × 8). Même composite → même ELO.
Score · composite 0-100 · moyenne pondérée également des 12 piliers.
Niveaux L0-L4 · L0 absent · L1 ad hoc · L2 documenté · L3 automatisé · L4 optimizing (auto-improve).
Composite = moyenne pondérée 4 clusters · ELO Bradley-Terry · ~70% audit déterministe · IRR 1.0 vérifié. Reference entries vérifiées dans le benchmark repo. Community submissions stockées en Vercel KV live · re-audit CI roadmap v0.5.
