On 2026-07-22, Lucas Cooper-Bey (San Diego) had his personal AI system audit itself against the three biggest open agentic-infrastructure projects in the world: Garry Tan's G-Brain, Nous Research's Hermes Agent, and OpenClaw. Then it spent the same day closing its biggest measured gap, and re-graded itself that night. This page is the record: identity cards, an interactive visual deck, a 12-dimension scorecard with receipts, and the before/after "ghost race."
An Intelligent OS is not one app. It is the whole operating layer a person builds around their AI tooling: memory with provenance, retrieval, a skill library, agents, scheduled autonomy, eval harnesses, security rails, and business integration. "Lucface IOS" is the author's instance of that idea.
Reading notes. The page is written in second person: "you" means the operator, because the system wrote it for him. GHOST is the system as it stood that morning; NOW is the same system that night. Scores are 0-10 per dimension, synthesized from receipts. N/A is never plotted as zero. Rival receipts come from their public repos, release ledgers, and published security research, date-stamped 2026-07-22 (stars and versions drift fast, the shapes drift slowly).
Honesty. This is a self-audit, with the incentives that implies. The counterweights: every score carries a receipt, the retrieval claim was settled by installing a real G-Brain and racing it on identical golden questions the same day, and the losing dimensions are printed with the same size font as the winning ones. Not affiliated with, or endorsed by, any of the three projects. → run this exact audit on your own system (section 8)
INGRESS HARNESS MEMORY EVALS SECURITY BUSINESS
& voice & skills & truth & meas. arch. & output
OpenClaw ████████ ██████ ███ ██ ██ ███
Hermes Agent ████████ ████████ ███ ███ ████ ████
G-Brain (host's) █████ ███████ ████████ █████ ██
LUCFACE ghost ░░░░ ░░░░░░░░ ░░░░░░░░ ░░░ ░░░░░░░░ ░░░░░░░░ this morning
LUCFACE NOW ████ ████████ ████████ ██████ ████████ ████████ tonight
▲ ▲
your thinnest wall the wall you BUILT today
(voice · deferred M6, (evals ░░░ → ██████ : substrate
plugins already loaded) 0.914 + retrieval 0.76 + census;
the --llm judge = last two blocks)
Each public project is tall in 1-2 columns. This morning you were
tall in 4. Tonight you are tall in 5, and the short one is chosen.
Nine charts, three logic trees, one timed day. Every chart answers one question and states its honest read under it; hover anything for exact values; click any heatmap cell to drill into its receipt. Charts are self-contained in this file (Apache ECharts inlined, zero external requests); every number also lives in section 3's bars and tables.
Twelve dimensions, 0-10, scored from the receipts in the table below each group. N/A means the dimension is out of the project's scope (a layer is not an assistant), never scored as zero. Scores are this audit's synthesis; every claim traces to a source.
Channels a human can reach the system through, and whether it can talk and listen live.
Does useful work happen without a human in the chair, and does dead work get loud.
Can the system prove where a fact came from, and does history survive updates.
Hybrid search quality and whether anyone has actually measured it.
How much packaged capability exists and how it is kept trustworthy.
Does the system get better on its own, and is that safe.
Is "does the right capability fire, and is the output good" measured, or vibes.
Fan-out, pipelines, cross-model adjudication, quarantine tiers.
Injection defense, blast-radius control, secrets, supply chain.
Local models, own hardware, no single vendor able to turn you off.
Finance, pricing gates, CRM, legal, client delivery, approval-gated publishing.
Who else improves the system while you sleep.
| Dimension | Ghost (am) | NOW (pm) | Δ | What moved it |
|---|---|---|---|---|
| D4 Retrieval, measured | 8 | 9 | ▲ +1 | hit@5 0.76 weekly-tracked + beat a fresh gbrain 0.90 to 0.80 on identical goldens; co-leads with G-Brain |
| D6 Self-improvement | 7 | 8 | ▲ +1 | librarian nightly consolidation live: 51 stale + 11 dupes + 9 dead paths proposed on day one |
| D7 Evals + dispatch | 4 | 7 | ▲ +3 | substrate evals + top1 emission metric + 130-orphan census, mutation-tested through 3 Codex passes, weekly job; --llm judge still spec |
| D2 Autonomy | 9 | 9 | = | +3 scheduled jobs (98 total), same guarantees |
| D3 Provenance | 9 | 9 | = | sha-bound eval ledger rows added; RFC-3161 still the missing point |
| D9 Security | 9 | 9 | = | gateway current + scan gate exercised; one internal hygiene item open |
| D1 / D5 / D8 / D10 / D11 / D12 | 5/9/9/9/10/2 | 5/9/9/9/10/2 | = | unchanged (D1 and D12 by explicit decision) |
| Leads or co-leads | 7 of 12 | 8 of 12 | ▲ | D4 joins as co-lead; D7 and D6 remain the two live gaps, both narrowed |
| Dimension | Lucface | G-Brain | OpenClaw | Hermes | Verdict |
|---|---|---|---|---|---|
| D1 Ingress + voice | 5 | N/A | 10 | 9 | Real gap: no live voice loop, thin mobile |
| D2 Autonomy + scheduling | 9 | 8 | 8 | 7 | You lead on depth and deadness-detection |
| D3 Truth + provenance | 9 | 7 | 4 | 3 | Your signature strength, field-wide |
| D4 Retrieval, measured | 8 | 9 | 6 | 5 | Par on architecture; receipts landed same-day (0.76 overall; 0.90 vs his 0.80 head-to-head, §6) |
| D5 Skill library | 9 | 6 | 8 | 9 | Co-lead: their breadth, your curation |
| D6 Self-improvement | 7 | 8 | 8 | 9 | They automate consolidation; you gate it |
| D7 Evals + dispatch | 4 | 9 | 3 | 4 | Was THE delta; closed same-day at substrate tier (§6), --llm judge = v2 |
| D8 Orchestration | 9 | 6 | 8 | 8 | Council + Workflow + lanes lead |
| D9 Security architecture | 9 | 6 | 3 | 5 | You chose architecture over speed; the field proves why |
| D10 Compute sovereignty | 9 | 2 | 8 | 8 | You + local stack lead; G-Brain is cloud-only |
| D11 Institution integration | 10 | 3 | 4 | 5 | Nobody else is playing this game |
| D12 Ecosystem | 2 | 7 | 10 | 9 | Structural: they have thousands, you have you |
The question that started this audit. G-Brain is a memory + eval layer, not an assistant, so the honest comparison is on the five axes where you both play.
| Axis | Your receipt | His state |
|---|---|---|
| Provenance discipline | Append-only spine (house standard), originSessionId on nodes, evidence notes that grade claims (your 07-20 note graded HIS keynote line by line) | Citations + gap analysis shipped; "librarian" and contradiction-checking still talk-track, per recon |
| Curation practice | MEMORY-ARCHIVE, ground-truths audits, keeper, shelf-review, dedupe-on-write memory rules | Preaches the hygiene triad on stage; dream cycle covers part of it |
| Corpus breadth | 61 collections spanning YC (~28.6K pts), Every (~22K), failure-intel DB (56K companies), sessions, wiki, personal | One brain repo, 220K pages, single-domain (his work life) |
| Sovereignty | Local embeddings, local Qdrant, local whisper/TTS/gen on the GPU box | Hosted APIs only, ~$100-150/mo, cloud-dependent |
| Everything downstream | Orchestration, security airlock, business controllers, output rails | Out of scope for G-Brain by design |
| Axis | You | Him |
|---|---|---|
| Hybrid retrieval | Dense + sparse twins (YC corpus hybrid-twinned), RRF-style merge via /recall doors | HNSW + BM25 + RRF + source-tier boost, one engine, two backends (PGLite / pgvector) |
| Knowledge graph | Wiki backlinks + [[memory links]] + graphify on demand | Self-wiring typed graph, zero LLM calls, +31.4pt P@5 attributable and MEASURED |
| Nightly maintenance | Pre-4am brief fleet, capability-pulse, session sweeps | Dream cycle: entity sweep, citation repair, consolidate, resync |
| Axis | His receipt | Your state | The move |
|---|---|---|---|
| Trigger / dispatch evals | gbrain routing-eval (intent→skill accuracy vs fixtures) + check-resolvable (MECE/DRY/orphan audit of the skill tree) | Was: H-042 open, ~90 capabilities measured going unremembered. Now: substrate evals SHIPPED same-day, 0.9143 over 70 fixtures, 130 orphans censused, weekly job live (§6) | M1: DONE |
| Retrieval benchmarks | BrainBench P@5 49.1 / R@5 97.9; LongMemEval 97.6% recall claim; regression-tracked per release | Was: zero benchmark numbers. Now: hit@5 0.76 across 6 collections, weekly job live, plus the 0.90-vs-0.80 head-to-head win (§6) | M2+M3: DONE |
| Skill optimization | skillopt treats SKILL.md as a trainable parameter, keeps only measured improvements (~$1, ~20 min) | skill-creator + definition-gate lint structure, nothing measures skill OUTPUT quality over versions | M2/M3 |
| Community leverage | 42-50 contributors incl. his own agents (garrytan-agents) opening PRs; a funded YC company building on it | n=1 plus your agents; deliberate, but it means every improvement costs your tokens | M7, optional |
| # | Move | Why now | Size |
|---|---|---|---|
| M1 | Ship dispatch evals (close H-042): build a fixtures file of task-utterance → expected-skill pairs (seed it from capability-nudge-map's mined keyword fingerprints), run it as a tier-1 deterministic check on a cron, report weekly hit-rate. Steal the shape of gbrain routing-eval + check-resolvable directly, MIT license. | It is the single axis where the primary rival is genuinely ahead, your own audit already tracks it, and the fixture data already exists on disk | S-M |
| M2 | Benchmark /recall, BrainBench-style: 20-40 golden questions across the 5 doors with known-correct sources, measure P@5/R@5, track per month. Reuse gbrain-evals' harness shape; your fuse-eval dir is prior art in-house. | "Par on architecture, zero receipts" is exactly the position your stack-reality rule exists to prevent | M |
| M3 | Run the G-Brain bake-off instead of guessing: point a PGLite G-Brain at a mirror of the business-brain markdown (~1 evening), run the same 20 golden questions against it and /recall, keep whichever wins per query class, steal the losers' tricks (source-tier boost, session demote, graph adjacency boost). | TOPIC-BANK already says "evaluate as prior art before building in this lane"; now it costs one evening and answers M2's calibration question for free | M |
| M4 | Upgrade the OpenClaw gateway to current stable, then re-harden (MCP session isolation, provenance-gated skill installs shipped upstream in July as direct CVE responses). It is your only internet-facing always-on agent surface; the pre-upgrade build (a stale June build behind a wrong unit label) already had TaskFlow but predated July's Skill Workshop, cloud workers, and isolation work. | Monthly CVE cadence upstream + 42,665 exposed instances found in the wild; staleness is exposure | S |
| M5 | Stand up your dream-cycle analog for backlogs: a nightly librarian pass with a pruning quota over INBOX (237 untriaged) and RUNBOOK (786 open), consolidating and closing, not just appending. The sentinels capture; nothing yet prunes. | Garry's "garbage dump with great search" failure mode is measurable in your own queues right now | M |
| M6 | Decide the voice-ingress play deliberately: Hermes ships full local voice (faster-whisper + local TTS both free); your phone-path design from 06-02 is archived. Either schedule it or write one line in ETHOS saying terminal-first is the choice. | D1 is your widest visible gap vs the field; an undecided gap costs attention every audit | decision |
| M7 | Optional flywheel play: publish the M2/M3 benchmark as SDAIS content ("I benchmarked my second brain against the YC CEO's"), with numbers and the harness repo. | It is the rare post where the receipts already exist and the audience is the studio's exact ICP | S after M2/M3 |
| Move | Result | Where it lives |
|---|---|---|
| M1 dispatch evals | SHIPPED + RUN: substrate rate 0.9143 (64/70 fixtures), 130 orphaned capabilities censused; 6 misses diagnosed; weekly Mon 06:20 | dispatch-evals harness · weekly job (private repo) |
| M2 retrieval bench | SHIPPED + RUN: hit@5 0.76 overall (llm-wiki 1.00, session-memory 1.00, YC library 1.00, YC blog 0.75, every-corpus 0.40, failure-intel 0.33); weekly Mon 06:35 | retrieval-bench harness · weekly job (private repo) |
| M3 G-Brain bake-off | RAN: supply-chain scan PASS, 400 public YC pages indexed on PGLite; gbrain 0.80 vs Lucface 0.90 on identical goldens; 1.00 vs 1.00 dead heat on the fully-covered slice; 3 steal-worthy tricks documented | bake-off + supply-chain scan notes (private repo) |
| M4 gateway | DONE: stale June build → current July stable, service unit made Node-bump-proof, messaging channel verified up; config backup taken | gateway box · systemd user unit · config backup |
| M5 librarian | LIVE: propose-only nightly 03:40; first pass found 51 stale, 11 duplicate pairs, 9 dead paths across INBOX (237 open) + RUNBOOK (786 open) | librarian harness · nightly job (private repo) |
| M6 voice ingress | DECIDED: deferred behind eval work, revisit post-SDAIS-launch; the upgraded gateway's new talk-voice + phone-control plugins are the cheapest path when revisited | decision recorded to memory |
| M7 content kit | BUILT: receipts-only writing kit (never ghostwritten), voice-score before posting | writing-kit dir (private repo) |
Scorecard note: section 3 now shows both runners per dimension — the GHOST (audit-time, translucent) and NOW (post-execution, solid). Re-graded NOW scores: D7 4→7, D4 8→9 (co-lead), D6 7→8; leads went 7 of 12 → 8 of 12 in one day. After the pp2p loop the three harnesses sit at v1.4.x: 4/4 mutations caught, 9/9 failure scenarios produce error rows instead of clean numerics, one final Codex-proven census bug fixed and live-verified. Infra note: the second box's Ollama is healthy; box-to-box embedding is closed by deliberate ACL design, so benches use the Mac-local fallback deliberately.
OpenClaw CVE-2026-25253 (RCE, 8.8) · Claw Chain 4-flaw · Zenity
prompt-injection persistent backdoor · ClawHavoc: 1,184
malicious skills / 247,693 installs · 42,665 exposed
Hermes 9 CVE-class findings in 4 days (May) · self-written
skills auto-loaded = memory-poisoning surface (Critical)
G-Brain smallest surface, best behavior of the three; still
ships "shell jobs read any file" as an accepted risk
Lucface quarantine-by-architecture BEFORE the field's incidents:
untrusted content cannot issue tool calls; leak guard on
push; supply-chain gate on install; destructive-command
hook (fired twice during this audit, correctly)
Lesson: the market optimized for growth and is now buying back
security. You paid for security first. Do not trade that away
for channel breadth.
Your on-disk evidence note records ~350K pages from the Lightcone captions; live recon could verify no figure above 220K (the 07-17 keynote, his highest public number; README says 146,646). Trajectory is the real signal either way: 17,888 pages on launch day to 220K in 3.5 months. Cite 220K.
Three quarantined research-proxy agents ran the web recon; six fetched pages tripped the injection sanitizer, all six triaged as false positives (site chrome, docs boilerplate, one page with 41 zero-width characters noted), zero instructions followed from fetched content. Hard numbers came from GitHub's REST API, not scraped HTML. Raw API responses parked in this session's scratchpad for re-verification.
Live: github.com/garrytan/gbrain (+ gbrain-evals, gstack) · github.com/NousResearch/hermes-agent (API + release ledger) · github.com/openclaw/openclaw/releases · openclaw Foundation coverage (thenewstack, computerworld, CNBC) · TechCrunch 07-13 (Nous valuation) · Cloud Security Alliance note (Hermes CVEs) · thehackernews (OpenClaw CVEs) · Orply + Tokenless (AI Engineer World's Fair closing keynote, 07-17) · vectorize.io (GBrain architecture reviews). Local: private evidence notes, corpora, ledgers, and the same-day capability inventory (internal).
The scorecard is a method, not a trophy. The prompt below runs the same 12-dimension, receipts-first audit on any machine with Claude Code (best on Fable or Opus at max effort). It inventories what actually exists on your disk, scores each dimension with a receipt, calibrates against the same reference field as this page, and renders you a dossier of your own.
1 · Open Claude Code at your home directory (or wherever your system lives). 2 · Paste the whole prompt. 3 · You get a terminal scorecard plus a self-contained HTML dossier. If you then build one of your ranked moves the same day, it re-grades you: your own ghost vs NOW.
Every run can go on the ledger: post your scorecard as an issue at Lucface/intelligent-os-field-audit/issues and it gets appended to LEDGER.md, the running record of harness shapes. Not a leaderboard: nobody wins a scorecard, the interesting part is how differently people build. One-person institutions, big-rig locals, file-system RAG purists, cloud maximalists, all of it belongs. The prompt also lives in the repo as prompt.md (MIT).
You are running the Intelligent-OS Field Audit v1 (method of 2026-07-22) on THIS machine.
An "Intelligent OS" is the whole operating layer a person builds around their AI tooling: memory, retrieval, skills, agents, scheduled autonomy, evals, security rails, business integration. Your job: audit the one on this machine, receipts first, and score it against the reference field below. Work at maximum effort. Do not ask the user anything; work from the disk and state assumptions inline.
RULES (non-negotiable)
1. Receipts before scores. Inventory by actually looking: count skills/commands/agents (list their directories), scheduled jobs (crontab -l; launchd or systemd user units), hooks, memory/notes files, vector collections, eval harnesses and their result ledgers. Run read-only commands freely; invent nothing. Anything you cannot verify is written "unverified" and scores as absent.
2. Measured beats claimed. A benchmark number in a ledger outranks an architecture that "should" work. If a dimension has zero measurements, the receipt must say so.
3. N/A is never 0. If a dimension is out of scope BY DESIGN (and you can point to where that was decided), grade it N/A and exclude it from lead counts.
4. No flattery. The gaps are the product. Classify every trailing dimension as CHOSEN (deliberate, cite the decision) or EARNED (real gap; name the smallest artifact that would close it).
5. Zero-regression honesty. If you re-grade later the same day, never silently lower a score; state what changed.
THE 12 DIMENSIONS (score 0-10)
D1 Ingress + voice — channels a human can reach it through; live voice loop. (2 one CLI · 5 CLI + one chat channel · 8 multi-channel + mobile · 10 many channels + wake word + live voice)
D2 Autonomy + scheduling — does useful work happen with nobody in the chair, and does dead work get loud? (2 manual · 5 a few crons · 8 dozens of jobs + failure detection · 10 self-healing fleet with heartbeats)
D3 Memory truth + provenance — can it PROVE where a fact came from; does history survive updates? (2 chat logs · 5 curated notes · 8 append-only + source IDs · 10 hashed lineage on every derived value)
D4 Retrieval, measured — hybrid search quality AND whether anyone measured it. Hard cap at 6 if zero benchmark numbers exist, regardless of architecture.
D5 Skill library quality × breadth — packaged capability, and how it is kept trustworthy (linting, curation, install gates).
D6 Self-improvement loop — does it get better on its own, and is that safe? Ungated autonomous self-modification also caps D9 at 5.
D7 Evals + dispatch measurement — is "does the right capability fire, and is the output good" measured, or vibes? (0-4 vibes · 5-7 deterministic fixtures on a schedule · 8-10 plus a semantic judge and regression history)
D8 Multi-agent orchestration — fan-out, pipelines, cross-model adjudication, quarantine tiers.
D9 Security architecture — injection defense, blast-radius control, secrets handling, supply chain. Score the architecture, not the absence of incidents.
D10 Compute sovereignty — local models, own hardware, no single vendor able to turn it off.
D11 Institution + business integration — finance, CRM, legal, client delivery, approval-gated publishing.
D12 Ecosystem + momentum — who else improves it while the operator sleeps. (n=1 by design is a legitimate CHOSEN 2.)
REFERENCE FIELD (as sourced 2026-07-22; D1..D12 in order; null = N/A; re-verify anything that matters to you)
{
"gbrain": {"scores":[null,8,7,9,6,8,9,6,6,2,3,7], "note":"Garry Tan's memory/eval layer: hybrid retrieval w/ published benchmarks (BrainBench), routing-eval + resolvability audits, hosted-API only"},
"openclaw": {"scores":[10,8,4,6,8,8,3,8,3,8,4,10], "note":"biggest ecosystem (384K stars, 52K skills) + best ingress; monthly CVE cadence and a poisoned-registry incident"},
"hermes": {"scores":[9,7,3,5,9,9,4,8,5,8,5,9], "note":"~20 channels + full local voice, autonomous skill Curator; zero memory provenance"},
"lucface_now": {"scores":[5,9,9,9,9,8,7,9,9,9,10,2], "note":"the author's one-person system, post-execution 2026-07-22: dispatch evals 0.914 weekly, retrieval hit@5 0.76 weekly, beat a live gbrain install 0.90-0.80 on identical goldens; leads/co-leads 8 of 12"}
}
DO, IN ORDER
1. Inventory sweep (read-only). Enumerate and count the primitives above. Note the 3 most surprising findings.
2. Score all 12 dimensions with a one-line receipt each. The receipt is the evidence itself, not a justification.
3. Comparison table: your row against the reference field. Mark lead / co-lead / trail per dimension; count leads (N/A dims excluded).
4. Verdict in three honest sentences. Then classify every trailing dimension CHOSEN vs EARNED.
5. Ranked moves: the 3 smallest artifacts that would most move the EARNED gaps. Steal proven shapes: a routing-fixtures file scored weekly (does the right capability fire), a golden-set retrieval bench (hit@5 over 20+ questions with known answers), a nightly propose-only librarian over your backlogs. Size each S/M/L.
6. Render a single-file dark-theme HTML dossier (inline CSS, zero external requests, robots noindex): a scorecard section with one bar row per runner per dimension and the receipt beside it, the comparison table, the verdict, the ranked moves. Save it and open it.
7. Ghost mechanic (optional, recommended): if the operator executes a move today, re-run scoring and show GHOST (before) vs NOW (after) rows for the changed dimensions.
Print the scorecard table and verdict in the terminal as well, not only the file.
Method + reference dossier: https://lucface.github.io/intelligent-os-field-audit/ — post your scorecard as a GitHub issue there and it joins the ledger (LEDGER.md), the running record of harness shapes.