Field Audit · Lucface Intelligent OS vs the open-primitive field · 2026-07-22 · self-audit, receipts sourced

Where a one-person Intelligent OS stands vs G-Brain, Hermes, and OpenClaw

One line Your system is the only full-stack institution in the set; G-Brain leads you on exactly one axis you care about, measurement (shipped routing evals plus benchmarked retrieval), and one you deliberately do not play, open ecosystem. The gist 1) G-Brain = Garry Tan's MIT memory/eval layer: 26,762 stars, Postgres + pgvector, hybrid RRF, a self-wiring graph, and a real eval suite (routing-eval, check-resolvable, skillopt, BrainBench). It validates your spine thesis and ships the trigger-eval primitive your H-042 is missing. 2) Hermes Agent = 218,903 stars, ~2,156 contributors, v0.19 two days ago, full local voice, self-writing skills with an autonomous Curator, $1.5B valuation talks. Its memory has zero provenance (your 06-13 deep-dive still holds). 3) OpenClaw = 384K stars, now a 501(c)(3) Foundation with OpenAI/NVIDIA/Microsoft sponsors, and a monthly CVE cadence with 42,665 exposed instances found. The operator's own gateway had silently been running a stale June build behind an even staler version label; this session it was upgraded to current stable, verified active, picking up July's Skill Workshop, cloud workers, and MCP session-isolation hardening. 4) You lead or co-lead 7 of 12 dimensions (autonomy, provenance, skill curation, orchestration, security architecture, compute sovereignty, business integration). You trail on dispatch evals, voice/mobile ingress, and community leverage. 5) The close, EXECUTED same-day (section 6): trigger evals shipped and measured (substrate rate 0.914, 130 orphans found), retrieval benchmarked (hit@5 0.76 overall), gateway upgraded and verified, librarian live (51 stale + 11 dupes + 9 dead paths found), and the G-Brain bake-off ran: Lucface 0.90 vs gbrain 0.80 on identical goldens, dead heat 1.00 each on the fully-covered slice. 6) The ghost race, same day: D7 evals 4 → 7 (substrate tier shipped), D4 retrieval 8 → 9 (receipts landed, co-lead with G-Brain), D6 self-improvement 7 → 8 (librarian live). Leads or co-leads went 7 of 12 this morning to 8 of 12 tonight. The map Section 2 is the visual deck: nine interactive charts, three logic trees, and the timed day, with click-to-drill receipts. Section 3 is the 12-dimension scorecard, each dimension racing the GHOST (this morning, translucent) against NOW (tonight, solid), plus the ghost-race and audit-time tables. Section 4 is the G-Brain delta. Section 5 the ranked moves, section 6 the executed receipts, section 7 field notes and sources.
recon: 3 quarantined research agents, live-verified 2026-07-22local ground truth: capability inventory regenerated same-day, live ssh receipts

0 · What is this

On 2026-07-22, Lucas Cooper-Bey (San Diego) had his personal AI system audit itself against the three biggest open agentic-infrastructure projects in the world: Garry Tan's G-Brain, Nous Research's Hermes Agent, and OpenClaw. Then it spent the same day closing its biggest measured gap, and re-graded itself that night. This page is the record: identity cards, an interactive visual deck, a 12-dimension scorecard with receipts, and the before/after "ghost race."

An Intelligent OS is not one app. It is the whole operating layer a person builds around their AI tooling: memory with provenance, retrieval, a skill library, agents, scheduled autonomy, eval harnesses, security rails, and business integration. "Lucface IOS" is the author's instance of that idea.

Reading notes. The page is written in second person: "you" means the operator, because the system wrote it for him. GHOST is the system as it stood that morning; NOW is the same system that night. Scores are 0-10 per dimension, synthesized from receipts. N/A is never plotted as zero. Rival receipts come from their public repos, release ledgers, and published security research, date-stamped 2026-07-22 (stars and versions drift fast, the shapes drift slowly).

Honesty. This is a self-audit, with the incentives that implies. The counterweights: every score carries a receipt, the retrieval claim was settled by installing a real G-Brain and racing it on identical golden questions the same day, and the losing dimensions are printed with the same size font as the winning ones. Not affiliated with, or endorsed by, any of the three projects. → run this exact audit on your own system (section 8)

1 · The five runners (four systems, plus your ghost)

Lucface IOS · THE GHOST
as audited this morning, pre-execution
147 skills · 35 agents · 118 commands
95 scheduled jobs (64 launchd + 31 cron)
65 hook commands · 316 memory files
61 Qdrant collections · wiki 271
zero benchmarks · H-042 open · gateway label lying
G-Brain (gbrain)
Garry Tan (YC CEO), MIT, launched 2026-04-05
26,762★ · 3,891 forks · 42-50 contributors
Postgres+pgvector / PGLite · hybrid RRF + graph
eval suite: routing-eval · skillopt · BrainBench
Tan's brain: 17.9K → 220K pages in 3.5 months
near-daily releases → v0.42.64.0 (07-20)
OpenClaw
OpenClaw Foundation 501(c)(3), steipete at OpenAI
384K★ · 80.6K forks · ~2,737 contributors
ClawHub registry: 52,652 skills (Jun 2026)
iOS + Watch + Voice Wake · cloud workers
monthly CVE cadence · 42,665 exposed instances
v2026.7.2-beta.3 (07-18) · yours: current July stable (upgraded this session from a stale June build)
Hermes Agent
Nous Research, MIT, $1.5B valuation talks
218,903★ · 41,455 forks · ~2,156 contributors
~20 chat channels · full local voice (5 STT/10 TTS)
self-writing skills + autonomous Curator
FTS5 memory, zero provenance (verified 06-13)
v0.19.0 "Quicksilver" (07-20)
Lucface IOS · NOW
same system, tonight, after the /go wave + pp2p
3 eval harnesses (v1.4.x) · 98 scheduled jobs
dispatch 0.914 · top1 0.529 · 130 orphans censused
retrieval 0.76 weekly-tracked · beat gbrain 0.90-0.80
gateway current July stable · librarian 51/11/9 nightly
pp2p: 3 waves · 3 Codex passes · 4/4 mutations caught
The layer picture, one glance · ghost (░) vs NOW (█)
                INGRESS   HARNESS   MEMORY    EVALS    SECURITY  BUSINESS
                & voice   & skills  & truth   & meas.  arch.     & output
 OpenClaw       ████████  ██████    ███       ██       ██        ███
 Hermes Agent   ████████  ████████  ███       ███      ████      ████
 G-Brain        (host's)  █████     ███████   ████████ █████     ██
 LUCFACE ghost  ░░░░      ░░░░░░░░  ░░░░░░░░  ░░░      ░░░░░░░░  ░░░░░░░░   this morning
 LUCFACE NOW    ████      ████████  ████████  ██████   ████████  ████████   tonight
                   ▲                             ▲
             your thinnest wall           the wall you BUILT today
             (voice · deferred M6,        (evals ░░░ → ██████ : substrate
              plugins already loaded)      0.914 + retrieval 0.76 + census;
                                           the --llm judge = last two blocks)
 Each public project is tall in 1-2 columns. This morning you were
 tall in 4. Tonight you are tall in 5, and the short one is chosen.

2 · The visual deck

Nine charts, three logic trees, one timed day. Every chart answers one question and states its honest read under it; hover anything for exact values; click any heatmap cell to drill into its receipt. Charts are self-contained in this file (Apache ECharts inlined, zero external requests); every number also lives in section 3's bars and tables.

V1 · the shape of each runner
Who is built like what, across all 12 dimensions?
Honest read: the rivals spike in one or two places; both of your outlines are wide. The gold G-Brain line breaks at D1 because ingress is N/A for a memory layer, and N/A is never plotted as zero. Click legend names to isolate runners, e.g. GHOST vs NOW alone shows the day's work.
V2 · every score, one grid
Where exactly does each runner stand? Click any cell for its receipt.
Drill
Click a cell above. You get the score, the receipt behind it, and a jump link to the full scorecard row.
V3 · the delta tornado
Where do you lead, and by how much?
Honest read: vs G-Brain you lead 7 dimensions, tie 2, trail 2 (D7 evals by 2, D12 ecosystem by 5), with D1 N/A. Vs the whole field at once, your worst number is the one you chose: D12, minus 8 against a 384K-star foundation.
V4 · what one day moved
Ghost this morning, NOW tonight: which dimensions actually moved?
Honest read: three climbed (D7 +3, D4 +1, D6 +1), nine held, zero regressed. The two flat lines at the bottom, D1 voice and D12 ecosystem, are flat by decision, not neglect.
V5 · head-to-head vs a real G-Brain
Same golden questions, same day, your stack vs a live gbrain install: who retrieves better?
Honest read: dead heat 1.00 on the slice both fully ingested. Your 0.90 vs 0.80 overall win is real but partly corpus coverage: the 400-file cap left 3 of 4 blog targets out of gbrain's index. Both stacks share the same single true miss (Employee #1 at Reddit). Receipts: §6.
V6 · retrieval by collection
Where is your own hit@5 strong, and where is it weak?
Honest read: perfect on wiki, session memory, and the YC library; the two weak collections are diagnosed data-shape problems (prose-set and title-only points), not harness bugs. Overall 0.76, tracked by the weekly retrieval-bench job.
V7 · the dispatch meters
When you type a prompt, does the right capability surface? Two different questions, two very different numbers.
signal present · the right capability is IN the hook's candidate set (64/70 fixtures)
argmax right · the hook's single surfaced pick is the right one (37/70); faint arc = its 0.914 ceiling
Honest read: the 38.6-point gap between the two gauges is the routing story: the signal is nearly always there, the emission policy picks wrong half the time. That is exactly what the specced --llm judge (v2) attacks. By fixture source: mined nudge-map 0.957, capability matrix 0.946, hand-written 0.700. The 6 misses are named in the ledger: legal-entity, deep-extract, council, cold-email, c4-architect, security-lead.
V8 · the orphan census
Of 299 registered capabilities, how many can the dispatch layer actually reach?
Honest read: 130 orphans sounds bad until you see the split: 109 are command aliases (wtd, sv, tabs-*), cheap fixture debt. The 15 skills and 6 agents are the real routing debt, now NAMED in the ledger instead of unknown.
V9 · scale vs dimensions led
How does a zero-star, one-person system stand against foundations and unicorns?
Honest read: log scale on stars. Every rival, 26K to 384K stars, leads exactly 2 of 12 dimensions. You lead or co-lead 8, up from the ghost's 7, because you play a full-stack institution game they structurally do not.

The logic trees

T1 · "why am I still lagging in anything?"
T2 · how a prompt routes through IOS, with tonight's measurements attached
T3 · which runner wins which job (the chooser)

The day, timed (real ledger timestamps where they exist)

3 · The scorecard

Twelve dimensions, 0-10, scored from the receipts in the table below each group. N/A means the dimension is out of the project's scope (a layer is not an assistant), never scored as zero. Scores are this audit's synthesis; every claim traces to a source.

Lucface GHOST (this morning) Lucface NOW (tonight) G-Brain OpenClaw Hermes
D1Ingress breadth and voice▼ you trail

Channels a human can reach the system through, and whether it can talk and listen live.

Luc ghost
5 terminal-first + one messaging gateway, no live voice loop
Luc NOW
5 unchanged by decision (M6); talk-voice + phone-control plugins now loaded on the gateway, unwired
G-Brain
N/A layer, ingress belongs to the host agent
OpenClaw
10 10+ channels, iOS + Watch, Voice Wake
Hermes
9 ~20 channels, PTT + live Discord VC, 5 STT / 10 TTS
D2Autonomy and scheduling▲ you lead

Does useful work happen without a human in the chair, and does dead work get loud.

Luc ghost
9 95 jobs, sentinels, heartbeats, pre-4am brief fleet
Luc NOW
9 98 jobs: +dispatch-eval (Mon 06:20), +retrieval-bench (Mon 06:35), +librarian (nightly 03:40)
G-Brain
8 Minions durable queue, 66 crons in Tan's instance, dream cycle
OpenClaw
8 cron ledger, Automations, cloud workers
Hermes
7 daemon + cron + durable delivery ledger
D3Memory truth and provenance▲ you lead

Can the system prove where a fact came from, and does history survive updates.

Luc ghost
9 append-only spine, originSessionId, graded evidence notes
Luc NOW
9 plus sha-bound eval ledgers (fixture/golden/input hashes, params, versions per row); RFC-3161 still unshipped
G-Brain
7 citations + gap analysis + git history; librarian still aspirational
OpenClaw
4 markdown+git, but untrusted data flattened into prompts (Imperva)
Hermes
3 zero crypto provenance across 11 tiers; update_fact overwrites
D4Retrieval power, measured▲ ghost 8 → NOW 9, co-lead, proven

Hybrid search quality and whether anyone has actually measured it.

Luc ghost
8 61 collections, hybrid twins, /recall federation; zero benchmarks at audit time
Luc NOW
9 receipts live: hit@5 0.76 weekly-tracked + 0.90 vs gbrain's 0.80 head-to-head; co-leads the dimension (§6)
G-Brain
9 HNSW+BM25+RRF+graph; BrainBench P@5 49.1 / R@5 97.9; LongMemEval 97.6 claim
OpenClaw
6 FTS5 + sqlite-vec hybrid (70/30), local embed fallback chain
Hermes
5 FTS5 default, pluggable vector providers
D5Skill library quality × breadth▲ you co-lead

How much packaged capability exists and how it is kept trustworthy.

Luc ghost
9 147 curated + definition-gate + daily inventory regen
Luc NOW
9 unchanged, but the 130-orphan census now NAMES the dispatch debt instead of guessing at it
G-Brain
6 34-60 curated, RESOLVER.md routing, skillify 10-item audit
OpenClaw
8 ClawHub 52,652 skills; ClawHavoc found 1,184 malicious across 12 publishers, 247,693 installs
Hermes
9 89 built-in + Skills Hub meta-search; install-time security scan
D6Self-improvement loop▲ ghost 7 → NOW 8; still trails Hermes 9, gap halved

Does the system get better on its own, and is that safe.

Luc ghost
7 feedback memory + nightly capability-pulse + resolve-radar + audits; human-gated by choice
Luc NOW
8 librarian nightly consolidation live (the dream-cycle analog: 51 stale + 11 dupes + 9 dead found day one), still propose-only by design
G-Brain
8 nightly dream cycle + skillopt (SKILL.md as trainable parameter)
OpenClaw
8 Skill Workshop mines own sessions, approval-gated
Hermes
9 self-writes skills + GEPA/DSPy evolution + autonomous Curator
D7Evals and dispatch measurement▲ ghost 4 → NOW 7; was THE delta, gap now 2 not 5

Is "does the right capability fire, and is the output good" measured, or vibes.

Luc ghost
4 output evals real (mutation gate, golden sets, agent-eval); dispatch evals absent, H-042 open
Luc NOW
7 substrate evals live weekly (0.914 + top1 0.529 + 130-orphan census, mutation-tested, 3 Codex passes); --llm judge still v2 spec, so not a 9 yet
G-Brain
9 routing-eval + check-resolvable + cross-modal 3×3 gate + BrainBench regression suite
OpenClaw
3 skill security scans; no eval framework found
Hermes
4 Kanban ops checks; community reports unreliable self-evaluation
D8Multi-agent orchestration▲ you lead

Fan-out, pipelines, cross-model adjudication, quarantine tiers.

Luc ghost
9 Workflow pipelines, 5-model council w/ blind judge, fuse, architect lanes, 35 agents
Luc NOW
9 unchanged; tonight's proof: 3 recon + 4 builder + 4 critic + 1 fixer agents plus 3 cross-family Codex passes ran this very audit
G-Brain
6 Minions control plane + durable agents; execution engine out of scope by design
OpenClaw
8 paired-node CLI agents, cloud workers, Workboard
Hermes
8 durable Kanban board + live subagent view
D9Security architecture▲ you lead

Injection defense, blast-radius control, secrets, supply chain.

Luc ghost
9 quarantine-by-architecture, leak guard, /scan-tool, block-destructive fired live twice today; one internal hygiene item tracked
Luc NOW
9 gateway current + Node-proof unit + supply-chain scan exercised for real on gbrain; one internal hygiene item still open, so still 9 not 10
G-Brain
6 query sanitizer, SSRF blocks, OAuth 2.1; shell-job fs reads unsandboxed, accepted risk
OpenClaw
3 monthly CVEs, Claw Chain, Zenity prompt-injection backdoor, 42,665 exposed instances, ClawHub malware
Hermes
5 9 CVE-class findings May; Smart Approvals + vault secrets since; self-written skills remain a poisoning surface
D10Compute sovereignty▲ you lead

Local models, own hardware, no single vendor able to turn you off.

Luc ghost
9 Mac + a local GPU box, meshed; Ollama, Qdrant, ComfyUI, whisper, Kokoro, TRELLIS local; an always-on fleet on the second box
Luc NOW
9 unchanged; the box-to-box ACL quirk got root-caused (a deliberate tightening) and every fallback proved itself live
G-Brain
2 hosted APIs only; no local-model story, recon-confirmed
OpenClaw
8 nodes concept, Ollama/LM Studio, NVIDIA NemoClaw on-prem
Hermes
8 Ollama/vLLM/llama.cpp + fully local voice stack
D11Institution and business integration▲ you lead, nobody close

Finance, pricing gates, CRM, legal, client delivery, approval-gated publishing.

Luc ghost
10 S-Corp controllers, cap-table agent, pricing/roadmap gates, CRM, legal-entity, voice editors, syndicate w/ approval
Luc NOW
10 unchanged; still the game nobody else is playing
G-Brain
3 calendar/email read recipes; payments + CRM confirmed absent
OpenClaw
4 skill-mediated Stripe/CRM via Composio; real SMB lead-gen adoption
Hermes
5 native Stripe + Composio reach
D12Ecosystem and momentum▼ structural, by design

Who else improves the system while you sleep.

Luc ghost
2 n=1 + your own agents; spine-pattern and blind-council OSS'd; private moat by choice
Luc NOW
2 unchanged by design; M7 kit exists if the benchmark story ever goes public
G-Brain
7 26.8K★ in 3.5 months, YC megaphone, Hyper (YC S26) building on it, his agents commit as garrytan-agents
OpenClaw
10 384K★, Foundation, OpenAI/NVIDIA/Microsoft/Tencent sponsors
Hermes
9 218.9K★, 2,156 contributors, $1.5B talks, passed OpenClaw on OpenRouter tokens
The ghost race — this morning vs tonight, same system
DimensionGhost (am)NOW (pm)ΔWhat moved it
D4 Retrieval, measured89▲ +1hit@5 0.76 weekly-tracked + beat a fresh gbrain 0.90 to 0.80 on identical goldens; co-leads with G-Brain
D6 Self-improvement78▲ +1librarian nightly consolidation live: 51 stale + 11 dupes + 9 dead paths proposed on day one
D7 Evals + dispatch47▲ +3substrate evals + top1 emission metric + 130-orphan census, mutation-tested through 3 Codex passes, weekly job; --llm judge still spec
D2 Autonomy99=+3 scheduled jobs (98 total), same guarantees
D3 Provenance99=sha-bound eval ledger rows added; RFC-3161 still the missing point
D9 Security99=gateway current + scan gate exercised; one internal hygiene item open
D1 / D5 / D8 / D10 / D11 / D125/9/9/9/10/25/9/9/9/10/2=unchanged (D1 and D12 by explicit decision)
Leads or co-leads7 of 128 of 12D4 joins as co-lead; D7 and D6 remain the two live gaps, both narrowed
Score table — audit-time snapshot (the ghost's receipts)
DimensionLucfaceG-BrainOpenClawHermesVerdict
D1 Ingress + voice5N/A109Real gap: no live voice loop, thin mobile
D2 Autonomy + scheduling9887You lead on depth and deadness-detection
D3 Truth + provenance9743Your signature strength, field-wide
D4 Retrieval, measured8965Par on architecture; receipts landed same-day (0.76 overall; 0.90 vs his 0.80 head-to-head, §6)
D5 Skill library9689Co-lead: their breadth, your curation
D6 Self-improvement7889They automate consolidation; you gate it
D7 Evals + dispatch4934Was THE delta; closed same-day at substrate tier (§6), --llm judge = v2
D8 Orchestration9688Council + Workflow + lanes lead
D9 Security architecture9635You chose architecture over speed; the field proves why
D10 Compute sovereignty9288You + local stack lead; G-Brain is cloud-only
D11 Institution integration10345Nobody else is playing this game
D12 Ecosystem27109Structural: they have thousands, you have you

4 · The G-Brain delta, precisely

The question that started this audit. G-Brain is a memory + eval layer, not an assistant, so the honest comparison is on the five axes where you both play.

Ahead · where your system beats G-Brain today
AxisYour receiptHis state
Provenance disciplineAppend-only spine (house standard), originSessionId on nodes, evidence notes that grade claims (your 07-20 note graded HIS keynote line by line)Citations + gap analysis shipped; "librarian" and contradiction-checking still talk-track, per recon
Curation practiceMEMORY-ARCHIVE, ground-truths audits, keeper, shelf-review, dedupe-on-write memory rulesPreaches the hygiene triad on stage; dream cycle covers part of it
Corpus breadth61 collections spanning YC (~28.6K pts), Every (~22K), failure-intel DB (56K companies), sessions, wiki, personalOne brain repo, 220K pages, single-domain (his work life)
SovereigntyLocal embeddings, local Qdrant, local whisper/TTS/gen on the GPU boxHosted APIs only, ~$100-150/mo, cloud-dependent
Everything downstreamOrchestration, security airlock, business controllers, output railsOut of scope for G-Brain by design
Par · same idea, different maturity
AxisYouHim
Hybrid retrievalDense + sparse twins (YC corpus hybrid-twinned), RRF-style merge via /recall doorsHNSW + BM25 + RRF + source-tier boost, one engine, two backends (PGLite / pgvector)
Knowledge graphWiki backlinks + [[memory links]] + graphify on demandSelf-wiring typed graph, zero LLM calls, +31.4pt P@5 attributable and MEASURED
Nightly maintenancePre-4am brief fleet, capability-pulse, session sweepsDream cycle: entity sweep, citation repair, consolidate, resync
Behind · where G-Brain is ahead of you today
AxisHis receiptYour stateThe move
Trigger / dispatch evalsgbrain routing-eval (intent→skill accuracy vs fixtures) + check-resolvable (MECE/DRY/orphan audit of the skill tree)Was: H-042 open, ~90 capabilities measured going unremembered. Now: substrate evals SHIPPED same-day, 0.9143 over 70 fixtures, 130 orphans censused, weekly job live (§6)M1: DONE
Retrieval benchmarksBrainBench P@5 49.1 / R@5 97.9; LongMemEval 97.6% recall claim; regression-tracked per releaseWas: zero benchmark numbers. Now: hit@5 0.76 across 6 collections, weekly job live, plus the 0.90-vs-0.80 head-to-head win (§6)M2+M3: DONE
Skill optimizationskillopt treats SKILL.md as a trainable parameter, keeps only measured improvements (~$1, ~20 min)skill-creator + definition-gate lint structure, nothing measures skill OUTPUT quality over versionsM2/M3
Community leverage42-50 contributors incl. his own agents (garrytan-agents) opening PRs; a funded YC company building on itn=1 plus your agents; deliberate, but it means every improvement costs your tokensM7, optional

5 · The moves, ranked (all executed same-day; results in §6)

#MoveWhy nowSize
M1Ship dispatch evals (close H-042): build a fixtures file of task-utterance → expected-skill pairs (seed it from capability-nudge-map's mined keyword fingerprints), run it as a tier-1 deterministic check on a cron, report weekly hit-rate. Steal the shape of gbrain routing-eval + check-resolvable directly, MIT license.It is the single axis where the primary rival is genuinely ahead, your own audit already tracks it, and the fixture data already exists on diskS-M
M2Benchmark /recall, BrainBench-style: 20-40 golden questions across the 5 doors with known-correct sources, measure P@5/R@5, track per month. Reuse gbrain-evals' harness shape; your fuse-eval dir is prior art in-house."Par on architecture, zero receipts" is exactly the position your stack-reality rule exists to preventM
M3Run the G-Brain bake-off instead of guessing: point a PGLite G-Brain at a mirror of the business-brain markdown (~1 evening), run the same 20 golden questions against it and /recall, keep whichever wins per query class, steal the losers' tricks (source-tier boost, session demote, graph adjacency boost).TOPIC-BANK already says "evaluate as prior art before building in this lane"; now it costs one evening and answers M2's calibration question for freeM
M4Upgrade the OpenClaw gateway to current stable, then re-harden (MCP session isolation, provenance-gated skill installs shipped upstream in July as direct CVE responses). It is your only internet-facing always-on agent surface; the pre-upgrade build (a stale June build behind a wrong unit label) already had TaskFlow but predated July's Skill Workshop, cloud workers, and isolation work.Monthly CVE cadence upstream + 42,665 exposed instances found in the wild; staleness is exposureS
M5Stand up your dream-cycle analog for backlogs: a nightly librarian pass with a pruning quota over INBOX (237 untriaged) and RUNBOOK (786 open), consolidating and closing, not just appending. The sentinels capture; nothing yet prunes.Garry's "garbage dump with great search" failure mode is measurable in your own queues right nowM
M6Decide the voice-ingress play deliberately: Hermes ships full local voice (faster-whisper + local TTS both free); your phone-path design from 06-02 is archived. Either schedule it or write one line in ETHOS saying terminal-first is the choice.D1 is your widest visible gap vs the field; an undecided gap costs attention every auditdecision
M7Optional flywheel play: publish the M2/M3 benchmark as SDAIS content ("I benchmarked my second brain against the YC CEO's"), with numbers and the harness repo.It is the rare post where the receipts already exist and the audience is the studio's exact ICPS after M2/M3

6 · Executed the same day (the receipts)

/go execution wave · 2026-07-22 · all verified live, ledgers append-only
MoveResultWhere it lives
M1 dispatch evalsSHIPPED + RUN: substrate rate 0.9143 (64/70 fixtures), 130 orphaned capabilities censused; 6 misses diagnosed; weekly Mon 06:20dispatch-evals harness · weekly job (private repo)
M2 retrieval benchSHIPPED + RUN: hit@5 0.76 overall (llm-wiki 1.00, session-memory 1.00, YC library 1.00, YC blog 0.75, every-corpus 0.40, failure-intel 0.33); weekly Mon 06:35retrieval-bench harness · weekly job (private repo)
M3 G-Brain bake-offRAN: supply-chain scan PASS, 400 public YC pages indexed on PGLite; gbrain 0.80 vs Lucface 0.90 on identical goldens; 1.00 vs 1.00 dead heat on the fully-covered slice; 3 steal-worthy tricks documentedbake-off + supply-chain scan notes (private repo)
M4 gatewayDONE: stale June build → current July stable, service unit made Node-bump-proof, messaging channel verified up; config backup takengateway box · systemd user unit · config backup
M5 librarianLIVE: propose-only nightly 03:40; first pass found 51 stale, 11 duplicate pairs, 9 dead paths across INBOX (237 open) + RUNBOOK (786 open)librarian harness · nightly job (private repo)
M6 voice ingressDECIDED: deferred behind eval work, revisit post-SDAIS-launch; the upgraded gateway's new talk-voice + phone-control plugins are the cheapest path when revisiteddecision recorded to memory
M7 content kitBUILT: receipts-only writing kit (never ghostwritten), voice-score before postingwriting-kit dir (private repo)

Scorecard note: section 3 now shows both runners per dimension — the GHOST (audit-time, translucent) and NOW (post-execution, solid). Re-graded NOW scores: D7 4→7, D4 8→9 (co-lead), D6 7→8; leads went 7 of 12 → 8 of 12 in one day. After the pp2p loop the three harnesses sit at v1.4.x: 4/4 mutations caught, 9/9 failure scenarios produce error rows instead of clean numerics, one final Codex-proven census bug fixed and live-verified. Infra note: the second box's Ollama is healthy; box-to-box embedding is closed by deliberate ACL design, so benches use the Mac-local fallback deliberately.

7 · Field notes

Security: the field is proving your architecture right
 OpenClaw   CVE-2026-25253 (RCE, 8.8) · Claw Chain 4-flaw · Zenity
            prompt-injection persistent backdoor · ClawHavoc: 1,184
            malicious skills / 247,693 installs · 42,665 exposed
 Hermes     9 CVE-class findings in 4 days (May) · self-written
            skills auto-loaded = memory-poisoning surface (Critical)
 G-Brain    smallest surface, best behavior of the three; still
            ships "shell jobs read any file" as an accepted risk
 Lucface    quarantine-by-architecture BEFORE the field's incidents:
            untrusted content cannot issue tool calls; leak guard on
            push; supply-chain gate on install; destructive-command
            hook (fired twice during this audit, correctly)
 Lesson: the market optimized for growth and is now buying back
 security. You paid for security first. Do not trade that away
 for channel breadth.
On the 350K pages figure

Your on-disk evidence note records ~350K pages from the Lightcone captions; live recon could verify no figure above 220K (the 07-17 keynote, his highest public number; README says 146,646). Trajectory is the real signal either way: 17,888 pages on launch day to 220K in 3.5 months. Cite 220K.

Recon provenance and injection transparency

Three quarantined research-proxy agents ran the web recon; six fetched pages tripped the injection sanitizer, all six triaged as false positives (site chrome, docs boilerplate, one page with 41 zero-width characters noted), zero instructions followed from fetched content. Hard numbers came from GitHub's REST API, not scraped HTML. Raw API responses parked in this session's scratchpad for re-verification.

Sources

Live: github.com/garrytan/gbrain (+ gbrain-evals, gstack) · github.com/NousResearch/hermes-agent (API + release ledger) · github.com/openclaw/openclaw/releases · openclaw Foundation coverage (thenewstack, computerworld, CNBC) · TechCrunch 07-13 (Nous valuation) · Cloud Security Alliance note (Hermes CVEs) · thehackernews (OpenClaw CVEs) · Orply + Tokenless (AI Engineer World's Fair closing keynote, 07-17) · vectorize.io (GBrain architecture reviews). Local: private evidence notes, corpora, ledgers, and the same-day capability inventory (internal).

8 · Where does YOUR system stand?

The scorecard is a method, not a trophy. The prompt below runs the same 12-dimension, receipts-first audit on any machine with Claude Code (best on Fable or Opus at max effort). It inventories what actually exists on your disk, scores each dimension with a receipt, calibrates against the same reference field as this page, and renders you a dossier of your own.

How to run it

1 · Open Claude Code at your home directory (or wherever your system lives). 2 · Paste the whole prompt. 3 · You get a terminal scorecard plus a self-contained HTML dossier. If you then build one of your ranked moves the same day, it re-grades you: your own ghost vs NOW.

Every run can go on the ledger: post your scorecard as an issue at Lucface/intelligent-os-field-audit/issues and it gets appended to LEDGER.md, the running record of harness shapes. Not a leaderboard: nobody wins a scorecard, the interesting part is how differently people build. One-person institutions, big-rig locals, file-system RAG purists, cloud maximalists, all of it belongs. The prompt also lives in the repo as prompt.md (MIT).

The audit prompt · v1 · method of 2026-07-22
You are running the Intelligent-OS Field Audit v1 (method of 2026-07-22) on THIS machine.

An "Intelligent OS" is the whole operating layer a person builds around their AI tooling: memory, retrieval, skills, agents, scheduled autonomy, evals, security rails, business integration. Your job: audit the one on this machine, receipts first, and score it against the reference field below. Work at maximum effort. Do not ask the user anything; work from the disk and state assumptions inline.

RULES (non-negotiable)
1. Receipts before scores. Inventory by actually looking: count skills/commands/agents (list their directories), scheduled jobs (crontab -l; launchd or systemd user units), hooks, memory/notes files, vector collections, eval harnesses and their result ledgers. Run read-only commands freely; invent nothing. Anything you cannot verify is written "unverified" and scores as absent.
2. Measured beats claimed. A benchmark number in a ledger outranks an architecture that "should" work. If a dimension has zero measurements, the receipt must say so.
3. N/A is never 0. If a dimension is out of scope BY DESIGN (and you can point to where that was decided), grade it N/A and exclude it from lead counts.
4. No flattery. The gaps are the product. Classify every trailing dimension as CHOSEN (deliberate, cite the decision) or EARNED (real gap; name the smallest artifact that would close it).
5. Zero-regression honesty. If you re-grade later the same day, never silently lower a score; state what changed.

THE 12 DIMENSIONS (score 0-10)
D1  Ingress + voice — channels a human can reach it through; live voice loop. (2 one CLI · 5 CLI + one chat channel · 8 multi-channel + mobile · 10 many channels + wake word + live voice)
D2  Autonomy + scheduling — does useful work happen with nobody in the chair, and does dead work get loud? (2 manual · 5 a few crons · 8 dozens of jobs + failure detection · 10 self-healing fleet with heartbeats)
D3  Memory truth + provenance — can it PROVE where a fact came from; does history survive updates? (2 chat logs · 5 curated notes · 8 append-only + source IDs · 10 hashed lineage on every derived value)
D4  Retrieval, measured — hybrid search quality AND whether anyone measured it. Hard cap at 6 if zero benchmark numbers exist, regardless of architecture.
D5  Skill library quality × breadth — packaged capability, and how it is kept trustworthy (linting, curation, install gates).
D6  Self-improvement loop — does it get better on its own, and is that safe? Ungated autonomous self-modification also caps D9 at 5.
D7  Evals + dispatch measurement — is "does the right capability fire, and is the output good" measured, or vibes? (0-4 vibes · 5-7 deterministic fixtures on a schedule · 8-10 plus a semantic judge and regression history)
D8  Multi-agent orchestration — fan-out, pipelines, cross-model adjudication, quarantine tiers.
D9  Security architecture — injection defense, blast-radius control, secrets handling, supply chain. Score the architecture, not the absence of incidents.
D10 Compute sovereignty — local models, own hardware, no single vendor able to turn it off.
D11 Institution + business integration — finance, CRM, legal, client delivery, approval-gated publishing.
D12 Ecosystem + momentum — who else improves it while the operator sleeps. (n=1 by design is a legitimate CHOSEN 2.)

REFERENCE FIELD (as sourced 2026-07-22; D1..D12 in order; null = N/A; re-verify anything that matters to you)
{
 "gbrain":      {"scores":[null,8,7,9,6,8,9,6,6,2,3,7],  "note":"Garry Tan's memory/eval layer: hybrid retrieval w/ published benchmarks (BrainBench), routing-eval + resolvability audits, hosted-API only"},
 "openclaw":    {"scores":[10,8,4,6,8,8,3,8,3,8,4,10],   "note":"biggest ecosystem (384K stars, 52K skills) + best ingress; monthly CVE cadence and a poisoned-registry incident"},
 "hermes":      {"scores":[9,7,3,5,9,9,4,8,5,8,5,9],     "note":"~20 channels + full local voice, autonomous skill Curator; zero memory provenance"},
 "lucface_now": {"scores":[5,9,9,9,9,8,7,9,9,9,10,2],    "note":"the author's one-person system, post-execution 2026-07-22: dispatch evals 0.914 weekly, retrieval hit@5 0.76 weekly, beat a live gbrain install 0.90-0.80 on identical goldens; leads/co-leads 8 of 12"}
}

DO, IN ORDER
1. Inventory sweep (read-only). Enumerate and count the primitives above. Note the 3 most surprising findings.
2. Score all 12 dimensions with a one-line receipt each. The receipt is the evidence itself, not a justification.
3. Comparison table: your row against the reference field. Mark lead / co-lead / trail per dimension; count leads (N/A dims excluded).
4. Verdict in three honest sentences. Then classify every trailing dimension CHOSEN vs EARNED.
5. Ranked moves: the 3 smallest artifacts that would most move the EARNED gaps. Steal proven shapes: a routing-fixtures file scored weekly (does the right capability fire), a golden-set retrieval bench (hit@5 over 20+ questions with known answers), a nightly propose-only librarian over your backlogs. Size each S/M/L.
6. Render a single-file dark-theme HTML dossier (inline CSS, zero external requests, robots noindex): a scorecard section with one bar row per runner per dimension and the receipt beside it, the comparison table, the verdict, the ranked moves. Save it and open it.
7. Ghost mechanic (optional, recommended): if the operator executes a move today, re-run scoring and show GHOST (before) vs NOW (after) rows for the changed dimensions.

Print the scorecard table and verdict in the terminal as well, not only the file.
Method + reference dossier: https://lucface.github.io/intelligent-os-field-audit/ — post your scorecard as a GitHub issue there and it joins the ledger (LEDGER.md), the running record of harness shapes.