Autonomous agents · comparisons · 2026-10-03

Agent architectures, drawn

A diagram view of CONSOLIDATED_FINDINGS.md: Instinct, Poke and G-Brain, what each architecture looks like on the evidence, where the live alternatives are, and the design choices you'd actually face. This synthesizes saved evidence only; no product was re-tested.

supported within scope reported · leading (moderate) plausible · unresolved · proposal unknown · unsupported attribution host must supply failure point · irreversible
BoundaryTwo assistants, one subsystemInstinct and Poke are whole assistants. G-Brain is a memory layer that needs a host.
InstinctMost concrete runtimeGit/Markdown, S3 bundles, Firecracker and browser leases are all reported. Who writes memory is still open.
PokeClearest suppliers and extension pointsBrowserbase, Parallel, PlanetScale and MCP/SDK are documented. Its memory internals aren't.
G-BrainMost inspectable memoryIts mechanisms are visible in pinned code. That doesn't show it gives better answers.
OverallNo winnerThere are no matched end-to-end results for quality, speed, reliability or cost.

01Compare the same system boundary

Every complete assistant needs these nine responsibilities. The colour shows how well the evidence covers each system at each layer.

INTEGRATED ASSISTANTS · internals mostly hidden MEMORY SUBSYSTEM · host does the rest Instinct Poke G-Brain core Channels & identity One identity · Workspace · Files Messaging · API · automations Host supplies Coordination Parent → delegated tasks Reusable contexts (historical) Host task queue Durable knowledge Git/Markdown /memory · S3 bundles Account memory · schema unknown Markdown/Git + DB-only state Capture & correction Deferred · writer unresolved Unmeasured Staged writes · provenance Retrieval & context Selected context + lexical recall Supplied context + source lookup Lexical · vector · graph · rerank Execution & tools Firecracker/E2B · browser leases Browserbase · Parallel · MCP · humans Host executor & tools Models Weights & operator unresolved Provider map unverified Host picks reasoning model Authority & secrets Vault · fill · TOTP OAuth · MCP · keys Memory grants Action approval Recovery & isolation Full restore unproven Persistence ≠ restore Needs DB + files + host
Reading it: to compare G-Brain with Instinct or Poke as a whole assistant, you have to include everything in the purple cells. These are differences in evidence, not an operational ranking.
G-Brain coreretained source · pinned commit Markdown / Git files PGLite or PostgreSQL CLI · MCP interfaces Access grants · maintenance jobs v0.60.31.0 · 2d8801b inside GBRAIN-ASSEMBLEDClaude report · documented host recipe Telegram OpenClaw / Hermes Hosted compute Supabase G-Brain core Brain & workspace Scheduled work approvals · recovery · ops: unknown not tested as a turnkey assistant ≠ don'tblend G-RefCodex report · proposal for testing Channels Executor Action records Effect receipts G-Brain core Recovery / replay Cost accounting required controls, stated explicitly none shipped or tested
R01: the recipe names more specific components, while G-Ref states the requirements more explicitly. Neither one picks up the other's capabilities.

02Instinct: the reported runtime

The most concrete personal-memory and runtime design among the two closed products. Every box here comes from investigator reports; none is a discovered call graph.

You Channelsone identity Conversation ownerlogical parent · todo ownership delegates task task task Modelstrained core + 3rd party · operator ? inference Firecracker guestE2B-associated · lifetime unknown CLI → GraphQL bridge runs in Browser profiles · leasesseparate identities Vault · fill · TOTPisolation not proven universal /memoryGit + Markdown reads S3 bundles publishes loads Background writer?identity · cadence unresolved writes ? Profile · summary · recapselected context injected · reported 4,250 + 8,750-token blocks
Context cost (R04): 4,250 + 8,750 is valid arithmetic for the blocks that were reported. Whether they're included every turn, what's actually billed, and the effect of caching are all unmeasured, so this is not a per-turn cost floor.
Who writes durable memory? A · Background-majority task reconciler memory reads transcript writes most leading · moderate B · Foreground + background task reconciler memory writes reconciles close plausible alternative C · Foreground-only, stale publish task no reconciler memory writes now publication /context lags less supported · not excluded What coordinates work? A · Persistent owner + delegated tasks ownerlong-lived task task task moderate B · Restored owner, fresh handlers owner (restored) shared state new handler new handler reload equally credible C · One permanent computer conversation · tasks · memory browser · tools all on one always-on machine less consistent with separated services
What would decide it: a trace running from message to writer to file to bundle to active context. A git author name and one ~23h16m delay can't settle it. The pips show relative fit, not probabilities.

03Poke: suppliers visible, core hidden

Execution suppliers and the extension surface have scoped evidence. The coordinator and personal memory don't.

Messaging channels API entry Automations Coordinatorcurrent orchestration unresolved Reusable task contextshistorical prompts · 2025 PlanetScalecore state — not a memory schema per-task routes · no fixed fallback order API · MCP RecipesSDK · MCP Browserbase — browser Parallel — monitoring Human fulfillment Site-build VM Vercel + Turso · per-artifact DB context: supplied + looked up Account memoryschema · admission unknown Supplied context + summarieshistorical provenance Connected-source lookup Semantic retrieval?plausible · unranked vector store · Git memory · graph: unsupported attributions
R11: the per-artifact Turso database belongs to the site-building path. It says nothing about how assistant memory or workers are isolated.
Current coordination: four candidates A · Coordinator + reusabletask contexts coord. context context context reused across turns moderate · historical B · Central agent + asynctools / workflows agent tool workflow tool launched, results return credible current alt. C · Persistent per-useractor user actorlives per user weak support D · Always-on VM perassistant VM · 24/7one per assistant unsupported Workflow substrate? all fit · none selected durable workflow engine DB-backed jobs actor runtime hybrid PlanetScale stores core state but doesn't determine which engine runs.
The purported internal prompts form one unauthenticated evidence family from September 2025. They aren't a current production snapshot.

04G-Brain: the inspectable memory core

Here the open questions are configuration choices rather than hidden architecture. The write path, read path, stores, and the two ways around the access filter are shown below.

CLI MCP stdiorestricted (remote: true) MCP remote Raw file / DB accessprivileged operator access filter · grants · transport WRITE PATH Explicit writewith provenance Optional capturehooks · extraction Admit → revise →supersederevision checks READ PATH index reads lexical vector graph rerankoptional context pack Markdown / Gitcanonical pages · inspectable PGLite / PostgreSQLDB-only facts · withdrawals · receiptsaccess & job state — indispensable Background maintenancedream · consolidationchanges validity · adds noise Embedding / aux providerremote · egress ≠ reader privacy sendstext bypasses app controls delivered to host model(may hold stale context) Host (not in core) channels · reasoning model · executor & tools · action approval · scheduling · notifications
R06: local stdio doesn't get past the grant filter. Only raw storage access does. Backup: Git alone isn't a complete backup, because the database holds state the files don't.

05Design space: what any architecture has to place

This is a reference composition built from the five lessons that transfer across systems. It's a map of responsibilities and choice points, not a recommendation or a build plan.

Channels & identitymessaging · app · API · notifications Durable task statequeue · leases · receipts · replay Coordinatorowns the conversation, not the state Context assemblerre-read right before acting Replaceable workersno durable state of their own Action gateaction record: recipient · amount · version Memory servicestaged lifecycle · files + DB · grants Execution routes API · MCP Browser (leased ID) Human fulfillment Artifact VM Model & aux providersegress policy per provider External world · irreversible effects Operationsback up DB + files + host assets · meter cost per accepted task · score monitoring obligations enqueue fresh lease propose re-check select approved, versioned effect receipts writes with provenance embed d a b c e 1 2 3 4 5 6
a explicit state ownership b fresh context at the point of action c scoped authority d recoverable task state e complete cost accounting choice point (below)
The edges to watch are the effect receipts going back to task state (recovery), the re-check from the assembler into the gate (freshness), and the gate sitting between workers and the outside world (authority).
1 Canonical knowledge storewhere the truth lives files / Gitdatabase InstinctPokeG-Brain ? memory schema unknown 2 Context deliverypush into the prompt vs pull on demand pushed (injected)pulled (retrieved) InstinctPokeG-Brain 3 Write timingwhen a statement becomes memory foregroundbackground InstinctPokeG-Brain ? unmeasured 4 Worker allocationhow compute is held per-task / pooleddedicated, always on InstinctPokeG-Brain host's choice 5 Placementwhere each component runs localhosted InstinctPokeG-Brain 6 Execution sourcingbuild vs rent capabilities buildbuy (suppliers) InstinctPokeG-Brain host's choice
Reading it: bar width shows uncertainty, and colour shows how strong the evidence is. Git vs SQL, push vs pull and foreground vs background are trade-offs that can coexist. G-Brain's wide bars mean it supports several settings, not that it's undecided. A local database can still send content to remote providers.

06Memory is a lifecycle, not a storage format

Every stage below can fail independently, so a correct store doesn't guarantee a correct outcome.

Instinct example: ~23h16m end-to-end · which stage lagged was never timed Statementreceived Admitted Durablyrecorded Published/ indexed Selected Deliveredto model Used Correctoutcome Explicit write ≠instant delivery Stored correctly, but a runningtask still uses the old prompt 53 / 500 answers wrongdespite complete retrieval Provenance records who,not whether it's true Packing loses detail:53 → 48 of 60 (holdout)
Use this as a diagnostic vocabulary. In real systems the stages can overlap or run out of order. A revision check can stop a stale write, but it can't decide which of two conflicting preferences is true. "At once or never" (R05) and "automatic memory everywhere" both oversimplify.

07Access, approval and disclosure are separate gates

Passing one gate tells you nothing about the next one.

Agent / task Account accessvault · OAuth · keyssaved browser identity Action approvalthis recipient · amount · versionbinding unproved anywhere External effectirreversible Receipt a credential ≠ permission for this action token expiry doesn't undo effects raw storage access bypasses the filter Memory storefiles + DB Reader access filtersource grants · transport Caller Provider egress policya separate control Embedding / aux provider reader privacy≠ egress privacy "private" labels and local storagedon't imply local-only processing
forgetG-Brain scoped withdrawal revokegrant · token cancela job or task erasestored copies Matching facts · overlays Paraphrases · other scopes Git history · backups · derived copies Context already loaded in a task Future access Pending job Committed external effect ✕ ✕ ✕ ✕ ?
Green means the operation acts on that state. ✕ means the evidence shows it doesn't reach it. Amber ? means it isn't established, since backup erasure isn't shown. A committed effect can only be offset by a compensating action.

08Persistence isn't recovery

Durable memory makes recovery possible. No system has shown a complete, correct restart with replay.

Memory DB · facts · receipts · grants Markdown / Git files Host assets · config · keys Task state & queue Pending approvals In-flight external effects persistence:evidence exists full restoreneeds all six— untested The test that matters interrupt between the external effect and its receipt start effect ⚡crash resume again? receipt never recorded memory persists fine; the task can't tell the effect happened → possible duplicate payment, message or order
R07: the persistence of memory says nothing about whether pending approvals or external effects can be recovered. Restoring from files alone isn't the same as a full restore.

09The only numbers, and how to count cost

These are G-Brain's published results, preserved locally and not reproduced here. Instinct and Poke have no matched denominators.

Retrieval recall@5n = 470 · LongMemEval-S no reranker 439 with reranker 449 +10 questions · 95.5% Answer correctnessn = 500 433 correct 53 wrong despite full retrieval 14 other misses Compressed context packetholdout n = 60 full packet 53 compressed 48 −5 · zero gains Background "dream" run29 transcripts useful signals (of 10) 310 junk pages 25 more recall, more noise, more spend Version v0.48.4.0, not the research commits · 30 abstention cases excluded from recall · tuning went beyond the dev slice · recall and correctness use different denominators
Key point: 53 answers were wrong even though retrieval had found all the evidence. Retrieval and the model's use of it fail in different places.
Service spend · one period · one currency foreground memory execution storage/net channels idlereserve human operations segments not to scale · idle is not automatically zero (R08) ÷ accepted original tasks ✓ ✓ ✕ ✓ ✓ ✕ ✓ ✕ ✓ ✕ ✓ counted once in the denominator · ✕ rejected work and the dotted retries behind it still add spend monitoring is scored per obligation and window, correct silence included; poll count is not value Full ownership adds amortized build hardware migration user repair time "G-Brain is free" and "Instinct builds, Poke buys" both fail at this boundary
R10: put every term into a single currency first, then divide the period's cost by the number of accepted tasks. Don't count retries twice, and don't count a subscription that's already included elsewhere.

10Conditional fit, with no forced winner

These are hypotheses about fit, not purchasing advice. Fill in the priority you hold, then read across.

If your priority is…InstinctPokeG-Brain comp.What could reverse it
Integrated delegation, little infra—Task acceptance, correction freshness, approval behaviour, tier, repair time. Nothing chooses between I and P.
A concrete runtime design to study——A current trace showing different internals. Not a reusable implementation.
Extend a managed assistant with your tools——Action scopes, connector lifecycle, key/version compatibility in your tier.
Own and inspect the knowledge lifecycle——Host integration burden, noisy transformations, context-delivery loss, ops cost.
Shared organizational knowledge——Employee identity semantics, privileged operators, egress, tenant isolation.
Fully local / strict data placement——Hardware, model and task feasibility; hidden remote dependencies.
Cheapest / most dependable unattended completionno selection justifiedMatched accepted outcomes, full costs, recovery and action-control evidence.
most relevant supported directionpermits investigation only

11Where the two reports disagree

R01–R15. Click any item for its resolution. Most disagreements are about how far to infer from the same evidence, not about conflicting observations.

scope difference source / confidence tension overstatement · overgeneralization method or accounting correction
R01G-Brain baseline: recipe vs G-Ref

Keep both, labelled separately. Neither is proof of a tested full assistant, so don't attribute G-Ref's guarantees to the recipe.

R02Instinct memory writer

There's moderate support for deferred writing but low certainty about who the writer is. A git author name can't settle it.

R03Instinct model ownership

Self-arranged compute is a candidate, not the default. Vendor-hosted customization can't be ranked materially lower.

R0413k-token "cost floor"

The arithmetic holds for the reported blocks. Fixed inclusion, billing, caching and dominance in overall cost are all unmeasured, so drop the cost-floor conclusion.

R05G-Brain "at once or never"

Use the staged account instead. Explicit writes can lag delivery, and hooks or background work can capture things later. There's no measured guarantee on freshness.

R06stdio grouped with raw access

Local stdio is restricted (remote: true, world-only). Raw storage access is a different kind of authority.

R07Persistence = recovery

Durable state supports recoverability, but restart and replay correctness are untested. You can't infer that approvals or effects would be recovered.

R08Idle cost ≈ zero

Idle cost can still include live containers, reserved compute, databases, browsers, monitors and channels. Zero is a conditional outcome, not an architectural property.

R09API < VM < browser < human

Reject a universal ordering. Task length, fees, parallelism, retries and acceptance rates can flip it.

R10Cost equation

Price everything in one currency, then divide the period's cost by accepted original tasks. Don't double-count retries or bundled subscriptions.

R11Finer stores = simple deletion

Dedicated objects can simplify some operations. They don't prove authorization or erase backups, and they don't make scaling trivial.

R12Patterns as "invariants"

Treat these as useful design principles, not empirical laws. Adoption claims don't validate demand.

R13"Hosted multi-tenant"

Hosted delivery is supported, but the exact tenancy isn't. Keep endpoint requirements separate from where each component runs.

R14Codex scenario numbers

These are evaluation scaffolding: illustrative deadlines, volumes and crossover algebra. They aren't measurements.

R15Source counts as confidence

Judge each claim against the evidence surface it depends on. Counts of publications don't measure reliability.

12Evidence that would redraw these diagrams

None of these has been run or scheduled. Which one comes first depends on the decision you're making.

→ §4G-Brain v31 → v32 diff
testDiff the disputed storage, access, retrieval and scheduling paths.
if failThe retained-code findings stop carrying forward.
→ §2Instinct writer & freshness
testRepeated traces from message to writer to file to bundle to active context.
if failMixed writers may fit better than a daily reconciler.
→ §3Poke coordinator & memory today
testCurrent task/context traces plus recall tests with sources isolated.
if failThe historical hypotheses need revising. Fluent answers don't count as evidence.
→ §6Correction across active tasks
testStart two tasks with A, correct it to B, and inspect both outputs and any later writes.
if failCorrect storage without propagation isn't enough.
→ §7Action authority
testInterleave drafts, scope an approval, edit one draft, cancel another, then check receipts.
if failConversational assent doesn't authorize a specific version.
→ §7Withdrawal vs erasure
testTry the exact fact, a paraphrase, a stale re-import, another scope, and context already loaded.
if failNarrow the promise to what's actually controlled.
→ §8Recovery without duplicates
testInterrupt between the effect and the receipt, resume, and compare a full restore with a files-only restore.
if failMemory persistence isn't task recovery.
→ §7Access vs provider disclosure
testUse a denied-source caller and capture provider payloads in a synthetic setup.
if failAccess filtering and egress policy need fixing separately.
→ §9Does added machinery pay?
testRun the same accepted tasks with maintenance and retrieval stages on and off.
if failSome stages may not earn their cost, or may help one metric while hurting another.