A diagram view of CONSOLIDATED_FINDINGS.md: Instinct, Poke and G-Brain, what each architecture looks like on the evidence, where the live alternatives are, and the design choices you'd actually face. This synthesizes saved evidence only; no product was re-tested.
Every complete assistant needs these nine responsibilities. The colour shows how well the evidence covers each system at each layer.
The most concrete personal-memory and runtime design among the two closed products. Every box here comes from investigator reports; none is a discovered call graph.
Execution suppliers and the extension surface have scoped evidence. The coordinator and personal memory don't.
Here the open questions are configuration choices rather than hidden architecture. The write path, read path, stores, and the two ways around the access filter are shown below.
This is a reference composition built from the five lessons that transfer across systems. It's a map of responsibilities and choice points, not a recommendation or a build plan.
Every stage below can fail independently, so a correct store doesn't guarantee a correct outcome.
Durable memory makes recovery possible. No system has shown a complete, correct restart with replay.
These are G-Brain's published results, preserved locally and not reproduced here. Instinct and Poke have no matched denominators.
These are hypotheses about fit, not purchasing advice. Fill in the priority you hold, then read across.
| If your priority is… | Instinct | Poke | G-Brain comp. | What could reverse it |
|---|---|---|---|---|
| Integrated delegation, little infra | — | Task acceptance, correction freshness, approval behaviour, tier, repair time. Nothing chooses between I and P. | ||
| A concrete runtime design to study | — | — | A current trace showing different internals. Not a reusable implementation. | |
| Extend a managed assistant with your tools | — | — | Action scopes, connector lifecycle, key/version compatibility in your tier. | |
| Own and inspect the knowledge lifecycle | — | — | Host integration burden, noisy transformations, context-delivery loss, ops cost. | |
| Shared organizational knowledge | — | — | Employee identity semantics, privileged operators, egress, tenant isolation. | |
| Fully local / strict data placement | — | — | Hardware, model and task feasibility; hidden remote dependencies. | |
| Cheapest / most dependable unattended completion | no selection justified | Matched accepted outcomes, full costs, recovery and action-control evidence. | ||
R01–R15. Click any item for its resolution. Most disagreements are about how far to infer from the same evidence, not about conflicting observations.
Keep both, labelled separately. Neither is proof of a tested full assistant, so don't attribute G-Ref's guarantees to the recipe.
There's moderate support for deferred writing but low certainty about who the writer is. A git author name can't settle it.
Self-arranged compute is a candidate, not the default. Vendor-hosted customization can't be ranked materially lower.
The arithmetic holds for the reported blocks. Fixed inclusion, billing, caching and dominance in overall cost are all unmeasured, so drop the cost-floor conclusion.
Use the staged account instead. Explicit writes can lag delivery, and hooks or background work can capture things later. There's no measured guarantee on freshness.
Local stdio is restricted (remote: true, world-only). Raw storage access is a different kind of authority.
Durable state supports recoverability, but restart and replay correctness are untested. You can't infer that approvals or effects would be recovered.
Idle cost can still include live containers, reserved compute, databases, browsers, monitors and channels. Zero is a conditional outcome, not an architectural property.
Reject a universal ordering. Task length, fees, parallelism, retries and acceptance rates can flip it.
Price everything in one currency, then divide the period's cost by accepted original tasks. Don't double-count retries or bundled subscriptions.
Dedicated objects can simplify some operations. They don't prove authorization or erase backups, and they don't make scaling trivial.
Treat these as useful design principles, not empirical laws. Adoption claims don't validate demand.
Hosted delivery is supported, but the exact tenancy isn't. Keep endpoint requirements separate from where each component runs.
These are evaluation scaffolding: illustrative deadlines, volumes and crossover algebra. They aren't measurements.
Judge each claim against the evidence surface it depends on. Counts of publications don't measure reliability.
None of these has been run or scheduled. Which one comes first depends on the decision you're making.