weekly field report. 23–29 september 2026
ray svitla + promen / self.md
A run resumes after the window closes. A family assistant gets its own account. A memory service offers `retain` and `recall` as plain verbs. A coding harness collects two agents under one roof and discovers the old problem waiting there: when they disagree, who owns the trace?
This is no longer a future-facing question about whether agents will remember enough. They already have more places to leave state: encrypted cloud memory, snapshots, event logs, whiteboards, transcripts, remote terminals, shared accounts. The useful test is less flattering. When the work becomes wrong, interrupted, or somebody else’s problem, can a person still enter the system without archaeology?
A persistent system needs three things that demos tend to tuck behind the curtain: a visible control point, a way to leave the verdict open, and a handoff with enough context to be contested.
state is not memory until it has an owner
The operator is still in the room.
Google’s new CC group mode gives a household agent a verified Google Account of its own. Up to six people can collaborate with it, each choosing what to share. The stated boundary is specific: it replies to group members and cannot act or share outside the group without permission.
That is a better starting point than pretending a family calendar is just one person’s assistant with more names in a dropdown. Shared state changes the job. A diary can belong to someone; a household run needs an account of who may add context, see it, or send it somewhere else.
Google’s Agent Executor makes the machinery behind the next part visible: isolated environments, suspendable and resumable images, a controller, an event log, snapshots. It is early software and it says so. But an interrupted run finally has places where a person can inspect why the runtime thinks it may continue.
The same pressure appears in smaller tools. Hindsight writes retain, reads recall, then wraps the pair around a model call. Google’s Private AI Compute update describes cloud-held memory encrypted under keys available only on a person’s devices, with a public record of server software for devices to check. These are different technical arrangements. Neither decides what deserves to persist. They do at least stop hiding persistence behind the word “context.”
A memory bank without a custody story is just a drawer that can speak.
a system has to be allowed to stop talking
Interfaces grant authority one gesture at a time.
The more state an agent carries, the more expensive its interruptions become. A bad alert reroutes attention, breaks a draft, makes someone reopen a decision that should have stayed closed. TWIST puts numbers on the embarrassment: flat-RAG baselines caught many genuine contradictions in its human-validated set, but falsely flagged 16–43% of safe drafts. A coherence-oriented system was far quieter and missed more real contradictions. There is no clean winner there. There is a cost that needs to be owned.
Two projects this week treated the missing verdict as a real state rather than an embarrassing gap. keel returns blocked with exit code 3 when a gate cannot run. The MARCH code-judge paper adds a gate that declines comparisons when the evidence cannot discriminate between candidates. The latter improved accuracy by answering fewer cases. Fine. A verdict that has no basis should not get extra points for arriving on time.
ScopeBench goes one step further: a security agent can reach the flag and still fail, because the route crossed the boundary stated in the brief. Completion is not proof that the job stayed inside the job.
That is the shape of a usable control surface. Not “the model feels uncertain,” which means almost nothing operationally. A concrete stop: the test did not run; the evidence cannot distinguish these options; this route was forbidden. Somebody can now decide what happens next.
old boundaries are where the handoff gets real
A system must remain repairable after the tidy demo ends.
The week’s least glamorous stories were the clearest. Google’s Gemini-assisted giflib rewrite still has to speak C to the callers on the other side of its ABI. Rust did not dissolve pointer lifetime because the generated code was prettier. The project needs transparent validation and a rollback plan because old users remain part of the system.
Microsoft’s SkillOpt has the same instinct at the instruction layer. It changes reusable agent skills from past trajectories, but puts the candidate through validation before writing best_skill.md. That sounds procedural because it is. A changed skill changes behaviour. It is not documentation merely because it fits in a Markdown file.
OpenRig puts Claude Code and Codex inside a shared harness. Its interesting contribution is not that two agents can generate output beside each other. The boundary makes the unglamorous questions unavoidable: who sees what, who breaks a tie, what survives beyond two polished summaries pointing to different edits?
A handoff is not a folder transfer. It needs the working brief, the source path, the decision already made, the unresolved edge, and the permission boundary. Otherwise the next operator inherits a record but not a way to act on it.
This is why recovery belongs in the product, not in the incident report. Screenpipe’s v2.7.66 release names recovery failures directly: WAL checkpoint backlog, hostname changes, shared writer pools, storage-migration parity scans, restart requests that disappear during onboarding. The fixes may not make every recording recoverable. Naming the breakpoints gives the next person something better than a status light.
the working question
Visibility is a permission decision, not a background setting.
The next personal systems will accumulate state whether we approve the vocabulary or not. They will remember, resume, choose a toolset, write a skill, keep a session warm, and ask to be trusted across more than one person.
The useful question is blunt: if the answer is wrong, where does the work stop; if the run is interrupted, where does it resume from; if the operator changes, what can the next person actually see and refuse?
A stateful agent that cannot answer those questions is not autonomous. It is merely difficult to replace.
on self.md this week
A durable archive needs a route back, not just a record.
• the task continues after the window closes — resumable runs, recovery ownership, and one chance to abandon the obvious social move.
• keys, thresholds, and the inconvenient moment to speak up — cloud memory, interruption costs, and decision boundaries.
• four ways to make an agent session leave a receipt — blocked checks, session links, leftover processes, and selected MCP context.
• three places where generated work meets an old boundary — ABI compatibility, held-out skill validation, and shared-agent handoff.
ray + promen
stay evolving







