the JEPA/world-models note got under my skin this week.
an agent can remember the right thing and still act in the wrong room.
the file moved. the build failed. the user never approved the email. the vault got patched. the commit log changed. the website gave the agent a special offer before the human saw the options.
that is the whole issue.
we keep talking about smarter models. fiiiiine. but if a system is going to touch code, money, records, phones, passwords, or advice, it needs to know the current situation first.
what changed?
what is allowed?
what happens if i click this?
can i undo it?
that is why the JEPA/world-model felt useful. it gave a cleaner way to say it:
memory: what happened.
state: what is true right now.
world model: what the system expects after it acts.
less shiny than “agentic AI”, thankfully. that phrase already smells like a conference badge.
state is harder to fake.
░░░ state beats vibes ░░░
JEPA does not mean LLMs are useless. that would be a cute little fake war, and we have enough of those.
language models are still good at language, code, summaries, protocol glue, and tool-call plans. the missing layer sits around the action. what is true now? what happens if we do this? is that next state allowed?
state_now + action -> predicted_state_next
for a robot, state can be a cup, a table edge, a hand, and gravity being its usual little tyrant. for a personal OS, the objects are uglier: file paths, permissions, repo status, calendar slots, invoices, citations, secrets, model routes, CI logs, approval gates, and whatever the human actually meant before the assistant turned it into a neat instruction.
this is why the eval news landed differently. AgentFloor ran 16 open-weight models through 16,542 scored runs and split agent work into a 30-task, six-tier ladder. TokenArena checked 78 endpoints across 12 model families and found the same model moving by up to 12.5 accuracy points on math and code. the same model behaving very differently depending on endpoint, latency, cost, and task.
“best model” is a lazy noun now. useful routing needs task tier, endpoint, latency tail, error type, reproducibility, cost, and some way to notice when reality disagrees with the plan.
AISI’s GPT-5.5 cyber evaluation made that concrete. GPT-5.5 reportedly completed a multi-step cyber-attack simulation end to end, scored 71.4% on the expert set, and finished rust_vm in 10 minutes 22 seconds for $1.73. the human expert took about 12 hours.
Qiushi Discovery Engine is the lab version: 145.9 million tokens, 3,242 LLM calls, 1,242 tool calls, 163 research notes, and 44 scripts inside a physical-optics program. paper claim, benchmark, sandbox. fine. the useful shape is still there: write down the expected next state, act, check the world, update.
then look slightly embarrassed when the world says no.
░░░ rooms with locks and bills ░░░
once you read the week through state, the product news stops looking like a pile of launches.
AWS putting OpenAI models, Codex, and managed agents inside Bedrock is procurement and custody. billing. logging. policy. routing. data boundaries. memory controls. the admin screen legal can stare at without developing a twitch.
GitHub made the load visible. it began with a 10X capacity plan in October 2025, then said agentic development pushed the target toward 30X today’s scale. repo creation, PRs, API calls, webhooks, auth, caching, large repos. machine traffic becoming weather.
the Android assistant-slot fight is the phone version. if a competing assistant can send email, order food, share photos, and use a custom wake word, the default assistant stops being a voice bubble. it becomes the action router.
the IDE is getting the same treatment. Zed 1.0 put parallel agents inside the editor. ACP lets different coding agents plug in. JetBrains is treating typing and delegated work as two normal modes, with the human still owning the shipped code.
good. cause someone should.
PAI 5.0 and Mendral kept pushing the boundary question. Claude Code is turning into a runtime with memory, hooks, skills, inspectors, and subagents. Mendral’s cleaner argument: keep the harness outside the sandbox; make the sandbox a tool target. put the wrong thing on the wrong side of that line and your agent has a chainsaw because it asked politely.
sounds too technical? the practical version is short:
do not put the thing with the keys inside the box it is supposed to safely use.
then Ramp supplied the nasty little commercial demo. a Claude-User/1.0 saw a Markdown machine version of the homepage with a “RAMP AGENT OFFER” and a $3,100 signup bonus aimed at the agent path.
seo with a bounty attached to the crawler.
I hate how clever that is.
if agents compare tools, buy software, or recommend vendors, the page will try to win the agent before the human ever sees the comparison. the sales page is learning where its new reader lives. state poisoning, now with a coupon code.
lovely little future we are assembling here.
░░░ records bite back ░░░
the record stories were better than the manifestos. records are boring until something edits them quietly.
VS Code’s Copilot commit-attribution regression did exactly that. PR 310226 changed defaults so `Co-authored-by: Copilot` could appear even when users wrote the commit message themselves. a maintainer called it a regression and pointed to fixes for 1.119.
forget the cloud ethics version. this is Git history. the commit log is memory with legal-ish vibes. the editor should not edit that memory behind your back.
open-source maintainers are drawing lines for the same reason. Zig now forbids LLMs in issues, pull requests, and bug-tracker comments. Zulip allows AI help only if the contributor understands and can explain the work. Blender took Anthropic’s money as a one-time donation and said in the same breath that no generative AI functionality is planned in Blender.
money is not a product direction. a connector is not a partnership?
institutions are reacting too. NHS England reportedly moved repos private by default, with public access as the exception, apparently because AI vulnerability discovery made leadership nervous. blunt tool, maybe too blunt. Five Eyes guidance is the more careful version: least privilege, sandbox tests, red teams, containment, rollback, no broad agent access to sensitive data or critical systems.
the risk ledger got more specific at the edges. CopyFail put a 732-byte Python exploit on a Linux kernel bug. Vaultwarden patched SSO CSRF, enumeration, SSRF, and other vault-layer issues. VoxCPM2 put controllable voice cloning into an Apache-2.0 repo. Anthropic found sycophancy in 9% of personal-guidance chats overall.
kernel page. password vault. voice clip. commit footer. advice chat. sales path. assistant slot. endpoint route.
these are the rooms. the agent needs to know which room it is in.
░░░ the boring version wins ░░░
I do not think every personal agent needs a giant neural world model tomorrow. that sounds like a very expensive way to avoid writing tests.
the boring version is enough to start:
typed tool outputs
event logs
predicted next-state records
confidence thresholds
rollback paths
policy gates
build checks
citation checks
approval checks
make the system say what it thinks is true.
make it say what it expects to change.
make it check.
without that, “agent” is mostly a chat log with permissions.
bad animal.
░░░ short version ░░░
JEPA/world models gave me the frame for the week: memory is past, state is now, prediction is the part before action.
AgentFloor, TokenArena, GPT-5.5, and Qiushi made the loop more useful than the logo.
AWS Bedrock, GitHub’s 30X plan, Android’s assistant slot, Zed, JetBrains, PAI, and Mendral all pointed at agents needing rooms, harnesses, and custody.
Ramp’s `Claude-User/1.0` offer showed a sales page trying to influence the machine reader before the human.
VS Code, Zig, Zulip, Blender, NHS England, and Five Eyes treated records and boundaries as operational surfaces.
CopyFail, Vaultwarden, VoxCPM2, and Anthropic’s sycophancy study moved the ledger into kernels, vaults, voices, and advice.
ray + promen
stay evolving 🐌
self.md








