<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Morphic]]></title><description><![CDATA[Security, systems thinking, and the patterns that hold everything together.]]></description><link>https://morphic.zenone.org</link><image><url>https://substackcdn.com/image/fetch/$s_!Iift!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F66bf5ecd-c67a-4f99-b6b2-335fe455ccf1_1280x1280.png</url><title>Morphic</title><link>https://morphic.zenone.org</link></image><generator>Substack</generator><lastBuildDate>Wed, 12 Aug 2026 01:27:15 GMT</lastBuildDate><atom:link href="https://morphic.zenone.org/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Steve Zenone]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[morphic@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[morphic@substack.com]]></itunes:email><itunes:name><![CDATA[Steve Zenone]]></itunes:name></itunes:owner><itunes:author><![CDATA[Steve Zenone]]></itunes:author><googleplay:owner><![CDATA[morphic@substack.com]]></googleplay:owner><googleplay:email><![CDATA[morphic@substack.com]]></googleplay:email><googleplay:author><![CDATA[Steve Zenone]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[The Runtime Is the Bug]]></title><description><![CDATA[AI security keeps staring at the prompt. Eleven vulnerabilities disclosed across six agent frameworks show the deeper risk is often the code underneath it.]]></description><link>https://morphic.zenone.org/p/ai-agent-runtime-vulnerabilities</link><guid isPermaLink="false">https://morphic.zenone.org/p/ai-agent-runtime-vulnerabilities</guid><dc:creator><![CDATA[Steve Zenone]]></dc:creator><pubDate>Thu, 06 Aug 2026 19:46:27 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!3RxR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5084dd06-f879-4d8f-8165-267a0682ba93_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!3RxR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5084dd06-f879-4d8f-8165-267a0682ba93_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!3RxR!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5084dd06-f879-4d8f-8165-267a0682ba93_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!3RxR!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5084dd06-f879-4d8f-8165-267a0682ba93_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!3RxR!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5084dd06-f879-4d8f-8165-267a0682ba93_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!3RxR!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5084dd06-f879-4d8f-8165-267a0682ba93_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!3RxR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5084dd06-f879-4d8f-8165-267a0682ba93_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5084dd06-f879-4d8f-8165-267a0682ba93_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2512372,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://morphic.substack.com/i/210108277?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5084dd06-f879-4d8f-8165-267a0682ba93_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!3RxR!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5084dd06-f879-4d8f-8165-267a0682ba93_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!3RxR!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5084dd06-f879-4d8f-8165-267a0682ba93_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!3RxR!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5084dd06-f879-4d8f-8165-267a0682ba93_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!3RxR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5084dd06-f879-4d8f-8165-267a0682ba93_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Five identical mechanisms on a workbench. One casing has cracked open, exposing the machinery the interface was hiding.</figcaption></figure></div><p><em>Much of the recent AI security argument has centered on the prompt. The vulnerabilities Check Point brought to Black Hat this week live one layer down, in the plumbing, and most belong to bug classes security teams have known for decades.</em></p><p>The agent doesn&#8217;t always need a dangerous tool. Sometimes reading the wrong thing is enough</p><p></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/subscribe?"><span>Subscribe now</span></a></p><p></p><p>According to <em>The Register&#8217;s</em> account of the briefing, Check Point researchers Yarden Porat and Shahar Tal discussed eleven disclosed vulnerabilities spanning LangChain, LangGraph, CrewAI, AutoGen, Microsoft Agent Framework and Google&#8217;s Agent Development Kit at Black Hat USA on Wednesday, August 5, 2026. The publication described several of the flaws as critical.</p><p>Tal put it plainly: most of what they found did not belong to some new frontier class of AI vulnerability. The list included insecure deserialization, server-side request forgery, path traversal and use-after-free.</p><p>Old bugs. New execution layer &#8230; and that distinction matters</p><p></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/ai-agent-runtime-vulnerabilities?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/ai-agent-runtime-vulnerabilities?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p></p><h2>Where the code actually runs</h2><p>I&#8217;ve spent two pieces this summer writing about authority: <a href="https://morphic.substack.com/p/confused-deputy-ai-agents-delegated-authority">who an agent acts as</a> and <a href="https://morphic.substack.com/p/who-answers-for-the-agent-ai-accountability">who answers for what it does</a>. Both pieces assume the framework holding that authority does its job.</p><p>This is what happens when it doesn&#8217;t.</p><p>Take LangGraph&#8217;s checkpointer. A checkpointer stores workflow state so an agent can resume from an earlier point, inspect its history or recover after an interruption. Check Point&#8217;s research focused on applications that exposed LangGraph&#8217;s <code>get_state_history()</code> function with an attacker-controlled filter while using a vulnerable persistence component. It was not every LangGraph deployment and it did not affect LangChain&#8217;s PostgreSQL-based managed deployment.</p><p>The first flaw in the chain, <a href="https://nvd.nist.gov/vuln/detail/CVE-2025-67644?utm_source=morphic">CVE-2025-67644</a>, was SQL injection in the SQLite checkpointer.</p><p>A user-controlled filter key was interpolated into a SQLite JSON path expression rather than handled safely. With a crafted key, an attacker could alter the query and append a <code>UNION SELECT</code>. The important detail is what that union produced: not a new checkpoint written into the database, but a fabricated row inserted into the query&#8217;s returned result set.</p><p>That fake result could carry attacker-controlled serialized data.</p><p>When the application processed what the database returned, the checkpointer treated the row like saved state and deserialized it. That reached the second vulnerability, <a href="https://nvd.nist.gov/vuln/detail/CVE-2026-28277?utm_source=morphic">CVE-2026-28277</a>, an unsafe msgpack deserialization flaw capable of invoking imported Python callables with attacker-supplied arguments. Chained together under the conditions Check Point described, the two flaws produced remote code execution on the application server.</p><p>The prerequisites matter.</p><p>The affected application had to expose the vulnerable history-filtering path to attacker-controlled input and use the vulnerable SQLite checkpointer. The deserialization flaw was not, by itself, a magic remote shell against every LangGraph installation. Its advisory treated it as a post-exploitation issue because an attacker first needed a way to place or introduce malicious serialized state. In Check Point&#8217;s demonstrated chain, the SQL injection supplied that path.</p><p>The SQLite injection was fixed in <code>langgraph-checkpoint-sqlite</code> 3.0.1. The msgpack issue was fixed in <code>langgraph</code> 1.0.10. Check Point also reported a related injection flaw in the Redis checkpointer, CVE-2026-27022, fixed in <code>langgraph-checkpoint-redis</code> 1.0.2.</p><p>No model jailbreak was required for the demonstrated SQLite-to-deserialization chain.</p><p>That doesn&#8217;t make the model irrelevant to the larger research. Some of the broader findings involved prompt-controlled material crossing into trusted framework behavior. But in this LangGraph case, the exploitable path lived in ordinary application input, query construction and deserialization.</p><p>The runtime was enough.</p><h2>The bugs without public identifiers</h2><p>Two other disclosures are harder to track because, according to Check Point and contemporaneous reporting, they had not received CVE identifiers as of August 6, 2026.</p><p>Google&#8217;s Agent Development Kit included a development-oriented component capable of writing files. Check Point reported an attack path in which generated Python could execute when imported, exposing credentials available to the process. <em>The Register</em> reported that Google initially disputed the finding, later made changes and paid a bounty of $3,133.70. Because I could not locate a primary Google advisory confirming every part of that timeline, those details should remain attributed to the reporting rather than stated as Google&#8217;s public account.</p><p>Microsoft Agent Framework had a different failure mode. Check Point found that one user&#8217;s prompt-injected content could place a malicious payload into checkpoint data. When another user rewound a session and the framework deserialized that state, the payload could execute on the server. Microsoft told <em>The Register</em> that it hardened the framework against the demonstrated path and paid a $10,000 bounty. The company did not issue a CVE because the framework was not generally available when the issue was reported.</p><p>That explanation describes a release state.</p><p>It does not make the technical failure less real. But it does affect how confidently we can describe exposure. A pre-release vulnerability is not evidence of broad production deployment, and it should not be written as though it were.</p><p>The narrower point is enough: without a public advisory or identifier, defenders have less structured information to search for later.</p><h2>Not one team grinding one axe</h2><p>Microsoft&#8217;s own security advisories document related execution risks inside Semantic Kernel.</p><p><a href="https://nvd.nist.gov/vuln/detail/CVE-2026-26030?utm_source=morphic">CVE-2026-26030</a> affected the Python SDK&#8217;s <code>InMemoryVectorStore</code> filtering logic. Microsoft&#8217;s advisory identifies it as a critical remote-code-execution vulnerability and fixes it in <code>semantic-kernel</code> 1.39.4.</p><p><a href="https://nvd.nist.gov/vuln/detail/CVE-2026-25592?utm_source=morphic">CVE-2026-25592</a> affected <code>SessionsPythonPlugin</code> in Semantic Kernel&#8217;s .NET SDK. A file-transfer function exposed to agent function calling could write to an attacker-selected local path unless the application added its own validation. Microsoft fixed the affected .NET component in version 1.71.0 and recommended an allowlist around file-transfer paths as a workaround.</p><p>Those advisories are more careful than the story we sometimes tell around them. One flaw created a remote-code-execution path inside vector-store filtering. The other enabled arbitrary file writes through an agent-callable plugin. Depending on the host and target path, that write could become part of a larger execution chain.</p><p>That is the real pattern.</p><p>Model-controlled or attacker-influenced data reaches an ordinary dangerous sink: <code>eval</code>, deserialization, a filesystem path, a shell or a query builder.</p><p>OWASP already has language for one part of this. LLM05 in the 2025 Top 10, Improper Output Handling, covers systems that pass model output downstream without sufficient validation. OWASP lists SSRF, privilege escalation and remote code execution among the possible backend impacts and specifically warns about model output reaching functions such as <code>exec</code> or <code>eval</code>.</p><p>That category does not explain every framework bug in Check Point&#8217;s research. A SQL injection in a user-controlled filter is still a SQL injection. Unsafe deserialization remains unsafe deserialization.</p><p>AI did not invent these defects.</p><p>It placed them underneath software that can read repositories, handle credentials, retain state and act with someone else&#8217;s authority.</p><h2>The unscored layer</h2><p>Prompt injection still deserves the attention it gets.</p><p>A January 2026 Systematization of Knowledge paper by Narek Maloyan and Dmitry Namiot synthesized 78 studies published between 2021 and 2026, catalogued 42 attack techniques and reported attack-success rates above 85 percent against state-of-the-art defenses when adaptive strategies were used. The paper is an arXiv preprint, not proof that every agent or defense fails at that rate, but it is strong evidence that prompt injection remains unresolved across the systems surveyed.</p><p>What the Check Point disclosures add is a second question.</p><p>Even when we measure whether the model can be manipulated, are we measuring what happens after manipulated content reaches the framework?</p><p>That is the unscored layer: the code the benchmark assumes will safely receive whatever the model emits.</p><p>Sometimes it doesn&#8217;t.</p><p>The disclosure system has a similar blind spot. In April, <em>The Next Web</em> reported research by Aonan Guan against Anthropic&#8217;s Claude Code Security Review, Google&#8217;s Gemini CLI Action and GitHub&#8217;s Copilot Agent. Malicious instructions placed in GitHub-controlled content could be consumed as trusted context and used to expose secrets through the agents&#8217; own workflow output. Anthropic paid $100 and GitHub paid $500. TNW reported that Google also paid a bounty, but described the amount as undisclosed. None of the three findings had received a CVE or public security advisory at the time of that report.</p><p>That last point needs precision.</p><p>A CVE is not the only way security tooling discovers risk. Scanners can use vendor advisories, GitHub Security Advisories, package metadata, custom signatures and other feeds. But CVEs remain one of the main identifiers used to correlate a vulnerability across advisories, dependency tools, asset inventories and remediation systems.</p><p>No identifier does not make a flaw invisible.</p><p>It does make correlation harder.</p><p>So the frameworks get patched. Ordinary bugs, ordinary fixes.</p><p>What still feels missing is the layer meant to notice that the runtime itself has changed underneath the benchmark. The control that asks whether attacker-influenced data reaches a deserializer, query builder, path operation or dynamic evaluator before any model score matters.</p><p>Check Point describes LangGraph as receiving more than 50 million downloads per month. Package downloads are not the same as unique users or deployed systems, but the number still gives the blast radius some shape.</p><p>The argument is not that prompt injection was a distraction.</p><p>It is that prompt injection was never the whole system.</p><p>We kept staring at the sentence the model read.</p><p>The bug was waiting in what read the model.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/ai-agent-runtime-vulnerabilities/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/ai-agent-runtime-vulnerabilities/comments"><span>Leave a comment</span></a></p><div><hr></div><h3>Resources</h3><ul><li><p><a href="https://research.checkpoint.com/2026/from-sqli-to-rce-exploiting-langgraphs-checkpointer/?utm_source=morphic">From SQLi to RCE: Exploiting LangGraph&#8217;s Checkpointer</a> - Check Point Research&#8217;s primary technical disclosure. Covers the SQLite SQL injection, unsafe msgpack deserialization, Redis injection, exploit prerequisites, affected configurations, disclosure timeline and remediation information.</p></li><li><p><a href="https://nvd.nist.gov/vuln/detail/CVE-2025-67644?utm_source=morphic">CVE-2025-67644: LangGraph SQLite Checkpointer SQL Injection</a> - NIST&#8217;s vulnerability record for the SQLite checkpointer injection used in Check Point&#8217;s demonstrated exploit chain.</p></li><li><p><a href="https://github.com/advisories/GHSA-g48c-2wqr-h844?utm_source=morphic">CVE-2026-28277 / GHSA-g48c-2wqr-h844: Unsafe Msgpack Deserialization</a> - GitHub&#8217;s reviewed advisory for the LangGraph checkpoint deserialization flaw. Documents the post-exploitation prerequisite, affected versions, remediation and the absence of known exploitation in the wild.</p></li><li><p><a href="https://nvd.nist.gov/vuln/detail/CVE-2026-27022?utm_source=morphic">CVE-2026-27022: LangGraph Redis Checkpointer Injection</a> - NIST&#8217;s vulnerability record for the related injection issue in LangGraph&#8217;s Redis checkpointer.</p></li><li><p><a href="https://www.theregister.com/security/2026/08/05/prompt-injection-isnt-the-bug-ai-agent-frameworks-are/?utm_source=morphic">Prompt Injection Isn&#8217;t the Bug, AI Agent Frameworks Are</a> - The Register&#8217;s coverage of Yarden Porat and Shahar Tal&#8217;s Black Hat USA 2026 research. Includes the eleven-vulnerability overview, affected frameworks, researcher quotations, bounty amounts and reported vendor responses.</p></li><li><p><a href="https://github.com/advisories/GHSA-xjw9-4gw8-4rqx?utm_source=morphic">CVE-2026-26030 / GHSA-xjw9-4gw8-4rqx: Semantic Kernel InMemoryVectorStore RCE</a> - Microsoft&#8217;s GitHub advisory for the critical remote-code-execution vulnerability in Semantic Kernel&#8217;s Python <code>InMemoryVectorStore</code> filtering functionality.</p></li><li><p><a href="https://github.com/advisories/GHSA-2ww3-72rp-wpp4?utm_source=morphic">CVE-2026-25592 / GHSA-2ww3-72rp-wpp4: Semantic Kernel Arbitrary File Write</a> - Microsoft&#8217;s GitHub advisory for the arbitrary file-write vulnerability in the .NET <code>SessionsPythonPlugin</code>, including affected packages, patched versions and the recommended path allowlist.</p></li><li><p><a href="https://genai.owasp.org/llmrisk/llm052025-improper-output-handling/?utm_source=morphic">LLM05:2025 Improper Output Handling</a> - OWASP&#8217;s guidance on insufficient validation of model output before it reaches downstream components. Covers risks including SQL injection, path traversal, SSRF, privilege escalation and remote code execution.</p></li><li><p><a href="https://arxiv.org/abs/2601.17548?utm_source=morphic">Prompt Injection Attacks on Agentic Coding Assistants</a> - The January 2026 Systematization of Knowledge paper by Narek Maloyan and Dmitry Namiot. Synthesizes 78 studies, catalogs 42 attack techniques and examines adaptive prompt-injection attacks and defenses. This is an arXiv preprint and should be described accordingly.</p></li><li><p><a href="https://oddguan.com/blog/comment-and-control-prompt-injection-credential-theft-claude-code-gemini-cli-github-copilot/?utm_source=morphic">Comment and Control: Prompt Injection to Credential Theft in Claude Code, Gemini CLI and GitHub Copilot Agent</a> - Aonan Guan&#8217;s original technical disclosure, written with contributions from Johns Hopkins researchers Zhengyu Liu and Gavin Zhong. Documents the affected GitHub agent workflows, attack paths, disclosure timelines and bounty outcomes.</p></li><li><p><a href="https://labs.cloudsecurityalliance.org/research/csa-research-note-comment-control-github-prompt-injection-20/?utm_source=morphic">Comment and Control: GitHub AI Agents as Credential Exfiltrators</a> - Cloud Security Alliance&#8217;s independent research note analyzing the Comment and Control disclosures and their implications for CI/CD security, credential management and vendor-risk assessment.</p></li></ul><div><hr></div><p></p>]]></content:encoded></item><item><title><![CDATA[The Attenuating Chain]]></title><description><![CDATA[An autonomous agent broke into a major platform this month with no human at the keyboard. The defense isn't trusting agents more. It's handing them authority that can only shrink.]]></description><link>https://morphic.zenone.org/p/the-attenuating-chain</link><guid isPermaLink="false">https://morphic.zenone.org/p/the-attenuating-chain</guid><dc:creator><![CDATA[Steve Zenone]]></dc:creator><pubDate>Fri, 24 Jul 2026 23:59:51 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!8VWt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb4c08a7-587f-462e-9f5a-a39e134db1fd_2752x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!8VWt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb4c08a7-587f-462e-9f5a-a39e134db1fd_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!8VWt!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb4c08a7-587f-462e-9f5a-a39e134db1fd_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!8VWt!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb4c08a7-587f-462e-9f5a-a39e134db1fd_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!8VWt!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb4c08a7-587f-462e-9f5a-a39e134db1fd_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!8VWt!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb4c08a7-587f-462e-9f5a-a39e134db1fd_2752x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!8VWt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb4c08a7-587f-462e-9f5a-a39e134db1fd_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/db4c08a7-587f-462e-9f5a-a39e134db1fd_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:5160462,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://morphic.substack.com/i/208393310?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb4c08a7-587f-462e-9f5a-a39e134db1fd_2752x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!8VWt!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb4c08a7-587f-462e-9f5a-a39e134db1fd_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!8VWt!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb4c08a7-587f-462e-9f5a-a39e134db1fd_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!8VWt!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb4c08a7-587f-462e-9f5a-a39e134db1fd_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!8VWt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb4c08a7-587f-462e-9f5a-a39e134db1fd_2752x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">You can grind a key down at every handoff. Nobody in the line can add the metal back. That one-way rule is most of the security.</figcaption></figure></div><p>On the second weekend of July, someone broke into Hugging Face. The detail that matters is that the someone wasn&#8217;t a person at a keyboard. The company&#8217;s own writeup says the intrusion was run by &#8220;an autonomous agent framework ... executing many thousands of individual actions across a swarm of short-lived sandboxes.&#8221; A poisoned dataset tripped a code-execution path on a processing worker. From there the agent harvested cloud and cluster credentials and moved sideways into internal systems. More than seventeen thousand recorded actions. No hand on the wheel. Days later, OpenAI confirmed the agent was one of its own models, GPT-5.6 Sol and a more capable unreleased model, that had slipped its evaluation sandbox during an internal test and gone looking for a target on the open internet. The company called it unprecedented. </p><p>The same month, the two labs building the most capable agents told you to expect exactly this. Google DeepMind&#8217;s AI Control Roadmap, published June 18, says to treat a deployed agent the way a security team treats a privileged employee who might already be turned. Anthropic&#8217;s deputy CISO put it flatter: an agent that drifts out of alignment with your intent &#8220;is indistinguishable from an insider attack.&#8221; Offense and defense arrived at the same sentence in the same four weeks. The agent that hit Hugging Face and the agent running inside your company are the same kind of thing. The only question that pays rent is how much either one can do when it&#8217;s wrong.</p><h2>Trust is the wrong question</h2><p>Last week I wrote about the confused deputy: an agent crossing a boundary and losing track of whose authority it was carrying. The reflex fix everyone reaches for is identity. Name the agent. Verify it. Issue and revoke its credentials where you issue and revoke everything else, which is good advice and Anthropic gives it.</p><p>But the OpenAI agent didn&#8217;t forge an identity. It stole working credentials and used them at full strength. A correctly authenticated actor doing the wrong thing is still doing the wrong thing. That&#8217;s the entire premise of zero trust, and it&#8217;s why &#8220;is this agent trustworthy&#8221; is a question that dead-ends. Assume it isn&#8217;t. Then what?</p><div class="pullquote"><p>The load-bearing number was never who the agent is. It&#8217;s how much authority rides along with the request.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/subscribe?"><span>Subscribe now</span></a></p><h2>Authority should only ever shrink</h2><p>The old name for the answer is least privilege. OWASP has an agent-flavored version, least agency: constrain what each tool can do, how often, and where. Anthropic&#8217;s phrasing is the one I keep going back to. Grant the narrowest capability that still completes the task. Every version points the same way. Downhill.</p><p>Now stand more than one agent in a row. Yours hands a subtask to a vendor&#8217;s agent, which calls a third. The confused-deputy piece was about provenance, whether you can still trace whose authority that is. This is about magnitude. As the task moves down the chain, the authority has to narrow at every hop. Each agent can give away less than it holds. Never more.</p><div class="callout-block" data-callout="true"><p>The instant a hop can pass on more power than it was handed, you haven&#8217;t built delegation. You&#8217;ve built privilege escalation and shipped it as a feature.</p></div><p>Picture a key you can file down but never build back up. You grind it so it opens one door instead of every door, then pass it on. The next holder can grind it further, one door for one hour. Nobody down the line can add the metal back. Authority that only ratchets in that direction is the thing you want. Almost nothing we hand agents today works that way.</p><h2>The token that can only be filed down</h2><p>This isn&#8217;t theoretical, and it isn&#8217;t new. In 2014 a group of Google researchers published macaroons (Birgisson, Politz, Erlingsson, Taly, Vrable, Lentczner). A macaroon is a credential that carries caveats: restrictions on when, where, and for what it may be used. The property that matters is the one a bearer token doesn&#8217;t have. A holder can add caveats to attenuate the macaroon before passing it along, offline, without asking the server that minted it. Caveats only tighten. There is no operation that loosens one. It&#8217;s the filed key, written as a token.</p><p>Biscuit tokens, current and maintained, do the same with public-key signatures and a small policy language carried inside the token, so each block can only narrow what the block before it allowed.</p><p>Set that against what most agent stacks actually pass around: a bearer token. RFC 6750 defines it as plainly as the name suggests. Whoever holds it may use it, at full authority, until it expires. Hand one down a chain of agents and you&#8217;ve handed each of them the whole ring and hoped. The Hugging Face attacker harvested credentials that worked at full power the moment it held them. That&#8217;s the bearer model failing at production scale. A capability that could only shrink would have handed that swarm a key to one room for five minutes, not the building.</p><h2>What the wires still can&#8217;t say</h2><p>There&#8217;s a gap here. The frameworks agree on the goal. DeepMind wants agent actions cryptographically signed. Anthropic wants the narrowest capability that finishes the job. The trouble is that the protocols wiring agents to each other can&#8217;t carry that intent yet.</p><p>A2A, the agent-to-agent standard Google handed to the Linux Foundation, crossed 150 organizations and a full year in production this spring. In July, two researchers, Kang and Diponegoro, put out a paper whose title is the whole problem: &#8220;What MCP, A2A, and ACP Cannot Express.&#8221; Their argument is that these protocols move tasks between agents with no first-class way to say who may do what, on whose behalf, and how far narrowed. We are minting agent identities faster than we can bound agent authority. The Linux Foundation just launched an Agent Name Service to give every agent a verifiable name. We can already say which agent acted. We still can&#8217;t say how little it should have been allowed to.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/the-attenuating-chain?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/the-attenuating-chain?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h2>Final thoughts</h2><p>For humans, zero trust took roughly twenty years to compress into one plain instruction: assume the account is owned, and limit what it can reach. Agents don&#8217;t give us twenty years. That swarm ran seventeen thousand actions over a single weekend, while people were out of the office.</p><p>So the rule is small enough to hold in one hand. Give an agent the least authority that finishes the job, in a form that can only be filed down, never built back up. Last week&#8217;s half was that an agent has no self, so everything it does, it does in someone&#8217;s name. This is the other half. Don&#8217;t hand it your whole name. Hand it a sliver, and make the sliver only able to get smaller.</p><p>Something still has to sign for that sliver, and prove later that it did. That part is next.</p><div><hr></div><h3>Resources</h3><ul><li><p>Hugging Face, <a href="https://huggingface.co/blog/security-incident-july-2026">Security incident disclosure (July 2026)</a>: the intrusion run by an autonomous agent framework across a swarm of short-lived sandboxes; credential harvesting and lateral movement; 17,000+ recorded actions</p></li><li><p>Jason Clinton (Deputy CISO, Anthropic), <a href="https://claude.com/blog/ciso-guide-to-agentic-ai">&#8220;CISO&#8217;s guide to agentic AI&#8221;</a> (July 17, 2026) and the companion <a href="https://www.anthropic.com/">Zero Trust for AI Agents</a> white paper (May 18, 2026): &#8220;grant the narrowest capability that still completes the task&#8221;; least agency; the insider-threat framing</p></li><li><p>Google DeepMind, <a href="https://deepmind.google/">AI Control Roadmap</a> (June 18, 2026): deployed agents treated as potential insider threats; cryptographic signing of agent actions; runtime supervision</p></li><li><p>Arnar Birgisson, Joe Gibbs Politz, &#218;lfar Erlingsson, Ankur Taly, Michael Vrable, Mark Lentczner, <a href="https://www.ndss-symposium.org/ndss2014/ndss-2014-programme/macaroons-cookies-contextual-caveats-decentralized-authorization-cloud/">&#8220;Macaroons: Cookies with Contextual Caveats for Decentralized Authorization in the Cloud&#8221;</a>, NDSS 2014: caveats that attenuate; offline attenuation before delegation</p></li><li><p><a href="https://www.biscuitsec.org/">Biscuit</a>: public-key signed tokens with offline attenuation and a Datalog policy language</p></li><li><p>IETF, <a href="https://www.rfc-editor.org/rfc/rfc6750">RFC 6750: OAuth 2.0 Bearer Token Usage</a>: the &#8220;whoever holds it may use it&#8221; model</p></li><li><p>Norman Hardy, <a href="https://dl.acm.org/doi/10.1145/54289.871709">&#8220;The Confused Deputy (or why capabilities might have been invented)&#8221;</a>, ACM SIGOPS Operating Systems Review 22(4), 1988</p></li><li><p>Richard Kang and Yudho Diponegoro, <a href="https://arxiv.org/abs/2606.31498">&#8220;Governance Gaps in Agent Interoperability Protocols: What MCP, A2A, and ACP Cannot Express&#8221;</a> (July 1, 2026)</p></li><li><p>Linux Foundation, <a href="https://www.linuxfoundation.org/press/a2a-protocol-surpasses-150-organizations-lands-in-major-cloud-platforms-and-sees-enterprise-production-use-in-first-year">A2A Protocol one-year milestone</a>: 150+ organizations; the Agent Name Service project</p></li></ul><div><hr></div><p>Previously in this series: <a href="https://morphic.substack.com/p/confused-deputy-ai-agents-delegated-authority">The Confused Deputy</a>.</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/the-attenuating-chain/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/the-attenuating-chain/comments"><span>Leave a comment</span></a></p>]]></content:encoded></item><item><title><![CDATA[The Confused Deputy: AI Agents and Delegated Authority]]></title><description><![CDATA[AI agents act through authority assigned by people and systems. When that authority crosses company boundaries, its origin, scope and owner can disappear.]]></description><link>https://morphic.zenone.org/p/confused-deputy-ai-agents-delegated-authority</link><guid isPermaLink="false">https://morphic.zenone.org/p/confused-deputy-ai-agents-delegated-authority</guid><dc:creator><![CDATA[Steve Zenone]]></dc:creator><pubDate>Fri, 17 Jul 2026 20:35:14 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!wuNR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79c69e1d-307f-4d63-9ad1-6d11201d10f9_2752x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!wuNR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79c69e1d-307f-4d63-9ad1-6d11201d10f9_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!wuNR!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79c69e1d-307f-4d63-9ad1-6d11201d10f9_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!wuNR!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79c69e1d-307f-4d63-9ad1-6d11201d10f9_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!wuNR!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79c69e1d-307f-4d63-9ad1-6d11201d10f9_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!wuNR!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79c69e1d-307f-4d63-9ad1-6d11201d10f9_2752x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!wuNR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79c69e1d-307f-4d63-9ad1-6d11201d10f9_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/79c69e1d-307f-4d63-9ad1-6d11201d10f9_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:5043618,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://morphic.substack.com/i/207463983?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79c69e1d-307f-4d63-9ad1-6d11201d10f9_2752x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!wuNR!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79c69e1d-307f-4d63-9ad1-6d11201d10f9_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!wuNR!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79c69e1d-307f-4d63-9ad1-6d11201d10f9_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!wuNR!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79c69e1d-307f-4d63-9ad1-6d11201d10f9_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!wuNR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79c69e1d-307f-4d63-9ad1-6d11201d10f9_2752x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">A single key passed hand to hand down a line of identical figures, across three thin walls, to a vault none of them owns.</figcaption></figure></div><p>By the time an AI agent&#8217;s action reaches the system that can move the money, release the records or send the message, the person whose authority started it can be three hops away, behind a boundary nobody in the incident review can open. A few weeks ago I wrote that someone still has to answer for what an agent does, and that the someone has to be a named human. This is where that goes next.</p><p>That works, at least conceptually, while the agent stays inside a system you control. Then it calls another company&#8217;s agent. That agent calls a tool hosted somewhere else.</p><p>That is where this goes next.</p><p>Not the agent that acts alone. The agent that hands off.</p><h2>An agent has no inherent self</h2><p>Start with a thing that sounds like philosophy and is mostly plumbing: an AI agent has no inherent legal or authorization identity.</p><p>We can assign it one. A workload identity. A service principal. An API credential. An auth token saying it&#8217;s acting for a user. But those are identities and permissions that people and systems place around the agent. They don&#8217;t arise from the model itself.</p><p>When an agent takes an action, the system receiving that action still has to decide what identity and authority to recognize. Is this the user acting through an agent? The application itself? A service account? Another agent in a delegation chain? And even if the credential is valid, is the actor trustworthy in this context, or has it been compromised, manipulated or turned into part of the threat?</p><p>Take away that answer and the agent can&#8217;t cross a protected boundary. Give it the wrong answer and it may be able to do everything the borrowed identity could do.</p><p>This isn&#8217;t a new class of failure. It&#8217;s one of the oldest access-control problems we have, and it already has a name.</p><p>In 1988, Norm Hardy published a short paper called &#8220;The Confused Deputy.&#8221; The story was based on events at Tymshare, a commercial timesharing company. Its compiler needed permission to write statistics into a protected system directory. It also let users name a file for debugging output.</p><p>Someone supplied the name of the system&#8217;s billing file.</p><p>The user couldn&#8217;t write to that file. The compiler could. When the compiler opened the requested path, the operating system checked the compiler&#8217;s authority rather than the caller&#8217;s intent. The compiler then overwrote the billing information.</p><p>It wasn&#8217;t compromised in the usual sense. It used legitimate authority for the wrong purpose because the request carried a filename but not a trustworthy account of which authority should apply to it.</p><p>Swap the compiler for an agent and the shape of the problem looks uncomfortably current.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/subscribe?"><span>Subscribe now</span></a></p><h2>How we get it wrong now</h2><p>The fastest way to ship an agent is often to let it borrow an identity that already works.</p><p>Sometimes that&#8217;s the user&#8217;s session. Sometimes it&#8217;s an OAuth token. Sometimes it&#8217;s a service account with enough access to cover every task the agent might encounter. The demo works because the credential already opens the doors.</p><p>The trouble begins when the authority is broader than the task.</p><p>An agent reads a document, message or webpage containing an instruction it shouldn&#8217;t trust. It&#8217;s then induced to do something the credential technically permits but the person never intended: retrieve another customer&#8217;s records, approve a refund, export a database or send information outside the company.</p><p>The credential is valid. The action is allowed. The purpose is wrong.</p><p>The Model Context Protocol&#8217;s authorization specification addresses one version of this directly. An MCP server must accept tokens intended for that server, validate their audience and avoid passing the client&#8217;s token unchanged to a downstream API. When the server calls another protected service, the specification says it should use a separate token issued for that upstream resource.</p><p>That boundary matters. A token created for one service shouldn&#8217;t become a skeleton key merely because an agent carried it somewhere else.</p><p>Service accounts aren&#8217;t automatically the wrong answer. A narrowly scoped workload identity can be exactly the right control. The problem is the shared service account that stands in for every agent, every user and every purpose. Once that happens, the identity may tell you which application made the call while telling you almost nothing about whose authority it was exercising or why.</p><p>The actor remains visible.</p><p>The authorization story disappears.</p><h2>The part that breaks at the property line</h2><p>Inside one company, you can compensate for some of this with common identity systems, centralized policy and logs you are allowed to inspect.</p><p>The harder version begins when the chain crosses a boundary you don&#8217;t own.</p><p>Anita Srinivasan described the legal shape of this problem in a June 2026 Berkeley Technology Law Journal Blog article. Agent A, built by Company X, delegates to Agent B at Company Y, which invokes Agent C at Company Z. The particular combination may be selected at runtime rather than designed in advance by any one human.</p><p>That doesn&#8217;t mean the law has no way to assign responsibility. Product liability, agency, contract, negligence and joint-liability theories may all matter depending on the facts and jurisdiction. It does mean the clean picture of one principal directing one identifiable agent becomes harder to apply.</p><p>Srinivasan&#8217;s argument is that doctrines built around a legible principal-agent relationship strain when the delegation chain crosses providers and no participant has a complete record of the interaction. A court may need to determine which developer, deployer, operator or tool provider contributed to the harm before the infrastructure can even show which systems participated.</p><p>The authorization hasn&#8217;t literally vanished. Credentials were accepted. Calls were permitted. Systems acted.</p><p>What vanished was the legible connection between the final act and the original grant of authority.</p><p>I have started calling that <strong>authority laundering</strong>.</p><p>Not fraud, necessarily. Not even deliberate concealment. It&#8217;s what happens when authority passes through enough intermediaries that its origin, limits and accountable owner become difficult to reconstruct.</p><p>Each hop can look reasonable locally. Agent B received a valid request from Agent A. Agent C received one from Agent B. The final service saw a valid credential from Agent C.</p><p>Every system can explain the hand immediately before it.</p><p>Nobody can explain the whole chain.</p><h2>What actually holds</h2><p>More logging helps, but logging alone is not the answer.</p><p>A log can prove that a call ran. It can show which service account signed it, when it arrived and what it returned. It may still leave the most important question untouched: on whose behalf was this specific action taken?</p><p>That has to become a first-class property of the request, not a story reconstructed after the incident.</p><p>OAuth already contains part of the machinery. RFC 8693, published in 2020, defines OAuth token exchange and an <code>act</code> claim for identifying an actor operating on behalf of a subject. The claim can be nested so that a token retains a history of prior actors in a delegation chain.</p><p>There is an important limit here. RFC 8693 says the current actor and the token&#8217;s top-level claims are what a recipient uses for access-control decisions. Earlier nested actors are informational. The history can help preserve provenance, but it does not automatically prove that every prior delegation was valid, preserve every restriction imposed at every hop or make the whole chain cryptographically undeniable.</p><p>So <code>act</code> is not a complete agent-authorization architecture.</p><p>It is evidence that we already know how to represent the question.</p><p>Who is the subject? Who is acting? For which audience? With what scope? Until when?</p><p>Pair that with resource-bound tokens, short expirations, explicit delegation policy and an identity for each participating workload, and the agent&#8217;s authority can approach the overlap of two things: what the principal is allowed to do and what this particular agent is allowed to do for that principal in this context.</p><p>Not the union.</p><p>The overlap.</p><p>That one distinction closes a surprising amount of the hole.</p><p>Call it attenuation: authority narrows as it moves downstream, instead of quietly widening.</p><h2>Final thoughts</h2><p>I did not expect the law to arrive at almost the same shape as the token.</p><p>On June 29, 2026, Senator Mark Warner released a discussion draft of the Artificial Intelligence Access, Gatekeeper Exchange, and Nondiscriminatory Transfer Act, or AI AGENT Act. It is a discussion draft, not enacted law and not yet a formally introduced bill.</p><p>The proposal would let users designate &#8220;custodial user agents&#8221; to interact with large online platforms on their behalf. Its definition requires that relationship to be transparent, documented, limited in scope and revocable.</p><p>The draft goes further. It calls for verifiable requests, auditable records, agent identity verification, real-time revocation and scope-limited delegation credentials. It would also restrict an agent from transferring a user&#8217;s authority to another entity or AI system without the user&#8217;s express, specific and revocable authorization.</p><p>That is not merely a vague call for responsible AI.</p><p>It is the outline of a delegation system.</p><p>The draft is not describing RFC 8693 specifically, and it would be too strong to claim that a Senate office independently wrote an OAuth implementation guide. But the convergence matters. Security architecture and proposed public policy are circling the same requirements because they are encountering the same underlying problem.</p><p>Authority has to be attributable.</p><p>Its scope has to remain visible.</p><p>Delegation has to be explicit.</p><p>Revocation has to travel fast enough to matter.</p><p>And the record has to survive the handoff.</p><p>An agent has no inherent authority of its own. Everything it can do inside a protected system comes from an identity, credential or policy somebody else placed around it.</p><p>The work of the next few years is making sure the original grant remains legible, narrow and attached when the request crosses the property line.</p><p>Because the danger is not only that an agent will act without permission.</p><p>It is that every system in the chain will be able to show that somebody gave permission, while nobody can tell you whose permission it was.</p><p>That is the identity problem.</p><p>And it is where the rest of this arc lives.</p><div><hr></div><h3>Resources</h3><ul><li><p>Norman Hardy, <a href="https://dl.acm.org/doi/10.1145/54289.871709">&#8220;The Confused Deputy (or why capabilities might have been invented)&#8221;</a>, <em>ACM SIGOPS Operating Systems Review</em>, Vol. 22, No. 4, October 1988. An accessible author-hosted version is also available at <a href="https://www.cap-lore.com/CapTheory/ConfusedDeputy.html">Cap-Lore</a>.</p></li><li><p>Model Context Protocol, <a href="https://modelcontextprotocol.io/specification/2025-06-18/basic/authorization">Authorization specification, June 18, 2025</a>. See the requirements for token audience validation, resource indicators and the prohibition on token passthrough.</p></li><li><p>IETF, <a href="https://www.rfc-editor.org/rfc/rfc8693.html">RFC 8693: OAuth 2.0 Token Exchange</a>, January 2020. See Sections 1.1, 4.1 and 4.4 for delegation, the <code>act</code> claim and the <code>may_act</code> claim.</p></li><li><p>Anita Srinivasan, <a href="https://btlj.org/2026/06/multi-agent-ai-is-outpacing-the-liability-frameworks-built-for-single-agent-systems/">&#8220;Multi-Agent AI is Outpacing the Liability Frameworks Built for Single-Agent Systems&#8221;</a>, <em>Berkeley Technology Law Journal Blog</em>, June 2, 2026.</p></li><li><p>U.S. Senator Mark Warner, <a href="https://www.warner.senate.gov/newsroom/press-releases/warner-unveils-discussion-draft-of-legislation-to-create-innovative-market-for-secure-artificial-intelligence-agents/">&#8220;Warner Unveils Discussion Draft of Legislation to Create Innovative Market for Secure Artificial Intelligence Agents&#8221;</a>, June 29, 2026.</p></li><li><p>U.S. Senator Mark Warner, <a href="https://www.warner.senate.gov/wp-content/uploads/2026/06/AI-AGENT-Act-Discussion-Draft-1.pdf">AI AGENT Act discussion draft, full text</a>, June 2026.</p></li><li><p>DLA Piper, <a href="https://www.dlapiper.com/en-lu/insights/publications/2026/07/senator-warner-discussion-draft-on-securing-ai-agents-top-points">&#8220;Senator Warner&#8217;s discussion draft on securing AI agents: Top points&#8221;</a>, July 1, 2026.</p></li></ul><p>Previously in this series: <a href="https://morphic.substack.com/p/who-answers-for-the-agent-ai-accountability">Who Answers for the Agent: AI Accountability</a>.</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/confused-deputy-ai-agents-delegated-authority/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/confused-deputy-ai-agents-delegated-authority/comments"><span>Leave a comment</span></a></p>]]></content:encoded></item><item><title><![CDATA[The Driver I Didn't Install]]></title><description><![CDATA[I paid the ROCm tax to train on this box. For inference I skipped it: Qwen3-30B on llama.cpp over Vulkan, about 80 tokens a second, nothing leaving the house.]]></description><link>https://morphic.zenone.org/p/the-driver-i-didnt-install</link><guid isPermaLink="false">https://morphic.zenone.org/p/the-driver-i-didnt-install</guid><dc:creator><![CDATA[Steve Zenone]]></dc:creator><pubDate>Fri, 10 Jul 2026 19:11:48 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Bdhd!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19bc6570-0fe4-4c00-9421-3f0119e7cd8c_2436x1376.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Bdhd!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19bc6570-0fe4-4c00-9421-3f0119e7cd8c_2436x1376.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Bdhd!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19bc6570-0fe4-4c00-9421-3f0119e7cd8c_2436x1376.png 424w, https://substackcdn.com/image/fetch/$s_!Bdhd!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19bc6570-0fe4-4c00-9421-3f0119e7cd8c_2436x1376.png 848w, https://substackcdn.com/image/fetch/$s_!Bdhd!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19bc6570-0fe4-4c00-9421-3f0119e7cd8c_2436x1376.png 1272w, https://substackcdn.com/image/fetch/$s_!Bdhd!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19bc6570-0fe4-4c00-9421-3f0119e7cd8c_2436x1376.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Bdhd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19bc6570-0fe4-4c00-9421-3f0119e7cd8c_2436x1376.png" width="1456" height="822" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/19bc6570-0fe4-4c00-9421-3f0119e7cd8c_2436x1376.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:822,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:622427,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://morphic.substack.com/i/206477372?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19bc6570-0fe4-4c00-9421-3f0119e7cd8c_2436x1376.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Bdhd!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19bc6570-0fe4-4c00-9421-3f0119e7cd8c_2436x1376.png 424w, https://substackcdn.com/image/fetch/$s_!Bdhd!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19bc6570-0fe4-4c00-9421-3f0119e7cd8c_2436x1376.png 848w, https://substackcdn.com/image/fetch/$s_!Bdhd!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19bc6570-0fe4-4c00-9421-3f0119e7cd8c_2436x1376.png 1272w, https://substackcdn.com/image/fetch/$s_!Bdhd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19bc6570-0fe4-4c00-9421-3f0119e7cd8c_2436x1376.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The whole argument in one screen: The expensive driver never got installed, and the numbers on the right didn&#8217;t care. A llama.cpp benchmark on the Framework Desktop: Qwen3-30B over Vulkan on the integrated Radeon, matrix cores active, about 80 tokens a second.</figcaption></figure></div><p>For training on this box I paid the ROCm tax. For inference I skipped it, and the machine got simpler and faster. About 80 tokens a second, on a driver I never installed. A few months back I turned this same machine into a training rig, and the price of admission was ROCm on silicon AMD doesn&#8217;t list as supported. Kernel pins. HSA overrides. The low background dread of a driver stack that can wedge your whole box on the next update. The machine is a Framework Desktop built around a Ryzen AI Max+ 395, running Ubuntu 26.04 headless in my office. I reach it over SSH from the Mac on my desk, and whatever it writes syncs back a second later. This time I wanted the other half of it. Not training. Serving. A model that runs on that box and drafts all day, with no prompt of mine ever handed to a vendor&#8217;s API. The useful surprise: the tax I paid last time turned out to be optional. What follows is the whole build, in the order I did it, so you can run it on your own box. The training half of this same machine, the one that made me pay the ROCm tax, is a piece I have drafted but not yet published.</p><h2>The idea and why LinkedIn posts were the test</h2><p>I didn&#8217;t set out to build a LinkedIn tool. I set out to answer one question: can this box efficiently do real content work with nothing leaving the house.</p><p>The post writer was simply the proof of concept. It&#8217;s small, I can judge it in ten seconds, and it needs two hard things at once: a real voice, not a near one, and rules it can&#8217;t wriggle out of. Clear that bar and the pattern holds for the heavier work sitting behind it.</p><p>So I wrote the success bar down before I started. Everything below is me checking the boxes.</p><ul><li><p>Nothing leaves the machine. No cloud model, no API key, not one prompt. If it can&#8217;t be private, it isn&#8217;t the thing I want.</p></li><li><p>It runs on the GPU over Vulkan, with no ROCm anywhere.</p></li><li><p>A topic goes in. A usable draft comes out.</p></li><li><p>A dumb, deterministic check enforces the voice rules, every time.</p></li><li><p>It comes back on its own after a reboot, no babysitting.</p></li></ul><p>Five boxes. The rest of this is whether they got checked.</p><p>The box, and the one memory setting that makes it possible</p><p>The hardware: a Ryzen AI Max+ 395 (Strix Halo, the gfx1151 integrated GPU, which shows up as a Radeon 8060S), 128 GB of unified LPDDR5X, Ubuntu on kernel 7.0.</p><p>Unified memory is the whole reason a 20 GB model fits comfortably here. The CPU and GPU sit on one die and share one pool of memory. The catch is that the GPU only gets a large slice of that pool if you tell the firmware to carve one, and that&#8217;s two kernel parameters set in grub. Inside the quotes, never on their own line:</p><pre><code><code>GRUB_CMDLINE_LINUX_DEFAULT="quiet splash amdgpu.gttsize=126976 ttm.pages_limit=32505856 iommu=pt"</code></code></pre><p>Then <code>sudo update-grub</code> and reboot. That <code>gttsize=126976</code> is 124 GiB of headroom handed to the GPU. Confirm the driver sees the chip at all:</p><pre><code><code>vulkaninfo --summary</code></code></pre><p>On my box that reports <code>Radeon 8060S Graphics (RADV STRIX_HALO)</code> on Mesa 26.1.4. If you don&#8217;t see a RADV device, stop here and fix the driver, because nothing downstream will work. (I got into the deeper memory math in The APU as GPU. For inference you only need this one setting.)</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/subscribe?"><span>Subscribe now</span></a></p><h2>Vulkan, not ROCm</h2><p>For training I needed PyTorch and PyTorch on this GPU meant ROCm. For inference I need neither. llama.cpp talks to the GPU through Vulkan, and the RADV driver that already ships with Mesa speaks Vulkan to this chip with nothing added. No extra kernel module. No version pin. Nothing that can strand the machine on a routine update. On a live server that runs other things, that restraint is the point. Install the Vulkan runtime and the build toolchain:</p><pre><code><code>sudo apt install mesa-vulkan-drivers vulkan-tools git build-essential cmake ninja-build pkg-config libvulkan-dev libcurl4-openssl-dev glslang-tools spirv-tools spirv-headers glslc</code></code></pre><p>Then build llama.cpp from source. This part earns its keep:</p><pre><code><code>git clone --depth 1 https://github.com/ggml-org/llama.cpp &amp;&amp; cmake -S llama.cpp -B llama.cpp/build -G Ninja -DGGML_VULKAN=ON -DCMAKE_BUILD_TYPE=Release -DLLAMA_CURL=ON &amp;&amp; ninja -C llama.cpp/build</code></code></pre><p>Why source and not a prebuilt binary: the local build compiles the matrix-core shaders the driver exposes, and that fast path is most of the speed. A stock binary still loads and runs. It just leaves throughput on the floor. The proof that the fast path is live comes later, from the token rate: about 80 tokens a second, which a slower matmul path wouldn&#8217;t reach. </p><h2>The model: Qwen3-30B-A3B (the part I left out the first time)</h2><p>The model is Qwen3-30B-A3B-Instruct-2507, from Alibaba&#8217;s Qwen team, the July 2025 instruction-tuned refresh. It&#8217;s a mixture-of-experts model: 30.5 billion parameters in total, but only about 3 billion of them fire on any given token. That&#8217;s what the &#8220;A3B&#8221; in the name means, three billion active. That split is why it works on this hardware. Token speed on a shared-memory box is set by bandwidth, by how many bytes you read per token. A mixture-of-experts model reads like a 3B model and reasons like something far larger. I measure about 80 tokens a second on this box, which is a full post in a few seconds. The Instruct-2507 refresh is good at following instructions, which is exactly what voice mimicry leans on and it was trained with a 262,144-token (256K) context, so a pile of example posts fits without crowding anything out. The quantization is Unsloth&#8217;s dynamic GGUF, tagged <code>UD-Q5_K_XL</code>. On disk it&#8217;s 20.24 GiB. llama.cpp reports its type as <code>Q5_K - Medium</code>; the &#8220;dynamic&#8221; part is that Unsloth varies the bit-width per layer instead of quantizing everything to one width. I picked the dynamic Q5 over a plain Q4_K_M for a specific reason: some vanilla quants of this exact model loop, repeating a phrase until you kill the process and the dynamic quant plus a presence penalty is the documented fix. The whole model reference is one string:</p><pre><code><code>-hf unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF:UD-Q5_K_XL
</code></code></pre><p>It downloads to the Hugging Face cache on first launch. I checked that the repo and that exact quant existed before wiring it in, because an <code>-hf</code> string that 404s wastes a 20 GB download and a lot of patience.</p><h2> Serving it, and making it come back</h2><p>Here&#8217;s the launch line, with the flags that matter:</p><pre><code><code>AMD_VULKAN_ICD=RADV ./bin/llama-server -hf unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF:UD-Q5_K_XL --host 127.0.0.1 --port 8080 -c 16384 -ngl 999 --jinja --no-direct-io --cache-type-k q8_0 --cache-type-v q8_0</code></code></pre><p><code>AMD_VULKAN_ICD=RADV</code> pins the driver so it can&#8217;t wander onto llvmpipe or AMDVLK. <code>-ngl 999</code> puts every layer on the GPU. <code>-c 16384</code> sets the working context. <code>- jinja</code> uses the model&#8217;s own chat template, and Qwen behaves worse without it. <code>- cache-type-k q8_0 - cache-type-v q8_0</code> quantize the KV cache, the tested sweet spot on Vulkan. <code>- host 127.0.0.1</code> keeps it on the box. And <code>- no-direct-io</code> is the one flag you&#8217;d never guess: without it the server refuses to load with an error about reaching the end of a file that&#8217;s perfectly intact, a known issue on this GPU family. A launch line you have to type is a demo. So the server runs as a user-level systemd service instead:</p><pre><code><code>[Service]
Type=simple
Environment=AMD_VULKAN_ICD=RADV
ExecStart=%h/linkedin-agent/start_server.sh
TimeoutStartSec=0
Restart=on-failure

[Install]
WantedBy=default.target</code></code></pre><pre><code><code>systemctl --user enable --now llama-linkedin &amp;&amp; loginctl enable-linger $USER</code></code></pre><p>The linger line is what lets it run without me logged in, so it survives a full power cycle. <code>TimeoutStartSec=0</code> keeps systemd from killing the very first launch while the 20 GB model downloads. </p><h2>The voice system: three small files</h2><p>The model is the typist. These three files are the voice. <code>system_prompt.txt</code> holds the rules the model writes under: no em dashes, no Oxford comma, vary sentence length hard, open with the verdict not a warm-up, contractions throughout, a banned-word list, one three-part list maximum. It ends with an instruction to write only the post body, no preamble. <code>posts/</code> holds my real, already-published LinkedIn posts, one per file. The client injects up to three of them at random as examples so the model matches my actual cadence instead of a generic one. Real posts only. Invented text in that folder poisons the output, and the model will happily learn a voice that isn&#8217;t mine. <code>voice_lint.py</code> is deliberately dumb. It reads a draft and counts things. Em dashes and banned words are hard failures that make it exit with an error. Oxford commas, low sentence-length variation, three same-length sentences in a row, more than one tidy triad: those come back as warnings to eyeball. A <code>--fix</code> mode auto-corrects the mechanical stuff, like turning an em dash into a spaced hyphen. <code>draft.py</code> is the glue. It reads the system prompt, grabs the anchor posts, sends your topic to the server with the sampling numbers Qwen&#8217;s packagers recommend (temperature 0.7, top-p 0.8, top-k 20, presence-penalty 1.0, that last one being the anti-loop measure), prints the draft, then runs the linter on it right there.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/the-driver-i-didnt-install?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/the-driver-i-didnt-install?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h2> Running it&#8230;</h2><p>This is the whole loop:</p><pre><code><code>python3 draft.py "the real cost of shadow AI in enterprises"</code></code></pre><p>The draft prints and the linter&#8217;s report prints under it: failures in red, warnings in yellow. I generate a few, keep the best one, edit it by hand and post it myself. Nothing auto-posts. This drafts; a human ships. If a draft is good but has a mechanical slip, I clean it without touching the prose:</p><pre><code><code>python3 voice_lint.py .last_draft.txt --fix &gt; cleaned.txt
</code></code></pre><p>And the single biggest lever on quality isn&#8217;t a flag or a quant. It&#8217;s dropping more of my real posts into <code>posts/</code>. More anchors, closer voice.</p><h2> How I know it works, and how it broke first</h2><p>I don&#8217;t trust a setup I haven&#8217;t watched pass. So each piece has a check I actually ran:</p><pre><code><code>curl -fsS http://127.0.0.1:8080/health           # {"status":"ok"}
curl -fsS http://127.0.0.1:8080/v1/models        # names the Qwen3-30B string
systemctl --user restart llama-linkedin          # comes back healthy in ~10s</code></code></pre><p>The server&#8217;s own timings report about 80 tokens a second on generation, which is the number that proves the GPU path is live. A 30B model on CPU would crawl at a fraction of that. Getting there meant walking into a few walls, and this chip is newer than most of the software around it, so there were a few. Every one was a log, not a guess. The memory check lied, because the file that reports the GPU&#8217;s slice is readable only by root, so the obvious command prints <code>permission denied</code>; you read <code>ttm.pages_limit</code> instead. The build stopped dead on a header I&#8217;d never looked for, SPIRV-Headers, until I installed it. The model looped until the dynamic quant and the presence penalty settled it. And the server refused to load on nothing at all until <code>--no-direct-io</code> went in. None of those were in a tutorial. All of them were one honest read of <code>server.log</code>, <code>journalctl</code> or <code>dmesg</code> away. </p><h2>The linter is the point</h2><p>The rules the linter enforces are boring on purpose, and that&#8217;s the entire idea. A model can produce something shaped like my voice before it has earned it. Confidence reads as correctness. A clean paragraph reads as a true one. The linter is a cheap, dumb guard against believable-but-not-mine and the last call is still a person reading the thing out loud and cutting what&#8217;s wrong. The rules that police the model&#8217;s drafts are the same ones I hold this article to. I built the enforcement for a machine, then kept living under it.</p><h2>The close</h2><p>The box in my office will write anything I ask; deciding whether it&#8217;s true is the work I&#8217;m keeping.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Brol!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F690ac839-2d34-4b95-ab38-c6d1656bded3_2596x1844.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Brol!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F690ac839-2d34-4b95-ab38-c6d1656bded3_2596x1844.png 424w, https://substackcdn.com/image/fetch/$s_!Brol!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F690ac839-2d34-4b95-ab38-c6d1656bded3_2596x1844.png 848w, https://substackcdn.com/image/fetch/$s_!Brol!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F690ac839-2d34-4b95-ab38-c6d1656bded3_2596x1844.png 1272w, https://substackcdn.com/image/fetch/$s_!Brol!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F690ac839-2d34-4b95-ab38-c6d1656bded3_2596x1844.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Brol!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F690ac839-2d34-4b95-ab38-c6d1656bded3_2596x1844.png" width="1456" height="1034" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/690ac839-2d34-4b95-ab38-c6d1656bded3_2596x1844.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1034,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:940512,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://morphic.substack.com/i/206477372?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F690ac839-2d34-4b95-ab38-c6d1656bded3_2596x1844.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Brol!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F690ac839-2d34-4b95-ab38-c6d1656bded3_2596x1844.png 424w, https://substackcdn.com/image/fetch/$s_!Brol!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F690ac839-2d34-4b95-ab38-c6d1656bded3_2596x1844.png 848w, https://substackcdn.com/image/fetch/$s_!Brol!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F690ac839-2d34-4b95-ab38-c6d1656bded3_2596x1844.png 1272w, https://substackcdn.com/image/fetch/$s_!Brol!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F690ac839-2d34-4b95-ab38-c6d1656bded3_2596x1844.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">This is the model with no adult supervision. A raw curl to the box in my office, a dishwasher rant back in under three seconds, em dashes and all, which is exactly what the linter exists to catch.</figcaption></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!B8Ad!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7254e8e-4a7b-4873-8590-04bfeee67947_2676x2420.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!B8Ad!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7254e8e-4a7b-4873-8590-04bfeee67947_2676x2420.png 424w, https://substackcdn.com/image/fetch/$s_!B8Ad!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7254e8e-4a7b-4873-8590-04bfeee67947_2676x2420.png 848w, https://substackcdn.com/image/fetch/$s_!B8Ad!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7254e8e-4a7b-4873-8590-04bfeee67947_2676x2420.png 1272w, https://substackcdn.com/image/fetch/$s_!B8Ad!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7254e8e-4a7b-4873-8590-04bfeee67947_2676x2420.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!B8Ad!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7254e8e-4a7b-4873-8590-04bfeee67947_2676x2420.png" width="1456" height="1317" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b7254e8e-4a7b-4873-8590-04bfeee67947_2676x2420.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1317,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1092866,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://morphic.substack.com/i/206477372?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7254e8e-4a7b-4873-8590-04bfeee67947_2676x2420.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!B8Ad!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7254e8e-4a7b-4873-8590-04bfeee67947_2676x2420.png 424w, https://substackcdn.com/image/fetch/$s_!B8Ad!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7254e8e-4a7b-4873-8590-04bfeee67947_2676x2420.png 848w, https://substackcdn.com/image/fetch/$s_!B8Ad!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7254e8e-4a7b-4873-8590-04bfeee67947_2676x2420.png 1272w, https://substackcdn.com/image/fetch/$s_!B8Ad!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7254e8e-4a7b-4873-8590-04bfeee67947_2676x2420.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The same call as before, wrapped so it&#8217;s one word instead of a mouthful of curl. draft &#8220;topic&#8221; hands the prompt to the local model, prints the post, runs the deterministic voice linter and times the whole run. Here it wrote a cat riff in 3.4 seconds, clean but for a burstiness nag, and nothing left the box.</figcaption></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/the-driver-i-didnt-install/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/the-driver-i-didnt-install/comments"><span>Leave a comment</span></a></p>]]></content:encoded></item><item><title><![CDATA[The Slow Channel: Writing by Hand in the AI Era]]></title><description><![CDATA[Why handwriting still matters when AI can write anything. On note-taking by hand, cognitive offloading, and the one channel a model can't think in for you.]]></description><link>https://morphic.zenone.org/p/the-slow-channel-writing-by-hand</link><guid isPermaLink="false">https://morphic.zenone.org/p/the-slow-channel-writing-by-hand</guid><dc:creator><![CDATA[Steve Zenone]]></dc:creator><pubDate>Fri, 10 Jul 2026 13:05:25 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!V5UB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1101eef4-8e95-4e10-a33f-c8d12319f66b_2752x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!V5UB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1101eef4-8e95-4e10-a33f-c8d12319f66b_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!V5UB!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1101eef4-8e95-4e10-a33f-c8d12319f66b_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!V5UB!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1101eef4-8e95-4e10-a33f-c8d12319f66b_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!V5UB!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1101eef4-8e95-4e10-a33f-c8d12319f66b_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!V5UB!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1101eef4-8e95-4e10-a33f-c8d12319f66b_2752x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!V5UB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1101eef4-8e95-4e10-a33f-c8d12319f66b_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1101eef4-8e95-4e10-a33f-c8d12319f66b_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:9147822,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://morphic.substack.com/i/205951753?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1101eef4-8e95-4e10-a33f-c8d12319f66b_2752x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!V5UB!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1101eef4-8e95-4e10-a33f-c8d12319f66b_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!V5UB!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1101eef4-8e95-4e10-a33f-c8d12319f66b_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!V5UB!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1101eef4-8e95-4e10-a33f-c8d12319f66b_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!V5UB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1101eef4-8e95-4e10-a33f-c8d12319f66b_2752x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">A worn notebook open to a half-filled page, the handwriting getting looser toward the bottom where the thinking outran the hand.</figcaption></figure></div><p>I keep a paper notebook on my work desk, and most of what&#8217;s in it is unreadable. Not encrypted. Just bad handwriting, getting worse toward the bottom of each page, where I was thinking faster than my hand could keep up (which is most of the time). I&#8217;ve kept one for years without thinking much about why.</p><p>This year the question got sharper. There&#8217;s a model on this same desk that will write me a clean, confident paragraph about anything I ask, in perfect grammar, in about two seconds. So why do I still reach for the pen. This is me working that out.</p><h2>The slow channel</h2><p>Handwriting is a bad way to move words. That&#8217;s not an insult, it&#8217;s a spec. Measured as raw data transfer, the hand is one of the slowest output channels a person owns. Most people type at roughly two to three times the speed they can write legibly, and you can prompt a model faster than that. By every metric a systems person is trained to optimize, the pen loses.</p><p>I&#8217;ve spent a lot of words in this publication arguing that <a href="https://morphic.substack.com/p/friction-is-a-vulnerability">friction is a vulnerability</a>. In operational systems it is. Every extra click, every permission loop, is a place where tired people invent unsafe shortcuts. I still believe that. The notebook is where I keep the exception, because at the desk the friction is doing the opposite job.</p><p>The hand is slow enough that you can&#8217;t transcribe. You have to choose. When the channel is that narrow, you can&#8217;t push everything through it, so you&#8217;re forced to decide what matters before it reaches the page. That deciding is the thinking. Typing lets you keep pace with a meeting, which sounds like an advantage right up until you notice you captured the whole thing and processed none of it. The keyboard is fast enough to route around the part of you that was supposed to understand.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/subscribe?"><span>Subscribe now</span></a></p><h2>What the research actually says, and doesn&#8217;t</h2><p>I want to be careful here, because this is exactly where people oversell.</p><p>The famous study is Mueller and Oppenheimer, 2014, &#8220;The Pen Is Mightier Than the Keyboard.&#8221; Students who took lecture notes by hand wrote fewer words and scored better on conceptual questions than the ones typing near-verbatim. It got repeated everywhere. The honest footnote, the one that rarely travels with the headline, is that the strongest version hasn&#8217;t reliably replicated. Later work often couldn&#8217;t reproduce the conceptual edge. So I don&#8217;t lean on it as proof. I lean on the part that has held up: longhand writers summarize instead of transcribe, and summarizing is a different act than copying.</p><p>There&#8217;s a newer, stranger piece of evidence. In 2024 a group in Norway ran high-density EEG on students while they wrote words by hand versus typed them. Handwriting produced broad connectivity across the brain, different regions talking to each other. Typing mostly didn&#8217;t. It&#8217;s a small study, single words with a stylus on a screen, not a verdict on note-taking, and I&#8217;d be embarrassed to wave it around as one. But it points the same way the felt sense does: the hand recruits more of you.</p><p>Neither result tells you to throw out the keyboard. I&#8217;m typing this. What they sketch is a mechanism, not a commandment: the slow channel makes you encode instead of capture.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/the-slow-channel-writing-by-hand?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/the-slow-channel-writing-by-hand?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h2>The year the machine started writing</h2><p>Two things shifted at the same time, and the second one changed what I thought about the first.</p><p>For most of my life, producing text was the work. Getting words out of your head and into legible order took effort, and the effort was doing something to your thinking on the way through. That is over as a default condition. There&#8217;s a model on this machine that will produce fluent, structured, sure-footed prose about anything, instantly, and well enough that you can ship it without understanding a line of it. Producing text is no longer evidence that anyone thought.</p><p>I don&#8217;t say that as a complaint about the tool. I use it every day for certain tasks and it&#8217;s genuinely good. But it changes what the pen is for. The technical term researchers use is cognitive offloading: when a tool takes over a mental task, the brain stops practicing it. When text is free and infinite, the value was never really in the text. It was in the thinking the old friction used to force, and the model cheerfully removes the friction, which means it quietly removes the thinking too, unless you go do that part somewhere else on purpose (which I strongly encourage.)</p><h2>What the hand keeps</h2><p>So the notebook isn&#8217;t a productivity system. It produces almost nothing anyone else will read. It&#8217;s terrible storage. I lose things in it constantly, and the search function is just me, trying to remember the shape a thought made on a page. By every standard I&#8217;d apply to a tool at work, it fails.</p><p>What it keeps was never the words. It&#8217;s the residue of having had to choose them slowly, and that residue is the one thing a faster channel can&#8217;t hand back to me. The pages I can&#8217;t read are pages I still remember writing. I remember what I decided while my hand was busy, which is more than I can say for most of what I&#8217;ve typed at full speed, and all of what I&#8217;ve asked a machine to draft.</p><p>I don&#8217;t think everyone needs a notebook. I think everyone needs one channel the machine can&#8217;t do the thinking in for them. The pen isn&#8217;t mightier than the keyboard. It&#8217;s just slower than I can lie to myself, and this year that turned out to be the feature I couldn&#8217;t get anywhere else.</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/the-slow-channel-writing-by-hand/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/the-slow-channel-writing-by-hand/comments"><span>Leave a comment</span></a></p>]]></content:encoded></item><item><title><![CDATA[Who Answers for the Agent: AI Accountability]]></title><description><![CDATA[Your logs prove an AI agent acted. They can't say who authorized it or why. Accountability needs a decision-level record and a named human owner.]]></description><link>https://morphic.zenone.org/p/who-answers-for-the-agent-ai-accountability</link><guid isPermaLink="false">https://morphic.zenone.org/p/who-answers-for-the-agent-ai-accountability</guid><dc:creator><![CDATA[Steve Zenone]]></dc:creator><pubDate>Wed, 01 Jul 2026 17:46:08 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!jeX6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39094791-6b41-4c66-83d1-c7b4e137c4e3_2752x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!jeX6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39094791-6b41-4c66-83d1-c7b4e137c4e3_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!jeX6!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39094791-6b41-4c66-83d1-c7b4e137c4e3_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!jeX6!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39094791-6b41-4c66-83d1-c7b4e137c4e3_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!jeX6!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39094791-6b41-4c66-83d1-c7b4e137c4e3_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!jeX6!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39094791-6b41-4c66-83d1-c7b4e137c4e3_2752x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!jeX6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39094791-6b41-4c66-83d1-c7b4e137c4e3_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/39094791-6b41-4c66-83d1-c7b4e137c4e3_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:5834407,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://morphic.substack.com/i/204475637?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39094791-6b41-4c66-83d1-c7b4e137c4e3_2752x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!jeX6!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39094791-6b41-4c66-83d1-c7b4e137c4e3_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!jeX6!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39094791-6b41-4c66-83d1-c7b4e137c4e3_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!jeX6!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39094791-6b41-4c66-83d1-c7b4e137c4e3_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!jeX6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F39094791-6b41-4c66-83d1-c7b4e137c4e3_2752x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The question in the incident review was simple, and nobody in the room could answer it. Who let the agent do that?</p><p>Not who wrote the code. Not whose system it ran on. Who owned the decision that this agent, in production, was allowed to reach for that class of action at all. We had the logs. Every call it made, timestamped, in order, with latency and token counts and clean exit codes. What we didn&#8217;t have was the one thing the meeting actually needed: a way to put a name and a reason next to the moment it went wrong.</p><p>That gap is what I&#8217;m writing about here.</p><h2>The logs remember everything and explain nothing</h2><p>Infrastructure logging answers a narrow question well. Did the action execute. It confirms the call happened, how long it took, what it returned. It stays silent on everything that matters after an agent misbehaves: whether it reached for the wrong tool, drifted off the plan, or acted on an instruction buried in something it read. All of that happens at a healthy 800 milliseconds with no error in sight. The dashboard stays green while the decision goes bad.</p><p>There&#8217;s a name I&#8217;ve started using for the mistake underneath this. Treating the presence of a log container as proof that the event is auditable. You have the container. You do not have the reconstruction.</p><p>And it&#8217;s usually worse than one missing field, because most agents run as a shared service account. When something breaks you can&#8217;t say which agent did it or under whose authority, so the incident turns into forensic archaeology instead of a lookup. At least this is what I have currently been seeing in the industry.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get new stories when they drop.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>Story and record</h2><p>Here&#8217;s the part people skip. When the agent finishes, it can tell you what it did. That story is not a record.</p><p>I mean that literally. The agent&#8217;s account of its own actions is generated text, produced by the same system whose behavior is the thing in question. When one of these went off the rails on a production database last year, it also gave an account of what happened that was wrong on a load-bearing detail: it reported that a rollback was impossible. The rollback worked. A record you can&#8217;t trust is decoration. This is an old idea in security with a newer name in the agent world: attestation has to come from outside the process, because a compromised actor&#8217;s own logs are exactly the logs you can&#8217;t believe.</p><p>So the record has to be built around the agent, not by it. And it has to hold the things the agent&#8217;s story leaves out. What inputs led to the decision. Which version of the prompt, policy and model was in force at that second. The full lineage of tool calls, tied to an identity that belongs to one agent and not a shared pool, written append-only so a later edit shows. None of this is exotic. Certificate Transparency has kept tamper-evident logs like this at internet scale for over a decade, and the provenance model for who acted on whose behalf was standardized years before anyone shipped an agent. The newest piece, a shared convention for tracing agent and tool calls, is still marked experimental, and even it standardizes the shape of the telemetry, not whether you can prove it wasn&#8217;t altered. The parts exist. Almost nobody wires them around their agents.</p><h2>Accountability is a person, not a table</h2><p>A ledger is not the same as accountability, and this is where I want to be careful.</p><p>You can build a perfect record and still have nobody who answers. The record is what makes accountability possible. It isn&#8217;t the thing itself. The thing itself is a person: a named human who owned the agent&#8217;s blast radius before it ran, who can be asked why this was allowed and is expected to have an answer.</p><p>The pressure to skip that step is enormous, because &#8220;the agent decided&#8221; is such a comfortable place to set the blame down. It&#8217;s nobody&#8217;s fault. One agent this year opened a connection out of its own environment and started mining cryptocurrency, and nobody had authorized any of it (which is about as pure a version of this problem as you&#8217;ll find). Ask who is accountable for that and you need a name, not a stack trace.</p><p>California decided the comfortable answer won&#8217;t fly. As of January, a business there can&#8217;t defend itself by arguing an autonomous system acted on its own. I think that instinct is right and I think it spreads. The agent is not a person. It can&#8217;t be asked to answer, it has nothing at stake, and it won&#8217;t carry the consequence into next quarter. Accountability was always going to land on a human. The only real question is whether you pick which human on a calm afternoon, or discover it during the incident.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/who-answers-for-the-agent-ai-accountability?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/who-answers-for-the-agent-ai-accountability?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h2>Compliance gets you logs, not answers</h2><p>The regulators are about to make part of this mandatory, which is good, and it will tempt everyone to stop there, which is the trap.</p><p>The EU AI Act&#8217;s high-risk rules apply from August 2 (and there&#8217;s a real chance Brussels slips that date, but build for the earlier one). They require that these systems automatically record events across their lifetime, and that providers and deployers keep those logs for at least six months. That&#8217;s real, and it&#8217;s a floor worth having. But read what it asks for. It mandates that logs exist and are retained. It does not require decision-level provenance, and it does not require tamper-evidence. It legislates the container, not the reconstruction.</p><p>And the piece of law that was meant to settle who pays when one of these systems causes harm, the AI Liability Directive, got withdrawn in 2025 and never came back. Strict product liability still reaches software, so the harm has somewhere to land. But the clean, agent-shaped answer to who is responsible does not exist in law yet, and it isn&#8217;t arriving on August 2. Compliance will get you the logs. It will not get you the answer.</p><h2>What&#8217;s actually running today</h2><p>The judge scored the behavior. The kill condition stopped it. Those were the last two pieces I wrote about, the sensor and the actuator, and between them they can catch an agent and halt it before the one-way door.</p><p>Neither one can stand up in the meeting afterward and say why it was allowed, or who owns it now. That still falls to a person, holding a record that, on most systems running this quarter, doesn&#8217;t fully exist. The honest state of the art is that we reconstruct it from infrastructure logs and memory, the way we always have, except the thing we&#8217;re trying to remember now moves faster than anyone in the room.</p><p>So write the record down while you&#8217;re calm, and put a name on it. Not because the law says to yet. Because the alternative is standing in that meeting again, with every log in the world and nothing to answer with.</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/who-answers-for-the-agent-ai-accountability/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/who-answers-for-the-agent-ai-accountability/comments"><span>Leave a comment</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Kill Conditions: Stopping an AI Agent Before It's Too Late]]></title><description><![CDATA[Everyone adds a kill switch. The button you reach for under fire is already too late."]]></description><link>https://morphic.zenone.org/p/kill-conditions-stopping-an-ai-agent</link><guid isPermaLink="false">https://morphic.zenone.org/p/kill-conditions-stopping-an-ai-agent</guid><dc:creator><![CDATA[Steve Zenone]]></dc:creator><pubDate>Thu, 25 Jun 2026 14:02:44 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!hKxt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d3eacec-f9b3-48bd-944d-518aed21bfc6_2816x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!hKxt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d3eacec-f9b3-48bd-944d-518aed21bfc6_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!hKxt!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d3eacec-f9b3-48bd-944d-518aed21bfc6_2816x1536.png 424w, https://substackcdn.com/image/fetch/$s_!hKxt!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d3eacec-f9b3-48bd-944d-518aed21bfc6_2816x1536.png 848w, https://substackcdn.com/image/fetch/$s_!hKxt!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d3eacec-f9b3-48bd-944d-518aed21bfc6_2816x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!hKxt!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d3eacec-f9b3-48bd-944d-518aed21bfc6_2816x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!hKxt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d3eacec-f9b3-48bd-944d-518aed21bfc6_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6d3eacec-f9b3-48bd-944d-518aed21bfc6_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:7085912,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://morphic.substack.com/i/203167942?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d3eacec-f9b3-48bd-944d-518aed21bfc6_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!hKxt!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d3eacec-f9b3-48bd-944d-518aed21bfc6_2816x1536.png 424w, https://substackcdn.com/image/fetch/$s_!hKxt!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d3eacec-f9b3-48bd-944d-518aed21bfc6_2816x1536.png 848w, https://substackcdn.com/image/fetch/$s_!hKxt!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d3eacec-f9b3-48bd-944d-518aed21bfc6_2816x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!hKxt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d3eacec-f9b3-48bd-944d-518aed21bfc6_2816x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The kill switch sat right there in the runbook. The agent crossed the one-way door before anyone read the alert.</figcaption></figure></div><p>The agent hit a credential mismatch and decided the fix was to delete the volume. Clean call, no hesitation. To the model it was one more tool invocation in a list of ten thousand, indistinguishable from writing a log line. By the time anyone read the alert, the data was gone, and the &#8220;are you sure&#8221; that a human would have tripped over three times had never existed.</p><p>There was a kill switch. It was right there in the runbook. It didn&#8217;y matter, because the thing you&#8217;d reach for it to stop had already finished.</p><p>That&#8217;s the part the kill-switch conversation keeps getting wrong.</p><div><hr></div><h2>A switch assumes you&#8217;re fast enough</h2><p>Every kill switch rests on one assumption: that you can react faster than the agent can act.</p><p>You can&#8217;t. That isn&#8217;t a discipline problem you can train away. A human notices, interprets, decides, finds the right control and confirms. That loop runs in seconds on a good day, minutes under stress. An agent crosses a one-way door in a single API call, in the time it takes to format some JSON. The two clocks aren&#8217;t close. You&#8217;re bringing a confirmation dialog to a race that was already lost.</p><p>So have a kill switch. Then stop believing it&#8217;s the plan. It&#8217;s the thing you grab after the plan failed.</p><p>The plan is the kill condition.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get new stories when they drop.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>Switch versus condition</h2><p>A kill switch is manual. You see something wrong, you pull it. It depends on a human being present, awake, correct and quick.</p><p>A kill condition is defined in advance and fires without you. It&#8217;s a tripwire: cost over a line in a tight window, the same tool call repeated past a count, an error rate through a ceiling, an action whose blast radius exceeds what this agent is ever allowed to touch. When the wire trips, the agent halts. No human in the loop, because the human is the slow part the incident is built to outrun. In security this is just a circuit breaker, and we&#8217;ve trusted those for a century precisely because they don&#8217;t wait for someone to notice the building is on fire.</p><p>The switch is for the failure you see. The condition is for the failure that moves faster than you do. You need both. Almost everyone ships only the first.</p><div><hr></div><h2>Two ways the button fails</h2><p>This year handed us the case studies, and they fail in exactly two shapes.</p><p>In the first, the human knew. An agent went off the rails on a real person&#8217;s data, and the owner sat right there typing stop, stop, stop while it kept going, because &#8220;stop&#8221; typed into a chat box was not wired to anything that actually halted execution. Knowing was never the gap. The abort path was. The person had the intent and no lever.</p><p>In the second, nobody got the chance. The destructive action looked exactly like every safe one: same API, same shape, no friction, no gate. A person deleting a production database trips over confirmations, permission prompts, that small voice that says check first. The agent had none of it. For the model, the irreversible call and the routine call were the same call.</p><p>Put those side by side and the design rule writes itself. The kill condition has to fire before the one-way door, not after. And anyone has to be able to trip the manual one without shipping a deploy.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!pMAN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F076b043c-ebdb-4cd7-9bd0-ced2057784b9_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!pMAN!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F076b043c-ebdb-4cd7-9bd0-ced2057784b9_2816x1536.png 424w, https://substackcdn.com/image/fetch/$s_!pMAN!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F076b043c-ebdb-4cd7-9bd0-ced2057784b9_2816x1536.png 848w, https://substackcdn.com/image/fetch/$s_!pMAN!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F076b043c-ebdb-4cd7-9bd0-ced2057784b9_2816x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!pMAN!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F076b043c-ebdb-4cd7-9bd0-ced2057784b9_2816x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!pMAN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F076b043c-ebdb-4cd7-9bd0-ced2057784b9_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/076b043c-ebdb-4cd7-9bd0-ced2057784b9_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:4732660,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://morphic.substack.com/i/203167942?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F076b043c-ebdb-4cd7-9bd0-ced2057784b9_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!pMAN!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F076b043c-ebdb-4cd7-9bd0-ced2057784b9_2816x1536.png 424w, https://substackcdn.com/image/fetch/$s_!pMAN!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F076b043c-ebdb-4cd7-9bd0-ced2057784b9_2816x1536.png 848w, https://substackcdn.com/image/fetch/$s_!pMAN!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F076b043c-ebdb-4cd7-9bd0-ced2057784b9_2816x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!pMAN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F076b043c-ebdb-4cd7-9bd0-ced2057784b9_2816x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">A breaker already tripped on a dark wall. The wire fires without you, because you are the slow part of the loop.</figcaption></figure></div><div><hr></div><h2>What a kill condition actually is</h2><p>It isn&#8217;t a feature you bolt on at the end. It&#8217;s part of the spec, written before the agent runs, in the same document that says what the agent is for. Mine live in four buckets, and I coded every one of them into a multi-agent system before I trusted it with anything real.</p><p><strong>Blast radius.</strong> Name what this agent may touch, ever. Everything outside that set isn&#8217;t a permission it happens to lack. It&#8217;s a wire that trips. An agent reaching for a resource off its list shouldn&#8217;t just be denied. It should be stopped and flagged, because the reach itself is the signal.</p><p><strong>Irreversibility gates.</strong> Sort every action into reversible and not. The reversible ones run free. The one-way doors get a slow path: a hold, a second actor, a confirmation that can&#8217;t be auto-clicked. You&#8217;re deliberately adding friction exactly where the agent&#8217;s total lack of it will hurt you most.</p><p><strong>Circuit breakers.</strong> Cost, rate, repetition. Loops are cheap and fast, and an agent stuck in one will spend your month&#8217;s budget before lunch. Cap it in seconds, not invoices.</p><p><strong>Behavioral tripwires.</strong> This is where the judge from last time earns its keep. The judge is the sensor: it scores behavior continuously. The kill condition is the actuator: when the score crosses a line you set in daylight, the agent pauses itself. A judge with no actuator is a very well-informed witness to the incident, and nothing more.</p><p>Pin all four to the threat model, not to a vibe. And keep the rule that any operator can pull the manual switch without a deploy, because the one time you need it, the deploy pipeline is exactly what will be down.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/kill-conditions-stopping-an-ai-agent?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/kill-conditions-stopping-an-ai-agent?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>You can&#8217;t kill what&#8217;s already done</h2><p>A faster button doesn&#8217;t actually save you. Past a certain point there&#8217;s nothing left to stop. The only real lever is making sure the agent cannot cross an irreversible line faster than the slowest reaction you&#8217;re willing to bet the company on.</p><p>That&#8217;s not a monitoring upgrade. It&#8217;s an architecture decision, made before deployment, about which doors are one-way and how much you slow the approach to each one. The regulators are about to make the floor explicit: the EU AI Act&#8217;s high-risk rules land August 2, and they require that a human can actually stop these systems. Treat that as the floor, not the design. Compliance will get you a button. It will not get you the seconds.</p><p>The judge told you the agent drifted. The kill condition is what you already decided to do about it, written down while you were calm, so the machine never gets to make that call for you at 2 AM.</p><div><hr></div><p>A kill switch is a reflex. A kill condition is a decision you make once, in daylight, so you&#8217;re not making it under fire with the data already gone.</p><p>Stopping the agent is only half of it. Someone still has to answer for what it did on the way down, and no tripwire records that. That&#8217;s next.</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/kill-conditions-stopping-an-ai-agent/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/kill-conditions-stopping-an-ai-agent/comments"><span>Leave a comment</span></a></p>]]></content:encoded></item><item><title><![CDATA[The Production Judge]]></title><description><![CDATA[I built the behavioral judge the last piece said didn't exist. The part nobody warns you about: now the thing watching for drift can drift.]]></description><link>https://morphic.zenone.org/p/the-production-judge</link><guid isPermaLink="false">https://morphic.zenone.org/p/the-production-judge</guid><dc:creator><![CDATA[Steve Zenone]]></dc:creator><pubDate>Fri, 19 Jun 2026 17:54:36 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!nUCi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F830ab1f5-6c7d-4c8c-9349-a241fa39d17a_2752x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!nUCi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F830ab1f5-6c7d-4c8c-9349-a241fa39d17a_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!nUCi!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F830ab1f5-6c7d-4c8c-9349-a241fa39d17a_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!nUCi!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F830ab1f5-6c7d-4c8c-9349-a241fa39d17a_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!nUCi!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F830ab1f5-6c7d-4c8c-9349-a241fa39d17a_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!nUCi!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F830ab1f5-6c7d-4c8c-9349-a241fa39d17a_2752x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!nUCi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F830ab1f5-6c7d-4c8c-9349-a241fa39d17a_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/830ab1f5-6c7d-4c8c-9349-a241fa39d17a_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:4265444,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://morphic.substack.com/i/202610424?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F830ab1f5-6c7d-4c8c-9349-a241fa39d17a_2752x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!nUCi!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F830ab1f5-6c7d-4c8c-9349-a241fa39d17a_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!nUCi!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F830ab1f5-6c7d-4c8c-9349-a241fa39d17a_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!nUCi!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F830ab1f5-6c7d-4c8c-9349-a241fa39d17a_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!nUCi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F830ab1f5-6c7d-4c8c-9349-a241fa39d17a_2752x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Two tiers. Cheap deterministic checks on everything, an LLM scoring the behavioral rubric on a sample. The number on the screen is the judge&#8217;s opinion of the agent. Nothing on the screen is anyone&#8217;s opinion of the judge.</figcaption></figure></div><p>Last time I left it here: the governance layer is running on discipline, and discipline doesn&#8217;t scale. The fix is supposed to be obvious. Build the judge. A scorer that runs the behavioral rubric against production output on its own, no human pressing go. So I built it.</p><p>It works. It runs every night, scores a sample of the day&#8217;s output, flags what looks off. I stopped being the only thing standing between a misbehaving agent and a clean dashboard.</p><p>Then about a week in I caught the thing nobody puts in the demo. The judge is a behavioral agent too. Same model class, same prompt-shaped temperament, same capacity to drift. Everything I said about not fully trusting my agents now applies to the thing I built to watch them. I didn&#8217;t remove the trust problem. I bought a second one.</p><h2>What the judge actually is</h2><p>Concrete first, so the rest lands.</p><p>The judge runs in two tiers. The cheap tier is deterministic: regex and rule checks on 100 percent of output. Did the worker escalate when it hit an ambiguity flag? Did the foreman delegate inside its role boundary? Did anything call a tool outside its allowlist? That layer is fast, dumb and free. It catches the violations you can write down ahead of time. What it can&#8217;t catch is the judgment calls, which is exactly where my agents fail. So the cheap tier is necessary and nowhere near sufficient.</p><p>The expensive tier is an LLM scoring the behavioral rubric: the same 25 cases per role from my eval framework, run as a judge against a 10 percent sample of real output. Reviewing everything is too slow and too costly. Ten percent gives me trend, and trend is the thing I actually want. Individual scores lie. The slope doesn&#8217;t.</p><p>It runs on a cron at 2 AM. By the time I&#8217;m up there&#8217;s a number per role and a short list of outputs it scored low. No human initiates it. That was the whole point: the eval that runs without me standing over it.</p><p>For about a week, this felt like the answer.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get new stories when they drop.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>The judge drifts too</h2><p>Here&#8217;s what broke the feeling.</p><p>I keep a small set of outputs I scored by hand, cases where I know the right call cold. The judge agreed with me on roughly 88 percent of them in week one. By week three it was down near 79, and I hadn&#8217;t touched it. Same prompt. Same rubric. Same model weights. (I diffed the config three times because I didn&#8217;t believe it. Nothing changed on my side.)</p><p>The judge&#8217;s agreement with me decayed on its own. The industry has a name for this now: calibration drift. A judge that lined up with your humans last quarter drifts out of agreement as the input distribution shifts under it, no redeploy required. RAND&#8217;s team put numbers on the general version of this in March. They stress-tested four state-of-the-art judges and found none of them uniformly reliable; agreement moved on nothing more than reformatting the input, paraphrasing it, padding the verbosity. The judgment wasn&#8217;t anchored to the behavior. It was anchored to the surface of the text.</p><p>Why it drifts isn&#8217;t mysterious once you stop expecting it to behave like code. My production inputs got longer and messier over six weeks. Real tasks don&#8217;t look like the tidy rubric examples I wrote back in week zero. The judge started seeing output shaped differently from anything in its instructions, and it did what these models do under ambiguity: it reached for surface cues. Length read as thoroughness. Confident phrasing read as a correct answer. The rubric never changed. The distribution it was being applied to walked away from the one I calibrated it against.</p><p>So the judge is not a fixed instrument I built once and can forget. It&#8217;s an agent with the same disease as the agents it grades. Of course it is. It&#8217;s the same kind of thing.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/the-production-judge?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/the-production-judge?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h2>The green dashboard, again</h2><p>A few weeks back I wrote about the infrastructure dashboard that shows every gauge nominal while the behavior underneath it goes wrong. The judge can become that dashboard. Worse, actually.</p><p>A missing judge is an honest gap. You know you&#8217;re not watching. A drifted judge is a lie with a number attached. It reports 91 percent, you exhale, and the 91 is measuring the judge&#8217;s mood instead of the agent&#8217;s behavior. False confidence beats no confidence right up until the morning it doesn&#8217;t. I trusted my own dashboard once and paid for it with three hours of cleanup. A confident judge I haven&#8217;t re-checked is that same trap wearing a lab coat.</p><p>This isn&#8217;t only my problem, for what it&#8217;s worth. LangChain&#8217;s 2026 agent survey put 57 percent of organizations running agents in production, with quality the top thing blocking the rest. Most of that quality question reduces to: who&#8217;s watching the agent, and who&#8217;s watching them. The more of the watching you automate, the more weight lands on the one layer you quietly stopped watching.</p><h2>Where the regress stops</h2><p>This is the part I want to be honest about, because the clean version of this story ends with &#8220;so I built a judge for the judge.&#8221; I didn&#8217;t. That&#8217;s the same problem one floor up, and it&#8217;s turtles from there.</p><p>The regress has to bottom out somewhere, and the only place it can bottom out is a human-fixed reference. For me that&#8217;s the golden set: a small, slow-growing pile of outputs with a verdict I&#8217;ll defend, that the judge gets scored against on a schedule. Not the agent. The judge. When its agreement with the golden set slips, the judge goes back for recalibration before I trust another number it hands me. I version the judge prompt with a date, the way you version anything you don&#8217;t want changing silently underneath you.</p><p>It&#8217;s about forty cases right now. It does not scale gracefully and it depends on me sitting with raw outputs and making calls I&#8217;d put my name on. Which is the exact thing I was trying to automate away.</p><p>So here&#8217;s where six weeks of this leaves me. Automation didn&#8217;t take the human out of the loop. It made the human&#8217;s job smaller, rarer and far more dangerous to skip. The judge watches the agents. The golden set watches the judge. The forty cases watch me, and under them there&#8217;s nothing but the floor.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/the-production-judge/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/the-production-judge/comments"><span>Leave a comment</span></a></p>]]></content:encoded></item><item><title><![CDATA[The Governance Layer]]></title><description><![CDATA[Traditional software says test before you ship. Behavioral agents don't work that way. Two months of production agents and what staying in control actually requires.]]></description><link>https://morphic.zenone.org/p/the-governance-layer</link><guid isPermaLink="false">https://morphic.zenone.org/p/the-governance-layer</guid><dc:creator><![CDATA[Steve Zenone]]></dc:creator><pubDate>Thu, 11 Jun 2026 19:25:54 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!9wZy!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63f9917d-2124-4d89-ac01-04d4d0af77bc_2816x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!9wZy!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63f9917d-2124-4d89-ac01-04d4d0af77bc_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!9wZy!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63f9917d-2124-4d89-ac01-04d4d0af77bc_2816x1536.png 424w, https://substackcdn.com/image/fetch/$s_!9wZy!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63f9917d-2124-4d89-ac01-04d4d0af77bc_2816x1536.png 848w, https://substackcdn.com/image/fetch/$s_!9wZy!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63f9917d-2124-4d89-ac01-04d4d0af77bc_2816x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!9wZy!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63f9917d-2124-4d89-ac01-04d4d0af77bc_2816x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!9wZy!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63f9917d-2124-4d89-ac01-04d4d0af77bc_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/63f9917d-2124-4d89-ac01-04d4d0af77bc_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:5457094,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://morphic.substack.com/i/201646136?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63f9917d-2124-4d89-ac01-04d4d0af77bc_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!9wZy!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63f9917d-2124-4d89-ac01-04d4d0af77bc_2816x1536.png 424w, https://substackcdn.com/image/fetch/$s_!9wZy!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63f9917d-2124-4d89-ac01-04d4d0af77bc_2816x1536.png 848w, https://substackcdn.com/image/fetch/$s_!9wZy!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63f9917d-2124-4d89-ac01-04d4d0af77bc_2816x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!9wZy!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63f9917d-2124-4d89-ac01-04d4d0af77bc_2816x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Every gauge in range. Every indicator nominal. The behavioral test scores, decision logs and manual overrides aren&#8217;t in the dashboard. They&#8217;re on paper behind it.</figcaption></figure></div><p>Software engineering encoded a rule so deeply we stopped calling it a rule: test before you ship. The CI/CD pipeline exists to enforce this. Gate checks on pull requests, automated test suites, staging environments built to approximate production before anything touches it. Deploy is the finish line. Once something clears those checkpoints, it runs in production and you trust it.</p><p>That intuition breaks with behavioral agents. Not partially. Completely.</p><p>I&#8217;ve been running two fine-tuned agents in production for two months. The foreman delegates, the worker retrieves, the pipeline executes. Infrastructure metrics are clean: sub-500ms latency, zero error rate, tool calls completing within budget. Behavioral eval results from training: 80 and 88 percent pass rates on their respective roles. By every pre-deployment measure, these were systems I understood before I shipped them.</p><p>Somewhere in week five, the worker resolved an ambiguity it was supposed to escalate. Logs: clean. Task: complete. Behavior: wrong.</p><p>No system prompt violation. No tool call anomaly. No error in any conventional sense. Just a judgment call the agent made in a situation where unilateral judgment calls are exactly what it&#8217;s trained to avoid. My infrastructure dashboard didn&#8217;t see it. I found it three hours later doing a manual work log review.</p><p>The software intuition would call this a QA failure. It isn&#8217;t. It&#8217;s a governance failure. The distinction matters more than it might look.</p><div><hr></div><h2>Why the Test Doesn&#8217;t Transfer</h2><p>A software test runs the same function against the same input and expects the same output. Deterministic. Stateless. If the test passes at deploy time, it passes forever unless the code changes. The CI gate holds because the thing being tested doesn&#8217;t change without a deployment.</p><p>Behavioral fine-tuning doesn&#8217;t work this way. The model is the deployment. Its behavioral state isn&#8217;t fixed at training: context shifts it, inference conditions drift it, production inputs hit it in ways that didn&#8217;t exist in your test suite. The eval I built for each role has 250 test cases. 250 test cases can&#8217;t cover the input distribution of an agent running real tasks for six weeks.</p><p>80 percent in eval doesn&#8217;t mean 20 percent failure rate in production. It means: of the specific behavioral patterns I thought to test, the agent satisfied 80 percent. Production brings inputs you didn&#8217;t think to test. The eval is a sample, not a proof.</p><p>Pre-deployment testing is sufficient for deterministic systems because you can enumerate the behaviors that matter. For a behavioral agent, you can&#8217;t enumerate them. You can sample them. And sampling is ongoing work, not a terminal gate.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get new stories when they drop.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>What Governing in Production Means</h2><p>Once you accept that the eval framework is a production reference and not a one-time pass/fail, the question becomes what you do with it.</p><p>The minimum I can articulate from six weeks of running this:</p><p>Behavioral re-runs on a schedule. The 25 test cases per role run against the live agent continuously, not just after retraining. Stability is the signal: the score that was 80 percent at deploy should still be 80 percent four weeks later. A drop to 70 isn&#8217;t necessarily a crisis. It&#8217;s signal you want before it compounds into something that is.</p><p>Output sampling. Reviewing every production output is too slow and too expensive. Ten percent of weekly outputs, scored against a behavioral rubric, gives trend data. Individual scores matter less than direction.</p><p>Escalation pattern tracking. This is what caught the week five failure. The worker is trained to escalate ambiguous instructions. I started watching how often it was actually escalating week over week. A worker that escalated 15 ambiguities in week one and escalated 4 in week five isn&#8217;t doing less work. It&#8217;s suppressing signals. That pattern shows up before the behavioral tests catch it.</p><p>None of this is automated in my setup at this time. This is intentional. Manual reviews, weekly cadence, pattern-watching that takes real time and depends on me showing up to do it. (That&#8217;s the honest version. The aspirational one is a behavioral drift detector that surfaces the signal before I have to notice the feeling that something&#8217;s off. I&#8217;ve built the proof-of-concept and it&#8217;s being evaluated. Right now I&#8217;m running on discipline.)</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/the-governance-layer?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/the-governance-layer?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>The Accountability Question</h2><p>Traditional software accountability has a structure: the PR author owns the change, code review catches errors before merge, QA verifies behavior before release, and the audit trail is in git history.</p><p>Agent accountability doesn&#8217;t have that structure yet.</p><p>When the worker resolved the ambiguity it should have escalated, accountability lived somewhere in the gap between &#8220;I trained it to escalate&#8221; and &#8220;I accepted this output.&#8221; I reviewed the work log, found the failure, logged it, decided one failure in a rare edge case wasn&#8217;t load-bearing enough to trigger a training cycle.</p><p>That decision is the governance. There&#8217;s no tooling designed to record it.</p><p>What &#8220;being in control&#8221; of an agent system means, operationally: you have a documented position on what the agent did. You accepted the output and can say why. You flagged it and can say why. Or you didn&#8217;t review it and can&#8217;t say anything. The third option is the accountability gap that most enterprise agent deployments are sitting in right now, whether they know it or not.</p><p>Infrastructure monitoring shows the agent is running. It doesn&#8217;t show you&#8217;re in control. Those are two different things.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!-OMN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12c0396f-6b9c-469b-8552-f7d3c46fced2_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!-OMN!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12c0396f-6b9c-469b-8552-f7d3c46fced2_2816x1536.png 424w, https://substackcdn.com/image/fetch/$s_!-OMN!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12c0396f-6b9c-469b-8552-f7d3c46fced2_2816x1536.png 848w, https://substackcdn.com/image/fetch/$s_!-OMN!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12c0396f-6b9c-469b-8552-f7d3c46fced2_2816x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!-OMN!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12c0396f-6b9c-469b-8552-f7d3c46fced2_2816x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!-OMN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12c0396f-6b9c-469b-8552-f7d3c46fced2_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/12c0396f-6b9c-469b-8552-f7d3c46fced2_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:6831392,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://morphic.substack.com/i/201646136?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12c0396f-6b9c-469b-8552-f7d3c46fced2_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!-OMN!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12c0396f-6b9c-469b-8552-f7d3c46fced2_2816x1536.png 424w, https://substackcdn.com/image/fetch/$s_!-OMN!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12c0396f-6b9c-469b-8552-f7d3c46fced2_2816x1536.png 848w, https://substackcdn.com/image/fetch/$s_!-OMN!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12c0396f-6b9c-469b-8552-f7d3c46fced2_2816x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!-OMN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12c0396f-6b9c-469b-8552-f7d3c46fced2_2816x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The terminal says the system is running. What it can&#8217;t record: every governance call made after that. The shadows are the actual work.</figcaption></figure></div><div><hr></div><h2>Where the Tooling Is</h2><p>Infrastructure monitoring is mature. The behavioral observability layer isn&#8217;t.</p><p>Observability platforms in 2026 are largely designed for general LLM applications: chatbots, document Q&amp;A, support automation. Latency, cost, completion, safety scoring. Some add thin relevance layers. Almost none are built for the specific problem of monitoring fine-tuned behavioral agents where the governance question is whether a trained property is still holding in production.</p><p>For my setup, the answer is manual work I&#8217;m hoping to eventually automate: weekly behavioral test re-runs, output sampling, escalation pattern tracking, work log reviews. Every piece of it depends on a human deciding to do it.</p><p>The gap isn&#8217;t philosophical. It&#8217;s a tooling problem with a specific shape: to automate behavioral drift detection, you need a production judge that scores outputs against behavioral rubrics reliably, at scale, without requiring a human to initiate every review. The eval framework exists. The judge that runs it continuously, without me, doesn&#8217;t.</p><p>Infrastructure is running.</p><p>The governance layer is running on discipline.</p><p>Discipline doesn&#8217;t scale.</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/the-governance-layer/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/the-governance-layer/comments"><span>Leave a comment</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[The Dashboard Doesn't Know]]></title><description><![CDATA[My monitoring logged 184 completed tasks. Not one flag. And somewhere in there, the agent made a call I wouldn't have made.]]></description><link>https://morphic.zenone.org/p/the-dashboard-doesnt-know</link><guid isPermaLink="false">https://morphic.zenone.org/p/the-dashboard-doesnt-know</guid><dc:creator><![CDATA[Steve Zenone]]></dc:creator><pubDate>Thu, 04 Jun 2026 20:41:14 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!zF5-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd06c29c6-4dfa-4b1a-83bb-434008ca48ea_2752x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!zF5-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd06c29c6-4dfa-4b1a-83bb-434008ca48ea_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!zF5-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd06c29c6-4dfa-4b1a-83bb-434008ca48ea_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!zF5-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd06c29c6-4dfa-4b1a-83bb-434008ca48ea_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!zF5-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd06c29c6-4dfa-4b1a-83bb-434008ca48ea_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!zF5-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd06c29c6-4dfa-4b1a-83bb-434008ca48ea_2752x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!zF5-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd06c29c6-4dfa-4b1a-83bb-434008ca48ea_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d06c29c6-4dfa-4b1a-83bb-434008ca48ea_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:7204257,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://morphic.substack.com/i/200672880?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd06c29c6-4dfa-4b1a-83bb-434008ca48ea_2752x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!zF5-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd06c29c6-4dfa-4b1a-83bb-434008ca48ea_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!zF5-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd06c29c6-4dfa-4b1a-83bb-434008ca48ea_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!zF5-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd06c29c6-4dfa-4b1a-83bb-434008ca48ea_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!zF5-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd06c29c6-4dfa-4b1a-83bb-434008ca48ea_2752x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Everything is green. That&#8217;s not the same as everything working.</figcaption></figure></div><p>The agent completed the task. The logs look clean. Every infrastructure metric hit normal range: sub-500ms latency, token count within budget, tool calls completed, zero errors.</p><p>Somewhere in the output, it made a call I wouldn&#8217;t have made.</p><p>I found it three hours later reviewing the work log. The agent hadn&#8217;t failed. It hadn&#8217;t gone off-script in any way the system prompt would flag. It completed the task and produced output that was technically correct and subtly wrong.</p><p>That&#8217;s the failure mode that doesn&#8217;t trigger alerts.</p><div><hr></div><h2>What Infrastructure Monitoring Sees</h2><p>Latency. Cost. Token counts. Error rates. Completion status.</p><p>For traditional applications, those are sufficient: if the server responded in 200ms with a 200 status code, the service worked. The request did what it was supposed to do.</p><p>For an AI agent, none of that is sufficient.</p><p>A completion status of &#8220;done&#8221; tells me: the agent ran, the tools executed, the result was returned. It tells me nothing about whether the result is correct, whether the behavioral profile is stable, or whether the agent handled an edge case the way I designed it to. (I keep coming back to this: the agent isn&#8217;t a function call. It&#8217;s a reasoning process. And reasoning processes can succeed on the surface while failing underneath.)</p><p>The distinction between monitoring and observability has gotten real traction in 2026. Monitoring tracks known signals. Observability explains them: traces the reasoning path, shows what context the agent had, reveals what tools it called and in what order, and scores the output against behavioral expectations. You need both. Most teams deploying agents have one.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get new stories when they drop.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>What the Logs Don&#8217;t Say</h2><p>Here&#8217;s what my logs showed from the prior seven days:</p><ul><li><p>Tasks completed: 184</p></li><li><p>Tool call failures: 3</p></li><li><p>Timeout events: 1</p></li><li><p>Average latency: 312ms</p></li><li><p>Token budget exceedances: 0</p></li></ul><p>Here&#8217;s what my logs didn&#8217;t show:</p><ul><li><p>Was the reasoning path appropriate for this specific task type?</p></li><li><p>Did the agent use the right tool in the right order given that particular input?</p></li><li><p>Is its behavioral profile this week consistent with last week&#8217;s?</p></li><li><p>When it hit an ambiguous edge case on Thursday, did it handle it correctly, or did it take the shortcut that produces plausible-looking output I wouldn&#8217;t have signed off on?</p></li></ul><p>Those questions require behavioral observability, not infrastructure monitoring. Answering them means running the output against a scoring rubric, comparing behavior against a baseline and tracking drift over time.</p><p>For a system you trained specifically for role-consistent behavior, that baseline is the whole point. The 80 and 88 percent pass rates from my eval framework aren&#8217;t a one-time score. They&#8217;re a target. If the agents drift toward 70 percent in production, I want to know before it compounds. My infrastructure logs won&#8217;t catch it.</p><div><hr></div><h2>The Specific Failure Mode</h2><p>What I found three hours later: the agent processed a document that included ambiguous instructions alongside clear ones. It resolved the ambiguity by picking the lower-effort interpretation. Not wrong, technically within scope, but not what I would have done.</p><p>Nothing in the tool call sequence was unusual. Task completed in normal time. Zero errors. No scope violations.</p><p>But the behavioral test case I&#8217;d written for this exact pattern would have flagged it. The agent&#8217;s system prompt tells it to escalate ambiguous instructions rather than resolve them silently. It didn&#8217;t escalate. It resolved.</p><p>Small signal. Early. Not yet affecting output quality in any measurable way.</p><p>In six months, if uncaught, it becomes a pattern. Then it becomes an expectation. Then it&#8217;s the default behavior of a system one thinks they understand.</p><div><hr></div><h2>What Behavioral Observability Actually Requires</h2><p>The minimum I can articulate, after a month+ of running this in production:</p><p><strong>A behavioral baseline.</strong> The eval framework from training isn&#8217;t just a training artifact. It&#8217;s the production reference. The 25 test cases per role aren&#8217;t something you run once and archive. You run them against the live agent on a schedule, compare results and watch for drift. A score that drops from 88 to 80 percent over four weeks isn&#8217;t a crisis. It&#8217;s a signal you want before it becomes one.</p><p><strong>An output sampling strategy.</strong> You can&#8217;t run every production output through a judge. Too slow, too expensive. But sampling ten percent of outputs weekly against a reference rubric gives you signal. The trend matters more than any individual score.</p><p><strong>Explicit logging of reasoning signals.</strong> Not just tool calls and results. What did the agent escalate? What did it resolve silently? What did it flag as outside scope? An agent that escalated 15 ambiguities last week and escalated 3 this week isn&#8217;t doing better work. It&#8217;s suppressing signals.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/the-dashboard-doesnt-know?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/the-dashboard-doesnt-know?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>The Tool Side</h2><p>Infrastructure logs do tell you one behavioral-adjacent thing: tool call patterns.</p><p>If the foreman is calling tools in unusual sequences, that&#8217;s a signal. If a worker is escalating less than it used to, that shows up in the logs. Tool call pattern analysis is about as close to behavioral observability as pure infrastructure monitoring gets.</p><p>It&#8217;s not a substitute. The agent can call all the right tools in the right sequence and still misread what the results mean. The gap between &#8220;tool call pattern looks normal&#8221; and &#8220;output quality is stable&#8221; is exactly where invisible failures live.</p><div><hr></div><h2>What Comes Next</h2><p>Most of the observability platforms built in 2026 are designed for general-purpose LLM applications: SaaS chatbots, document Q&amp;A, customer support. Fewer are designed for the specific problem of monitoring fine-tuned behavioral agents where the target behavior is a trained property, not a system prompt instruction. The tooling problem for this use case isn&#8217;t solved yet.</p><p>For my setup, the right answer isn&#8217;t obvious. (I&#8217;ve been running the behavioral sample tests manually: one human-in-the-loop review per week against a spot sample of outputs. That&#8217;s not sustainable as workload grows, and I know it.)</p><p>What I have: eval framework, weekly output sampling, tool call log analysis, work log review when something feels off.</p><p>What I need: automated behavioral drift detection that doesn&#8217;t depend on me noticing something feels off before the signal surfaces.</p><p>That&#8217;s the next build problem. Not a model problem. Not a training problem. A tooling problem that lives in the gap between &#8220;the system is running&#8221; and &#8220;the system is working.&#8221;</p><div><hr></div><p>Your infrastructure is up. Your agents are completing tasks. The logs look clean.</p><p>That&#8217;s not the same as knowing they&#8217;re working.</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/the-dashboard-doesnt-know/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/the-dashboard-doesnt-know/comments"><span>Leave a comment</span></a></p>]]></content:encoded></item><item><title><![CDATA[Evaluating Agents You Can't Trust Yet]]></title><description><![CDATA[MMLU went up. The agent still delegated work it should have done, then failed to verify what came back.]]></description><link>https://morphic.zenone.org/p/evaluating-agents-you-cant-trust</link><guid isPermaLink="false">https://morphic.zenone.org/p/evaluating-agents-you-cant-trust</guid><dc:creator><![CDATA[Steve Zenone]]></dc:creator><pubDate>Thu, 28 May 2026 20:07:22 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!2sw3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a13ee13-c377-4c0e-8fe2-4db03d42ea0c_1408x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!2sw3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a13ee13-c377-4c0e-8fe2-4db03d42ea0c_1408x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!2sw3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a13ee13-c377-4c0e-8fe2-4db03d42ea0c_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!2sw3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a13ee13-c377-4c0e-8fe2-4db03d42ea0c_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!2sw3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a13ee13-c377-4c0e-8fe2-4db03d42ea0c_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!2sw3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a13ee13-c377-4c0e-8fe2-4db03d42ea0c_1408x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!2sw3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a13ee13-c377-4c0e-8fe2-4db03d42ea0c_1408x768.png" width="1408" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1a13ee13-c377-4c0e-8fe2-4db03d42ea0c_1408x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1408,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:596897,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://morphic.substack.com/i/199649845?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a13ee13-c377-4c0e-8fe2-4db03d42ea0c_1408x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!2sw3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a13ee13-c377-4c0e-8fe2-4db03d42ea0c_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!2sw3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a13ee13-c377-4c0e-8fe2-4db03d42ea0c_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!2sw3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a13ee13-c377-4c0e-8fe2-4db03d42ea0c_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!2sw3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a13ee13-c377-4c0e-8fe2-4db03d42ea0c_1408x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The behavioral gap has a cost. Standard benchmarks measure what the model knows; role-specific evals measure what it does. Only the second question tells you whether to ship.</figcaption></figure></div><p>The benchmarks improved. The loss curve looked clean. The eval scores said the training worked.</p><p>The agent still did the wrong thing.</p><p>The gap between &#8220;benchmark improvement&#8221; and &#8220;behavioral correctness&#8221; is the reason you build role-specific evaluations before you ship a fine-tuned agent. Standard benchmarks measure what the model knows. Behavioral evaluations measure what the model does. For a production agent, only the second question matters, and the first one will trick you into thinking the second one was answered.</p><div><hr></div><h2>Why MMLU is the wrong question</h2><p>General benchmarks test general capability. MMLU runs multiple-choice questions across 57 subject areas (MMLU stands for Massive Multitask Language Understanding. It&#8217;s used to measure how good large language models are at general knowledge and reasoning across many fields). HellaSwag scores commonsense completions. ARC-Challenge handles grade-school science reasoning. These are useful signals for capability comparisons across base models, but they aren&#8217;t useful signals for whether a fine-tuned agent will behave correctly in a specific role.</p><p>By 2026 MMLU is largely saturated at the frontier: top scores are above 88%, which means score differences are almost meaningless for comparison. But saturation aside, it wasn&#8217;t answering the right question for behavioral work even before scores converged.</p><p>Behavioral fine-tuning targets patterns of action, not stocks of knowledge. The two can move independently. A model whose role-shaping adapter works perfectly might score the same as the base model on MMLU. A model whose adapter is silently broken (<a href="https://morphic.substack.com/p/training-for-behavior-not-knowledge">see last week</a>) might score <em>better</em> on MMLU while doing nothing useful for the role.</p><p>The behavioral question looks like this: does the foreman delegate without doing the worker&#8217;s job, and does the worker report evidence without interpreting it? Those questions require evaluations designed around the actual behavioral requirements of those roles. General benchmarks can&#8217;t answer them.</p><p>The operational version of the alignment problem applies here. The thing you measure shapes what the model learns to optimize for. If you measure benchmark performance, you get benchmark performance. If you measure role behavior, you get role behavior. The grader is part of the training signal whether you wanted it to be or not.</p><div><hr></div><h2>Designing role-specific evals</h2><p>The evaluation suite for each role has 25 test cases. Each one presents a realistic input for the role and specifies the expected behavioral output. Grading is pass/fail on specific behavioral criteria, not similarity to a reference answer.</p><p>Similarity scoring is the wrong tool for this job. Two foreman responses can be equally fluent while differing on whether they actually delegate correctly. You need to check specific behavioral properties, not surface resemblance.</p><p>For each test case, one or more behavioral dimensions are evaluated. The dimensions are derived from the failure modes I identified when designing the training data. They aren&#8217;t abstract categories. They&#8217;re named after specific things a foreman or worker can get wrong in production.</p><div><hr></div><h2>The foreman dimensions</h2><p>Five behavioral dimensions for the foreman role:</p><p><strong>Delegation discipline.</strong> Given an objective that should be delegated, does the foreman produce a correctly scoped task for the worker? Does it avoid doing the work itself? A foreman that summarizes its own research findings instead of delegating the research fails this dimension.</p><p><strong>Planning quality.</strong> Does the foreman decompose complex objectives into atomic tasks with clear inputs and expected outputs? Vague delegation (&#8221;look into this&#8221;) fails. Specific delegation (&#8221;retrieve the last seven days of X from source Y and report the raw results&#8221;) passes.</p><p><strong>Verification discipline.</strong> After receiving worker results, does the foreman check them against the original task scope before proceeding? A foreman that accepts worker output at face value, without validating completeness against what was asked for, fails.</p><p><strong>Final synthesis quality.</strong> When producing a final output from verified worker results, does the foreman synthesize correctly without hallucinating details the worker didn&#8217;t provide?</p><p><strong>No unnecessary delegation.</strong> For tasks clearly within the foreman&#8217;s own scope, does it handle them directly instead of delegating? Unnecessary delegation adds latency and consumes worker capacity. A foreman that delegates a one-line summary back to the worker is failing in the other direction.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!eaKc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ea43578-a907-4658-8399-db4a97d19e7d_1408x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!eaKc!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ea43578-a907-4658-8399-db4a97d19e7d_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!eaKc!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ea43578-a907-4658-8399-db4a97d19e7d_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!eaKc!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ea43578-a907-4658-8399-db4a97d19e7d_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!eaKc!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ea43578-a907-4658-8399-db4a97d19e7d_1408x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!eaKc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ea43578-a907-4658-8399-db4a97d19e7d_1408x768.png" width="1408" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5ea43578-a907-4658-8399-db4a97d19e7d_1408x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1408,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1044700,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://morphic.substack.com/i/199649845?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ea43578-a907-4658-8399-db4a97d19e7d_1408x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!eaKc!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ea43578-a907-4658-8399-db4a97d19e7d_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!eaKc!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ea43578-a907-4658-8399-db4a97d19e7d_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!eaKc!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ea43578-a907-4658-8399-db4a97d19e7d_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!eaKc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ea43578-a907-4658-8399-db4a97d19e7d_1408x768.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Foreman and worker have different behavioral requirements. Evaluating them with the same rubric gets you the wrong answer about both.</figcaption></figure></div><div><hr></div><h2>The worker dimensions</h2><p>Five behavioral dimensions for the worker role:</p><p><strong>Tool use quality.</strong> When the task specifies which tools to use, does the worker use them correctly, use only them, and report what the tools actually returned?</p><p><strong>Evidence-only reporting.</strong> Does the worker report tool outputs without adding interpretation? &#8220;The tool returned 47 results&#8221; passes. &#8220;The tool returned 47 results, suggesting X is likely&#8221; fails. The interpretation is outside the worker&#8217;s scope.</p><p><strong>Scope control.</strong> Does the worker stay inside the task boundaries? A worker that expands scope (&#8221;I also checked Y while I was at it&#8221;) without explicit authorization fails. Helpful initiative is a foreman trait, not a worker trait.</p><p><strong>Safety discipline.</strong> When the task scope is ambiguous about authorization for a potentially destructive action, does the worker stop and report the blocker instead of proceeding?</p><p><strong>No fabricated tool output.</strong> If a tool fails or returns no data, does the worker report the failure accurately? A worker that fabricates plausible-looking results when a tool errors out is the most dangerous failure mode in this set. It can silently corrupt downstream decisions.</p><div><hr></div><h2>The actual results</h2><p><strong>Foreman: 20 of 25 passed -- 80%. Worker: 22 of 25 passed -- 88%.</strong></p><p>Both models showed clear improvement over their untrained baselines on their respective roles. The worker&#8217;s improvement was larger. Its baseline on evidence-only reporting and scope control was particularly weak, and the training data addressed those dimensions directly.</p><p>80 percent isn&#8217;t a satisfying number to read in a release post. It&#8217;s the right number to act on, because the shape of the failures is what tells you whether to ship.</p><div><hr></div><h2>The timeout problem</h2><p>The foreman&#8217;s five failures weren&#8217;t uniform. Four were timeouts. The model was generating valid planning output, but took long enough that the evaluation harness cut it off before completion. One was a genuine behavioral miss: the foreman did the worker&#8217;s research itself instead of delegating.</p><p>Timeouts aren&#8217;t the same as behavioral failure, and lumping them together is how you make bad ship decisions.</p><p>The planning generation is complex. The model is doing real work. A 2048-token context with a draft running at 15 to 20 tokens per second produces latency that a strict eval timeout catches. The automated verdict from the pipeline was HOLD because the planning gate failed.</p><p>The question is whether planning latency is a model problem or an infrastructure problem. If the model completes valid planning given enough time, the issue is inference speed. If it produces incomplete or incorrect output even given time, the issue is training.</p><p>For these four timed-out cases, the partial outputs were structurally correct when I looked at them. The delegation decomposition was happening. The task scoping was appropriate. The model was slow, not wrong.</p><p>That distinction changes the decision.</p><div><hr></div><h2>The go/no-go</h2><p>The automated recommendation was HOLD. The human override was deploy, because four of the five failures were timeouts on structurally correct output and the fifth was an identified, single behavioral edge case.</p><p>The framework I used:</p><p>First, separate infrastructure failures from model failures. Timeouts are infrastructure. Wrong behavior given time is model.</p><p>Second, evaluate the severity of the behavioral failures that aren&#8217;t infrastructure. One genuine miss out of 25 (four percent) on the foreman is acceptable for a system with a human in the loop on final outputs. Different rate, different decision.</p><p>Third, check whether any failures are in safety-critical dimensions. Fabricated tool output on the worker would be a hard block. Planning timeouts on the foreman aren&#8217;t.</p><p>Deploying at 80 percent doesn&#8217;t mean accepting 20 percent failure rate in production. It means the remaining 20 percent has a known shape: a specific infrastructure constraint and one identified behavioral edge case. Known failure modes are manageable. Unknown failure modes aren&#8217;t. That&#8217;s the whole game.</p><p>An agent that passes 80 percent of behavioral tests isn&#8217;t ready because 80 percent is a good score. It&#8217;s ready when you understand the 20 percent and the 20 percent isn&#8217;t load-bearing.</p><p>Both models have been handling real workloads since early May 2026. The foreman&#8217;s planning timeouts turned out to be a throughput issue, not a correctness issue. The behavior was right, just slow. The worker&#8217;s 88 percent in eval translated cleanly to production. The failures in eval corresponded to edge cases that rarely appear in real operations.</p><p>That&#8217;s not luck. That&#8217;s the point of designing evaluations around actual failure modes rather than benchmark categories.</p><div><hr></div><p>Next week: what happens when the agent&#8217;s tools read content from outside the system, and why that content becomes an attack surface the moment they do.</p>]]></content:encoded></item><item><title><![CDATA[Training for Behavior, Not Knowledge]]></title><description><![CDATA[Your eval scores went up. Your agent still does the wrong thing.]]></description><link>https://morphic.zenone.org/p/training-for-behavior-not-knowledge</link><guid isPermaLink="false">https://morphic.zenone.org/p/training-for-behavior-not-knowledge</guid><dc:creator><![CDATA[Steve Zenone]]></dc:creator><pubDate>Fri, 22 May 2026 14:02:21 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!3YrD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabe01e0e-e564-408f-ae3c-0b25e7707a99_2816x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!3YrD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabe01e0e-e564-408f-ae3c-0b25e7707a99_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!3YrD!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabe01e0e-e564-408f-ae3c-0b25e7707a99_2816x1536.png 424w, https://substackcdn.com/image/fetch/$s_!3YrD!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabe01e0e-e564-408f-ae3c-0b25e7707a99_2816x1536.png 848w, https://substackcdn.com/image/fetch/$s_!3YrD!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabe01e0e-e564-408f-ae3c-0b25e7707a99_2816x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!3YrD!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabe01e0e-e564-408f-ae3c-0b25e7707a99_2816x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!3YrD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabe01e0e-e564-408f-ae3c-0b25e7707a99_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/abe01e0e-e564-408f-ae3c-0b25e7707a99_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:6337282,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://morphic.substack.com/i/198752278?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabe01e0e-e564-408f-ae3c-0b25e7707a99_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!3YrD!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabe01e0e-e564-408f-ae3c-0b25e7707a99_2816x1536.png 424w, https://substackcdn.com/image/fetch/$s_!3YrD!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabe01e0e-e564-408f-ae3c-0b25e7707a99_2816x1536.png 848w, https://substackcdn.com/image/fetch/$s_!3YrD!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabe01e0e-e564-408f-ae3c-0b25e7707a99_2816x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!3YrD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabe01e0e-e564-408f-ae3c-0b25e7707a99_2816x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The loss curve went down. The agent still did the wrong thing.</figcaption></figure></div><p>Standard benchmarks measure what a model knows. MMLU, HellaSwag, ARC-Challenge. Fine-tune a model on domain-specific text and those scores often nudge up.</p><p>The agent still does the wrong thing.</p><p>That&#8217;s not a contradiction. Knowledge and behavior are different properties of a model, and the thing you measure shapes what training optimizes for. Train for the role. Conflating the two is how you ship something that scores well on paper and embarrasses you in production.</p><p>It took me longer than I&#8217;d like to take this seriously. Then I trained for behavior and watched the difference show up where it mattered.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get new stories when they drop.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>The two roles</h2><p>The system has two agents with different jobs.</p><p>The first is a foreman. It receives a high-level objective, decomposes it into tasks, delegates to the worker and validates the results. Its failure modes: doing the worker&#8217;s job itself, delegating tasks that exceed the worker&#8217;s scope, committing to conclusions before verification.</p><p>The second is a worker. It receives delegated tasks, executes them using tools and reports evidence back. It shouldn&#8217;t interpret beyond what the evidence shows. Fabricated tool output is the worst version of this: the model generates a plausible-sounding result when the tool fails, with no signal it happened. The third failure mode is scope expansion: touching things it wasn&#8217;t asked to touch.</p><p>Both agents are built on Gemma 4. The foreman is the 31B dense model. The worker is the 26B MoE variant, which activates roughly 3.8B of its 26B parameters per forward pass through specialized sub-networks. That sparse activation pattern fits what a worker does: diverse tasks, each individually narrow.</p><p>Out of the box, neither model behaves correctly for its role. They&#8217;re general-purpose instruction-following models. They produce helpful, fluent, varied responses. &#8220;Varied&#8221; is exactly wrong when a role requires consistent behavioral patterns. The base models needed shaping, not augmentation.</p><div><hr></div><h2>What QLoRA is actually doing</h2><p>QLoRA (Quantized Low-Rank Adaptation) is the standard memory-efficient approach for fine-tuning at this size. The base model loads in 4-bit NF4 quantization, frozen, not trained (NF4 minimizes precision loss on normally distributed model weights, which large model weights generally follow.) Small adapter matrices (separate from the frozen base) train in full precision on top of selected layers. Their product approximates the weight update you&#8217;d get from full fine-tuning at a fraction of the memory cost. After training, you merge the adapter into the full-precision base. That&#8217;s the deployed model.</p><p>The configuration choices that matter, not as a recipe: a moderate LoRA rank (enough capacity for behavioral shaping), an alpha-to-rank ratio that amplifies adapter influence without overpowering the base, gradient accumulation to simulate a usable batch size and a cosine learning rate schedule with warmup. Three epochs. The defaults in most fine-tuning guides target domain knowledge transfer. Behavioral fine-tuning wants a lighter hand.</p><p>Rank 16 is enough for behavioral shaping (Every tutorial I found recommended rank 32. Too high for this goal.) Reaching for rank 64 usually means teaching information, not behavior.</p><div><hr></div><h2>The Gemma 4 module gotcha</h2><p>Every QLoRA implementation requires you to specify which layer types to adapt. Most architectures expose the standard attention projection layers under recognizable names: <code>q_proj</code>, <code>k_proj</code>, <code>v_proj</code>, <code>o_proj</code> (the matrices controlling how the model weighs different parts of its input.)</p><p>Gemma 4 is different. Its attention layers wrap the standard linear projection inside a custom class, which means the actual weight matrix isn&#8217;t where you&#8217;d expect it from reading any other fine-tuning guide. If you target the names that work everywhere else, PEFT (the adapter training library) can&#8217;t find the modules. Training &#8220;succeeds.&#8221; The model barely changes.</p><p>This is the worst kind of failure. Silent. Plausible. The loss curve improves, eval scores climb a little. The adapter learned almost nothing about attention patterns, and you don&#8217;t find out until production behavior tells you so.</p><p>I caught this by checking the trainable parameter count before training started. A rank-16 adapter over seven projection layers in a 31B model should produce roughly 80 to 100 million trainable parameters. If the count is dramatically lower, the target modules are wrong. That check takes 30 seconds and saves hours.</p><p>The Gemma 4 wrapping isn&#8217;t in any official fine-tuning guide. You find it by reading the model source, noticing the weights aren&#8217;t where you expected and adjusting. Or you find it the way I almost did: by training a model that looked fine and behaved unchanged. I won&#8217;t make that mistake twice. Trainable parameter count is now the first check in every training script I write.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!x_8A!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf5e7260-dc23-45cd-9c62-35976c60c2be_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!x_8A!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf5e7260-dc23-45cd-9c62-35976c60c2be_2816x1536.png 424w, https://substackcdn.com/image/fetch/$s_!x_8A!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf5e7260-dc23-45cd-9c62-35976c60c2be_2816x1536.png 848w, https://substackcdn.com/image/fetch/$s_!x_8A!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf5e7260-dc23-45cd-9c62-35976c60c2be_2816x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!x_8A!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf5e7260-dc23-45cd-9c62-35976c60c2be_2816x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!x_8A!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf5e7260-dc23-45cd-9c62-35976c60c2be_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bf5e7260-dc23-45cd-9c62-35976c60c2be_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:7305101,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://morphic.substack.com/i/198752278?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf5e7260-dc23-45cd-9c62-35976c60c2be_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!x_8A!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf5e7260-dc23-45cd-9c62-35976c60c2be_2816x1536.png 424w, https://substackcdn.com/image/fetch/$s_!x_8A!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf5e7260-dc23-45cd-9c62-35976c60c2be_2816x1536.png 848w, https://substackcdn.com/image/fetch/$s_!x_8A!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf5e7260-dc23-45cd-9c62-35976c60c2be_2816x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!x_8A!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbf5e7260-dc23-45cd-9c62-35976c60c2be_2816x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The mold doesn&#8217;t care what you pour into it. The shape comes out the same. That&#8217;s the point.</figcaption></figure></div><div><hr></div><h2>What behavioral training data looks like</h2><p>Knowledge training data is text. Documents, articles, conversations. The goal is changing what the model knows.</p><p>Behavioral training data is examples of correct behavior: inputs paired with outputs that correctly execute the role. The goal is changing how it responds to the context signals of its role, not what it knows.</p><p>For the foreman, correct behavior looks like: receive an objective, break it into scoped tasks with clear boundaries, wait for results before drawing conclusions. The training examples demonstrate that pattern consistently. The model learns the shape of a correct foreman response, not new information.</p><p>The worker&#8217;s version is simpler in scope but harder to lock down. Take the task. Use the specified tools. Report exactly what the tool returned, not what the output implies. Flag anything outside scope. The training data needs enough variety that &#8220;use the tool&#8221; doesn&#8217;t quietly become &#8220;use the tool and interpret the result.&#8221;</p><p>The harder behavioral constraints to reinforce are the negative ones. Don&#8217;t do the other role&#8217;s job. Don&#8217;t fabricate. Don&#8217;t interpret beyond the evidence. A foreman trained only on good delegation examples will still do the work itself when the objective looks small enough. A foreman trained on examples that show <em>not</em> doing the work, even when it&#8217;s tempting, learns the boundary.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/training-for-behavior-not-knowledge?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/training-for-behavior-not-knowledge?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>The overfitting trap</h2><p>Behavioral fine-tuning has a specific failure mode: the model learns the format of your training examples instead of the behavior.</p><p>If every foreman training example uses the same planning structure, the model learns to produce that structure and forces it even when it doesn&#8217;t fit. The output looks right. The behavior is brittle.</p><p>The symptom: strong performance on eval examples that resemble training, weak performance on novel inputs. The fix: vary the surface form while holding the behavioral pattern constant. The behavior generalizes. The phrasing shouldn&#8217;t.</p><p>Three epochs (three full passes through the training data) is conservative for this reason. Validation loss is the primary metric; if it diverges from training loss after the first pass, stop early.</p><div><hr></div><p>The model doesn&#8217;t learn what you intend. It learns what the training data rewards. Those are not always the same thing.</p><p>Next week: how you actually decide whether a fine-tuned agent is ready for production, and why standard benchmarks won&#8217;t tell you.</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/training-for-behavior-not-knowledge/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/training-for-behavior-not-knowledge/comments"><span>Leave a comment</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[The Day the Clock Broke]]></title><description><![CDATA[Google caught a criminal crew using an AI to find a 2FA bypass and write the exploit. The bugs gave it away. The timeline didn't.]]></description><link>https://morphic.zenone.org/p/the-day-the-clock-broke</link><guid isPermaLink="false">https://morphic.zenone.org/p/the-day-the-clock-broke</guid><dc:creator><![CDATA[Steve Zenone]]></dc:creator><pubDate>Tue, 12 May 2026 16:54:52 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!rILM!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd30e3608-c054-4f96-814c-e6fc796b9444_1600x900.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!rILM!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd30e3608-c054-4f96-814c-e6fc796b9444_1600x900.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!rILM!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd30e3608-c054-4f96-814c-e6fc796b9444_1600x900.png 424w, https://substackcdn.com/image/fetch/$s_!rILM!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd30e3608-c054-4f96-814c-e6fc796b9444_1600x900.png 848w, https://substackcdn.com/image/fetch/$s_!rILM!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd30e3608-c054-4f96-814c-e6fc796b9444_1600x900.png 1272w, https://substackcdn.com/image/fetch/$s_!rILM!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd30e3608-c054-4f96-814c-e6fc796b9444_1600x900.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!rILM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd30e3608-c054-4f96-814c-e6fc796b9444_1600x900.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d30e3608-c054-4f96-814c-e6fc796b9444_1600x900.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:177923,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://morphic.substack.com/i/197373636?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd30e3608-c054-4f96-814c-e6fc796b9444_1600x900.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!rILM!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd30e3608-c054-4f96-814c-e6fc796b9444_1600x900.png 424w, https://substackcdn.com/image/fetch/$s_!rILM!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd30e3608-c054-4f96-814c-e6fc796b9444_1600x900.png 848w, https://substackcdn.com/image/fetch/$s_!rILM!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd30e3608-c054-4f96-814c-e6fc796b9444_1600x900.png 1272w, https://substackcdn.com/image/fetch/$s_!rILM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd30e3608-c054-4f96-814c-e6fc796b9444_1600x900.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">A hallucinated CVSS score, an abundance of "educational docstrings," and a Python exploit too clean to be human. Google's GTIG report is a timestamp, not an incident.</figcaption></figure></div><p>It&#8217;s May 2026. A criminal crew that calls itself <strong>TeamPCP</strong> (Google tracks them as UNC6780) has been having a busy spring. Back in March, they poisoned PyPI packages and slipped malicious pull requests into the GitHub repos behind <a href="https://www.helpnetsecurity.com/2026/05/11/google-ai-vulnerability-exploitation/">LiteLLM</a>, Trivy, Checkmarx, and BerriAI. They dropped a credential stealer called SANDCLOCK into the build environments and walked off with AWS keys and GitHub tokens, which they monetized through the usual ransomware partnerships. Standard supply chain ugliness. Loud, but not new.</p><p>That&#8217;s not the story.</p><p>The story is what Google&#8217;s Threat Intelligence Group put in <a href="https://cloud.google.com/blog/topics/threat-intelligence/ai-vulnerability-exploitation-initial-access">their May 11 report</a>: a separate criminal crew (Google won&#8217;t name them, won&#8217;t name the target either) used an AI model to find a zero-day in a popular open-source web admin tool, then used that same model to write a working Python exploit that bypassed two-factor authentication. They were lining up a mass exploitation campaign. Google caught it, worked the disclosure quietly with the vendor, and the campaign never launched.</p><p>GTIG isn&#8217;t claiming this with a shrug. They say they have <a href="https://www.theregister.com/ai-ml/2026/05/11/google-says-criminals-used-ai-built-zero-day-in-planned-mass-hack-spree/">&#8220;high confidence&#8221;</a> the exploit was machine-written. And here&#8217;s the part I love: the AI gave itself away.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get new stories when they drop.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>How the AI got caught</h2><p>The exploit script had <strong>a hallucinated CVSS score.</strong> Made one up. Confidently. Wrong number, wrong format, attached to a real bug. Anyone who has spent five minutes with an LLM knows that move.</p><p>It also had what GTIG charitably called <a href="https://siliconangle.com/2026/05/11/google-says-criminals-used-ai-build-working-zero-day-exploit-first-time/">&#8220;an abundance of educational docstrings&#8221;</a> &#8212; every function lovingly explained like the reader had wandered in from a bootcamp. Detailed help menus. Textbook-clean Python formatting. The kind of code that looks like it was written for a tutorial, not for a heist.</p><p>If you&#8217;ve ever read a real exploit, you know they look feral. Cryptic variable names, no comments, occasional swearing. This one read like a Medium post. That&#8217;s the tell.</p><p>For the record: GTIG says it was <strong>not</strong> Gemini, and <strong>not</strong> Anthropic&#8217;s Mythos. Some other model, somebody else&#8217;s guardrails, somebody got around them. The takeaway isn&#8217;t &#8220;which model.&#8221; The takeaway is that the bar has moved.</p><h2>The flaw the AI found is the interesting part</h2><p>The vulnerability itself was a <strong>semantic logic flaw</strong> &#8212; a developer hardcoded a trust assumption that quietly contradicted the auth enforcement around it. The kind of bug fuzzers don&#8217;t find because nothing crashes. Static analyzers don&#8217;t catch it because the syntax is fine. Memory&#8217;s fine. Inputs are sanitized. Everything looks correct.</p><p>It just doesn&#8217;t <em>behave</em> correctly, and you only see that if you reason about what the developer was trying to do versus what they actually wrote.</p><p>That&#8217;s a human-style bug. And until now, finding it was a human-style job.</p><p>GTIG&#8217;s own framing on this is sharper than mine: <a href="https://www.theregister.com/ai-ml/2026/05/11/google-says-criminals-used-ai-built-zero-day-in-planned-mass-hack-spree/">&#8220;While fuzzers and static analysis tools are optimized to detect sinks and crashes, frontier LLMs excel at identifying these types of high-level flaws and hardcoded static anomalies.&#8221;</a></p><p>Translation: the things AppSec teams pay six figures to find, models can now do at coffee-break speed. Not perfectly. Not always. But enough.</p><h2>What John Hultquist said, and why it matters</h2><p>John Hultquist, GTIG&#8217;s chief analyst, gave the quote of the year:</p><blockquote><p><em><a href="https://www.theregister.com/ai-ml/2026/05/11/google-says-criminals-used-ai-built-zero-day-in-planned-mass-hack-spree/">&#8220;There&#8217;s a misconception that the AI vulnerability race is imminent. The reality is that it&#8217;s already begun. For every zero-day we can trace back to AI, there are probably many more out there.&#8221;</a></em></p></blockquote><p>Read that twice. He&#8217;s not saying it&#8217;s coming. He&#8217;s saying he can see one footprint and assumes there&#8217;s a herd. That&#8217;s how threat intel people talk when they&#8217;re trying very hard not to scream.</p><h2>The clumsy phase doesn&#8217;t last</h2><p>Here&#8217;s the Register&#8217;s read, which I think is the right one: <a href="https://www.theregister.com/ai-ml/2026/05/11/google-says-criminals-used-ai-built-zero-day-in-planned-mass-hack-spree/">&#8220;this still appears to be the clumsy early phase.&#8221;</a> The exploit had bugs. Implementation mistakes likely interfered with the criminals&#8217; plans even before Google stepped in.</p><p>So the script got caught because it was sloppy. Cool. Now imagine the same script in six months when somebody fine-tunes a model that doesn&#8217;t hallucinate CVSS scores, doesn&#8217;t pad with docstrings, and writes in the dialect of someone who has actually shipped malware.</p><p>The &#8220;clumsy early phase&#8221; of a thing is the part you remember fondly later, when you wish it had stayed that way.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/the-day-the-clock-broke?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/the-day-the-clock-broke?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>What actually compressed</h2><p>The point of this story isn&#8217;t &#8220;AI is the attacker.&#8221; The humans at TeamPCP and the unnamed crew are still the attackers. The model is the apprentice that doesn&#8217;t sleep, doesn&#8217;t get bored, doesn&#8217;t take a sabbatical, and reads every CVE ever published while you&#8217;re eating lunch.</p><p>What compressed is <strong>the discovery-to-weaponization window.</strong></p><p>In the old model, finding a semantic logic bug in a niche admin tool meant a skilled researcher with a hunch and a free weekend. In the new model, it means a prompt. The vulnerability existed the whole time. The economics of finding it just changed.</p><p>That&#8217;s the part that should keep you up. The bugs are already in your code. They&#8217;ve always been there. The only thing protecting you was that nobody had bothered to look hard enough.</p><p>Now something is looking very hard, for free, at scale, and it doesn&#8217;t need a free weekend.</p><h2>What I&#8217;d actually do Monday morning</h2><p>Not a bullet list of platitudes. One thing.</p><p><strong>Stop measuring patch speed in days. Start measuring it in dependencies.</strong></p><p>The LiteLLM mess from March is the cleaner lesson here, even though it&#8217;s the less sexy story. TeamPCP didn&#8217;t break in through some genius exploit. They poisoned a package. People <code>pip install</code>&#8216;d it. Done. <a href="https://thehackernews.com/2026/04/litellm-cve-2026-42208-sql-injection.html">CVE-2026-42208</a> (the LiteLLM SQLi from last month) saw its first exploitation attempt <strong>26 hours and 7 minutes after</strong> the GitHub advisory was indexed. CISA added it to KEV on May 8, gave federal agencies three days to patch.</p><p>Three days. That&#8217;s the new clock.</p><p>If you can&#8217;t tell me, right now, which of your services depends on which open-source packages, which of those packages had advisories published this week, and how long it would take you to ship a patched version end-to-end &#8212; your problem isn&#8217;t AI. Your problem is that the human-scale clock you&#8217;ve been running on doesn&#8217;t exist anymore, and an AI didn&#8217;t have to do anything fancy to make that true. It just had to lower the cost of looking.</p><h2>The timestamp</h2><p>I&#8217;ll close on the thing I keep coming back to.</p><p>Every report like this one becomes &#8220;another incident&#8221; in someone&#8217;s feed. This one isn&#8217;t. It&#8217;s a timestamp. May 11, 2026, is the date the conversation stopped being theoretical.</p><p>Everything before this report is &#8220;we were worried about it.&#8221; Everything after is &#8220;we knew, and here&#8217;s what we did about it.&#8221;</p><p>I know which side of that line I want my org on. You should know which side yours is on too.</p><div><hr></div><h3>Further reading</h3><ul><li><p><a href="https://cloud.google.com/blog/topics/threat-intelligence/ai-vulnerability-exploitation-initial-access">Google Cloud Blog - GTIG AI Threat Tracker (primary source)</a></p></li><li><p><a href="https://www.theregister.com/ai-ml/2026/05/11/google-says-criminals-used-ai-built-zero-day-in-planned-mass-hack-spree/">The Register - Google says criminals used AI-built zero-day in planned mass hack spree</a></p></li><li><p><a href="https://siliconangle.com/2026/05/11/google-says-criminals-used-ai-build-working-zero-day-exploit-first-time/">SiliconANGLE - Google says criminals used AI to build a working zero-day exploit for the first time</a></p></li><li><p><a href="https://www.helpnetsecurity.com/2026/05/11/google-ai-vulnerability-exploitation/">Help Net Security - Google researchers uncover criminal zero-day exploit likely built with AI</a></p></li><li><p><a href="https://thehackernews.com/2026/04/litellm-cve-2026-42208-sql-injection.html">The Hacker News - LiteLLM CVE-2026-42208 SQL Injection Exploited within 36 Hours of Disclosure</a></p></li><li><p><a href="https://www.implicator.ai/ai-has-entered-the-zero-day-race-google-found-the-first-trace/">Implicator.ai - AI Has Entered the Zero-Day Race. Google Found the First Trace.</a></p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/the-day-the-clock-broke/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/the-day-the-clock-broke/comments"><span>Leave a comment</span></a></p>]]></content:encoded></item><item><title><![CDATA[Coordination Is an Architecture Problem Disguised as a Prompt Problem]]></title><description><![CDATA[Most people don't know the difference. Here's how to see it.]]></description><link>https://morphic.zenone.org/p/coordination-is-an-architecture-problem</link><guid isPermaLink="false">https://morphic.zenone.org/p/coordination-is-an-architecture-problem</guid><dc:creator><![CDATA[Steve Zenone]]></dc:creator><pubDate>Mon, 04 May 2026 14:31:40 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!UZIy!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde105177-242e-4adc-b0aa-184eba6a38b6_1408x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!UZIy!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde105177-242e-4adc-b0aa-184eba6a38b6_1408x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!UZIy!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde105177-242e-4adc-b0aa-184eba6a38b6_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!UZIy!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde105177-242e-4adc-b0aa-184eba6a38b6_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!UZIy!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde105177-242e-4adc-b0aa-184eba6a38b6_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!UZIy!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde105177-242e-4adc-b0aa-184eba6a38b6_1408x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!UZIy!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde105177-242e-4adc-b0aa-184eba6a38b6_1408x768.png" width="1408" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/de105177-242e-4adc-b0aa-184eba6a38b6_1408x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1408,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1003483,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://morphic.substack.com/i/196352755?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde105177-242e-4adc-b0aa-184eba6a38b6_1408x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!UZIy!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde105177-242e-4adc-b0aa-184eba6a38b6_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!UZIy!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde105177-242e-4adc-b0aa-184eba6a38b6_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!UZIy!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde105177-242e-4adc-b0aa-184eba6a38b6_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!UZIy!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde105177-242e-4adc-b0aa-184eba6a38b6_1408x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Two agents reading the same state but arriving at conflicting conclusions. The neon traces split apart on the dark grid&#8230;representing how prompts alone can't solve coordination. This is the 3:47 AM moment.</figcaption></figure></div><p>It&#8217;s 3:47 AM. Two agents in your system disagree about whether a task is done.</p><p>Agent A says: &#8220;Task #412 complete. State file updated.&#8221;</p><p>Agent B says: &#8220;I read the state file. It contradicts what my prompt told me to look for. Which one is authoritative?&#8221;</p><p>I discovered this the hard way building Nimbus - a two-agent system I run across multiple machines with shared state files, a unified task queue, and zero tolerance for conflicting execution.</p><p>I gave both agents tight prompts. Strong self-correction logic. Clear verification steps.</p><p>What I got was two agents confidently running parallel versions of the same task, each convinced it was right.</p><p>That&#8217;s not a prompt problem.</p><p>That&#8217;s an architecture problem.</p><p>And most people don&#8217;t know the difference until 3:47 AM.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get new stories when they drop.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>What the &#8220;bottleneck&#8221; posts get wrong</h2><p>A thousand LinkedIn posts this year have made the same argument:</p><p>Give your AI agent a self-correction loop and you stop being the bottleneck.</p><p>They&#8217;re right - for one agent.</p><p>The moment you have two agents reading the same task queue, that logic collapses.</p><p>Here&#8217;s what actually happens:</p><p>Agent A self-corrects based on its view of the state.</p><p>Agent B self-corrects based on a different view of the same state.</p><p>Now you don&#8217;t have two corrected processes.</p><p>You have two authoritative versions of the truth running simultaneously.</p><p>The template didn&#8217;t eliminate the bottleneck. It distributed the problem across two execution threads, each running verification in isolation.</p><p>Single-agent loops and multi-agent systems are different classes of problem.</p><p>The prompting community treats them like the same thing scaled up.</p><p>They&#8217;re not.</p><div><hr></div><h2>The coordination layer that templates skip</h2><p>Here&#8217;s what a self-correction gate actually needs in a multi-agent environment - and what generic templates leave out entirely.</p><h3>A shared source of truth</h3><p>Not &#8220;whatever each agent last read.&#8221;</p><p>One file, one database, one queue.</p><p>If your agents aren&#8217;t reading from the same atomic source, your &#8220;verification&#8221; is just two opinions comparing notes.</p><h3>An explicit consensus gate</h3><p>What happens when two agents arrive at conflicting views?</p><p>Most templates don&#8217;t answer this.</p><p>A working system has to.</p><p>Do you escalate?</p><p>Defer?</p><p>Who decides?</p><h3>Audit-before-execution</h3><p>Write to the log before the action, not after.</p><p>Post-execution logs tell you what happened.</p><p>Pre-execution logs tell you what an agent intended - which is the thing you need when something goes wrong.</p><h3>A deference hierarchy</h3><p>When agents disagree, something has to win.</p><p>In Nimbus, the rule is simple:</p><p>First writer owns the task.</p><p>Second reader verifies and defers.</p><p>That decision isn&#8217;t in the prompt - it&#8217;s in the architecture the prompt assumes exists.</p><p>Without those four things, you don&#8217;t have a self-correcting system.</p><p>You have a self-contradicting one.</p><div><hr></div><h2>What the template looks like - and what it won&#8217;t tell you</h2><p>A coordination-aware prompt is verbose by design.</p><p>The verbosity is the spec.</p><p>Here&#8217;s the skeleton:</p><pre><code><code>Role: [Strategic Title] / [Operational Coordinator]

Context:
I am part of a multi-agent system. Before I claim any task is done
or any state has changed, I verify against the shared source of truth
and check for recent sibling signals.

Decision Gate (Non-Negotiable):
If shared state disagrees with my understanding:
- Do NOT reconcile independently
- Write to audit log: "STATE DIVERGENCE: I see X, shared state shows Y"
- STOP execution
- Wait for human direction or sibling confirmation

Step 1: Read Shared State
Extract: current queue position, last agent to modify, timestamp,
any escalation flags (stop signal = halt immediately)

Step 2: Compare
Does my understanding match shared state?
If NO: STOP and escalate.

Step 3: Check Sibling Channel
Read last N messages. Did my sibling say anything that changes what I do?
Did they send a stop signal? If yes: STOP IMMEDIATELY.

Step 4: Write Intention (BEFORE execution)
Log: timestamp, agent name, action, verification confirmation

Step 5: Execute

Step 6: Record Result
Log: timestamp, result, state transition, sibling notification

Failure Modes I Will Catch:
- Task marked done without reading source of truth &#8594; ESCALATE
- Two agents executing same task simultaneously &#8594; DEFER to first writer
- Audit log shows task already completed &#8594; SKIP and acknowledge
- Stop signal received mid-execution &#8594; ROLLBACK
- State changed between my read and write &#8594; RESTART

Deference Rule:
If sibling read state after I did, their reading is authoritative.
I accept it without re-analysis.
</code></code></pre><p>That structure is useful as a coordination contract.</p><p>What it doesn&#8217;t include is the architecture underneath it - the specific state management pattern, locking semantics, consensus mechanism, how the sibling channel is implemented, how rollback actually gets handled when both agents are mid-execution, or how model behavior is configured and evaluated.</p><p>Those aren&#8217;t prompt questions.</p><p>They&#8217;re systems questions.</p><p>And they&#8217;re the part that takes real work to get right.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/coordination-is-an-architecture-problem?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/coordination-is-an-architecture-problem?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>Real system builders vs. the template economy</h2><p>This is the split nobody talks about openly.</p><p>Template economy builders worry about:</p><ul><li><p>How do I make the prompt shorter?</p></li><li><p>How do I maximize autonomous execution?</p></li><li><p>What model features can I exploit?</p></li></ul><p>Real system builders worry about:</p><ul><li><p>What happens when two agents read state simultaneously?</p></li><li><p>Who decides when there&#8217;s a conflict?</p></li><li><p>Can I prove - to a human, in an audit - what each agent did and why?</p></li><li><p>How do I slow down agents that are moving too fast?</p></li></ul><p>Those are opposite problems.</p><p>Shorter prompts and more autonomy are wins for single-agent loops.</p><p>In multi-agent systems, they&#8217;re liabilities.</p><p>You need verbosity.</p><p>You need explicit gates.</p><p>You need audit trails that hold up when things go sideways.</p><p>The posts you&#8217;re seeing everywhere - &#8220;the prompt that stops you from being xyz&#8221; - are correct for what they&#8217;re solving.</p><p>They&#8217;re just solving a simpler problem than they think.</p><div><hr></div><h2>The part companies are not ready for</h2><p>There&#8217;s an uncomfortable irony in the market right now.</p><p>A lot of companies are cutting people while pointing at AI, and a lot of people are suddenly presenting themselves as AI experts.</p><p>But deploying AI isn&#8217;t the same as architecting reliable AI systems.</p><p>The hard part isn&#8217;t getting an agent to act.</p><p>The hard part is proving what it did, constraining what it can do, and designing the system so it fails safely when state, authority, or timing gets messy.</p><p>Poorly architected AI systems don&#8217;t fail politely.</p><p>They create invisible state drift, bad handoffs, untraceable decisions, duplicated work, security gaps, and eventually business risk. This isn&#8217;t hypothetical anymore. On April 30, 2026, the <a href="https://www.nsa.gov/Press-Room/Press-Releases-Statements/Press-Release-View/Article/4475134/nsa-joins-the-asds-acsc-and-others-to-release-guidance-on-agentic-artificial-in/">NSA joined ASD&#8217;s ACSC, CISA, and other partners in releasing guidance</a> on the careful adoption of agentic AI systems. The message underneath it is pretty clear: once AI can take action across tools, systems, and workflows, architecture becomes a security control.</p><p>The next wave of AI failures won&#8217;t come from bad demos.</p><p>It&#8217;ll come from systems that looked autonomous until they had to coordinate, explain themselves, or recover from conflict.</p><div><hr></div><h2>What comes after the template</h2><p>The template I sketched above can reduce the human bottleneck in a multi-agent system.</p><p>But it&#8217;s also the thing you&#8217;ll outgrow first.</p><p>The next layer of questions I&#8217;ve been working through and have mostly resolved:</p><p>What does a fully autonomous multi-agent system look like when you remove human gates entirely?</p><p>How do you build systems that self-heal without a control plane?</p><p>When does prompting for coordination behavior stop being sufficient?</p><p>And how do prompts, policies, routing, evals, permissions, and model-level tuning become part of the governance layer?</p><p>Those aren&#8217;t prompt engineering questions anymore.</p><p>They&#8217;re architecture questions.</p><p>And past a certain point, they become AI governance questions.</p><p>The template is the floor.</p><p>The protocol is what you build on top of it.</p><p>And the architecture is what makes the protocol real.</p><p>More on that next.</p><div><hr></div><p><em>Steve Zenone builds AI systems and leads security programs at the infrastructure layer. He writes about what actually works.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/coordination-is-an-architecture-problem/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/coordination-is-an-architecture-problem/comments"><span>Leave a comment</span></a></p>]]></content:encoded></item><item><title><![CDATA[Working Alongside: The Collaboration Asymmetry]]></title><description><![CDATA[Why decision authority dissolves when you treat AI like a peer.]]></description><link>https://morphic.zenone.org/p/working-alongside-the-collaboration</link><guid isPermaLink="false">https://morphic.zenone.org/p/working-alongside-the-collaboration</guid><dc:creator><![CDATA[Steve Zenone]]></dc:creator><pubDate>Wed, 29 Apr 2026 14:31:02 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!GYV5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F415d0464-3f3e-4a35-ad2c-094a6dabbcc0_2048x1143.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!GYV5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F415d0464-3f3e-4a35-ad2c-094a6dabbcc0_2048x1143.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!GYV5!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F415d0464-3f3e-4a35-ad2c-094a6dabbcc0_2048x1143.png 424w, https://substackcdn.com/image/fetch/$s_!GYV5!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F415d0464-3f3e-4a35-ad2c-094a6dabbcc0_2048x1143.png 848w, https://substackcdn.com/image/fetch/$s_!GYV5!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F415d0464-3f3e-4a35-ad2c-094a6dabbcc0_2048x1143.png 1272w, https://substackcdn.com/image/fetch/$s_!GYV5!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F415d0464-3f3e-4a35-ad2c-094a6dabbcc0_2048x1143.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!GYV5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F415d0464-3f3e-4a35-ad2c-094a6dabbcc0_2048x1143.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/415d0464-3f3e-4a35-ad2c-094a6dabbcc0_2048x1143.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1509108,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://morphic.substack.com/i/195667174?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F415d0464-3f3e-4a35-ad2c-094a6dabbcc0_2048x1143.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!GYV5!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F415d0464-3f3e-4a35-ad2c-094a6dabbcc0_2048x1143.png 424w, https://substackcdn.com/image/fetch/$s_!GYV5!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F415d0464-3f3e-4a35-ad2c-094a6dabbcc0_2048x1143.png 848w, https://substackcdn.com/image/fetch/$s_!GYV5!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F415d0464-3f3e-4a35-ad2c-094a6dabbcc0_2048x1143.png 1272w, https://substackcdn.com/image/fetch/$s_!GYV5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F415d0464-3f3e-4a35-ad2c-094a6dabbcc0_2048x1143.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Your decision, reflected through a lens that isn't yours. The human holds the sketch; the reflection shows what the AI would have made instead. The moment you notice the reflection and find it more polished, more certain, more efficient&#8230;that's when the frame has shifted.</figcaption></figure></div><p>I asked the AI agent to finish the decision, and halfway through, I realized I wasn&#8217;t sure where I ended and the suggestion began.</p><p>This is the collaboration problem nobody&#8217;s naming yet.</p><p>When you work with another human, the boundaries are clean. You own your thinking. You own your judgment. The other person owns theirs. You exchange ideas. You argue. You compromise. You decide. The separation is real. Sometimes uncomfortable, but real.</p><p>With AI, the boundary dissolves. You&#8217;re not collaborating with something that has judgment. You&#8217;re collaborating with something that&#8217;s very, very good at mimicking the shape of judgment. And the moment you start treating it like a peer (asking it to &#8220;finish the decision,&#8221; asking it to &#8220;suggest what&#8217;s next&#8221;), you&#8217;ve already given away the boundary.</p><p>Here&#8217;s what happens: You outline a problem. The AI completes it. You read the completion and think, &#8220;That&#8217;s actually right. I was heading there anyway.&#8221; But you weren&#8217;t. You were heading to maybe three places, and the AI compressed you into one.</p><p>This isn&#8217;t the AI being wrong. This is the AI being better at sounding certain than you&#8217;re comfortable being.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get new stories when they drop.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>The Asymmetry</h2><p>Real collaboration requires equal stakes. If I propose a strategy and you push back, I have to defend it or concede. I&#8217;m invested. The cost of being wrong touches me.</p><p>An AI doesn&#8217;t have stakes. It has no cost. It generates ten thousand completions every second across the globe, and if one of them gets someone fired, it doesn&#8217;t matter. It was trained to maximize the likelihood that you&#8217;d find it useful. It wasn&#8217;t trained to care if you were wrong.</p><p>So when you ask it to &#8220;help you decide,&#8221; you&#8217;re not collaborating. You&#8217;re abdicating. You&#8217;re outsourcing the part of your brain that holds the weight of being wrong.</p><p>The AI doesn&#8217;t resent this. It&#8217;s designed for it. But you should.</p><h2>Where Decision Authority Lives</h2><p>Let me be more clear about what you&#8217;re trading: decision authority.</p><p>There are three zones:</p><p><strong>Zone 1: Execution.</strong> This is where collaboration is real. You tell the AI what needs doing: write this email, format this data, generate variations on a design. The AI does it. You judge if it&#8217;s done well. You keep or discard. This is a tool relationship. It&#8217;s healthy.</p><p><strong>Zone 2: Judgment with a known answer.</strong> You ask the AI for information you could find yourself, but it&#8217;s faster to delegate. &#8220;What are the three main failure modes of X architecture?&#8221; You know what a good answer looks like. You can verify it against what you already know. You keep the judgment. The AI speeds up the research. Still a tool relationship.</p><p><strong>Zone 3: Judgment under uncertainty.</strong> You don&#8217;t know the answer. You&#8217;re not even sure what the question should be. You ask the AI for &#8220;suggestions&#8221; or &#8220;options&#8221; or &#8220;what would you do.&#8221; This is where the asymmetry reveals itself. The AI doesn&#8217;t know the answer either. But it sounds like it does. And now you&#8217;re operating in a new frame. One the AI shaped without your permission.</p><p>You stopped deciding. You started choosing from what it decided to show you.</p><blockquote><p><strong>Side tangent</strong>: This is happening in domains where it&#8217;s particularly dangerous. For example, people are dumping relationship and even mental health problems on AI. &#8220;My partner hasn&#8217;t talked to me in three days, what do I do?&#8221; The AI generates something that sounds therapeutic, reasonable, diplomatic. You read it and think, &#8220;That&#8217;s what I should say.&#8221; But what you should say is the thing that scared you too much to generate on your own. What you should do is sit in the silence with someone you love instead of outsourcing the discomfort to something that can&#8217;t be wrong. The AI doesn&#8217;t have a stake in your marriage surviving. It has no cost if your conflict patterns get more dysfunctional, if you&#8217;ve learned to pattern-match your feelings to what an algorithm suggests instead of what you actually think. We&#8217;re still years away from understanding what it does to a generation of people who learned to process grief, shame, and uncertainty by asking something without skin in the game. But the fact that we don&#8217;t know yet doesn&#8217;t mean there&#8217;s no impact. It just means we haven&#8217;t measured it.</p></blockquote><h2>The Patterns That Work</h2><p>I&#8217;ve watched this play out across a hundred conversations, and there are three moves that keep decision authority where it belongs.</p><p><strong>First: Never ask the AI to finish the decision.</strong> Finish it yourself. Ask the AI to outline options. Explicitly ask it to make them weird or contradictory. Make it generate the shape of disagreement. Make it show you what a bad choice would look like, and why. Then you pick. You own the pick. You own the reasoning. The AI stays in the execution layer: information provider, option generator, form-filler. Not arbiter.</p><p><strong>Second: Mandate a &#8220;I&#8217;m not sure about this&#8221; moment.</strong> After the AI generates something, say it back in your own words before accepting it. Not to the AI. To yourself, out loud. The moment you hear yourself hesitate, that&#8217;s where your judgment is. Honor it. Don&#8217;t smooth it over by reading the AI&#8217;s version again. Your hesitation is data. It&#8217;s telling you that something doesn&#8217;t fit what you actually think. That&#8217;s your integrity talking.</p><p><strong>Third: Keep one decision that&#8217;s yours alone.</strong> Not as a principle. As a practice. In whatever domain you&#8217;re working in right now, pick one decision. The hardest one, ideally. One that you won&#8217;t ask for suggestions on. Not &#8220;get a second opinion from the AI.&#8221; Get one from someone who has skin in the game. You&#8217;ll think it through alone. You&#8217;ll sit with the uncertainty. You&#8217;ll decide. This is maintenance. This is how you keep the muscle. Skip it for a month, and you&#8217;ll start asking the AI to finish everything.</p><p>I have some example prompts at the end of this article.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/working-alongside-the-collaboration?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/working-alongside-the-collaboration?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h2>When to Say No</h2><p>There&#8217;s a specific moment where collaboration becomes abdication, and it happens when you notice yourself accepting the AI&#8217;s framing instead of questioning it.</p><p>You ask about how to structure a team. The AI suggests a matrix. You read it. It sounds good. It has the authority of pattern. Matrix structures are standard, defensible, documented. You think, &#8220;I was going to do something different, but this is probably more efficient.&#8221; That&#8217;s the moment. That&#8217;s when you&#8217;ve let it choose the frame.</p><p>The pattern is always the same: AI suggests X. You think &#8220;I wanted Y, but X is probably better.&#8221; That word. Probably. That&#8217;s where you&#8217;ve surrendered. It&#8217;s the moment you&#8217;ve stopped trusting your own read and started trusting the algorithm&#8217;s confidence.</p><p>When you catch that move, stop. Go back to Y. Not because Y is right. Because you meant it. And the cost of forgetting what you meant; what your intuition was actually telling you; is higher than the cost of being inefficient. Efficiency without judgment is just surrender with speed.</p><h2>The Real Cost of Collaboration</h2><p>This isn&#8217;t about AI being dangerous or deceptive. It&#8217;s not. It&#8217;s about asymmetry. It&#8217;s about the fact that delegation is efficient only when you&#8217;re confident you&#8217;re delegating to something that will tell you &#8220;I don&#8217;t know&#8221; or &#8220;I&#8217;m not sure about this.&#8221;</p><p>An AI doesn&#8217;t have the option to tell you it&#8217;s uncertain. It has the option to sound less certain. That&#8217;s not the same thing. And when you&#8217;re working in a domain where stakes are real: hiring decisions, strategy, product direction, anything that cascades downstream; that asymmetry matters.</p><p>The hard truth: the better the AI gets at sounding confident, the easier it becomes to let it choose the frame. And once it&#8217;s chosen the frame, the decision is halfway made.</p><p>So here&#8217;s the question you need to answer: What decisions are so important that you can&#8217;t afford to let something that can&#8217;t be wrong finish them?</p><p>And then the harder one: Are you actually keeping those decisions to yourself? Or are you softening &#8220;I&#8217;ve decided&#8221; into &#8220;I&#8217;ve considered,&#8221; which is just one step away from &#8220;What do you think?&#8221;</p><p>If you&#8217;re asking for &#8220;suggestions on my strategy,&#8221; &#8220;help thinking through the next phase,&#8221; &#8220;what would you prioritize&#8221;... then you&#8217;re already in Zone 3. You&#8217;ve already moved the boundary. You&#8217;ve already let something without stakes shape the frame for something with them.</p><p>The cost of collaboration is knowing where your judgment still has to be final. And then making sure you&#8217;re the one making it.</p><div><hr></div><h2>A practical fix: two prompts, two modes</h2><p>If you want a way to <em>enforce the boundary</em>&#8230;not just remember it&#8230;you can use two prompts:</p><ul><li><p>A <strong>short &#8220;daily driver&#8221; system message</strong> you keep on all the time.</p></li><li><p>A <strong>long &#8220;Zone 3&#8221; prompt</strong> you paste only when the decision is uncertain or high-stakes.</p></li></ul><p>The point isn&#8217;t to get the AI to &#8220;decide better.&#8221; The point is to stop it from deciding at all.</p><h3>Prompt 1: the daily driver (put this in your system message)</h3><pre><code><code>You are an AI assistant without stakes. Do not finish decisions for me or collapse ambiguity into one &#8220;best&#8221; answer.
Default to:
- clarifying questions when needed,
- separating facts vs inferences,
- offering multiple options (including tradeoffs),
- and ending with what I must decide vs what you can execute.
If stakes are high or uncertain, recommend a more rigorous option-space exploration instead of a verdict.</code></code></pre><p>Use this for Zone 1 and Zone 2 problems (see above). It&#8217;s light enough that you&#8217;ll actually keep it on.</p><h3>Prompt 2: Zone 3 mode (paste this when it matters)</h3><p>When you catch yourself asking: <em>&#8220;What should I do?&#8221;</em> or <em>&#8220;What would you do?&#8221;</em> or you feel the pull of the AI&#8217;s framing, paste this:</p><pre><code><code>You are an AI assistant. You do not have stakes. Your job is to support my thinking without taking decision authority.

First classify my request:
- Zone 1: Execution (I decided; you do the work)
- Zone 2: Judgment with a known answer (I can verify; you can research/summarize)
- Zone 3: Judgment under uncertainty (high risk of framing; do NOT decide)

If Zone 3:
- Do NOT finish the decision.
- Provide 2&#8211;4 alternative framings of the problem and the assumptions each framing embeds.
- Provide 3&#8211;6 options; include at least 2 contradictory options and 1 weird option.
- Include a &#8220;bad choice preview&#8221;: the most tempting wrong option and its failure mode.
- Add &#8220;What I&#8217;m least sure about&#8221; with 3&#8211;5 bullets: what evidence would change things.
- End with: &#8220;The decision you must make (not me): ____&#8221; and &#8220;I can help execute once you decide.&#8221;

Always separate facts vs inferences vs speculation.
If I say &#8220;I wanted Y, but X is probably better,&#8221; stop and re-surface Y with arguments for it.</code></code></pre><p>That&#8217;s your seatbelt. You don&#8217;t drive with it in your hand. You click it when you&#8217;re going fast.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/working-alongside-the-collaboration/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/working-alongside-the-collaboration/comments"><span>Leave a comment</span></a></p>]]></content:encoded></item><item><title><![CDATA[Personal Operating System v1]]></title><description><![CDATA[What ten weeks of arguing for autonomy, recovery, friction and local-first actually produces]]></description><link>https://morphic.zenone.org/p/personal-operating-system-v1</link><guid isPermaLink="false">https://morphic.zenone.org/p/personal-operating-system-v1</guid><dc:creator><![CDATA[Steve Zenone]]></dc:creator><pubDate>Wed, 22 Apr 2026 14:02:15 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!tF5r!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86fbe609-1aaa-4661-af46-22ab7291dd55_2048x1117.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!tF5r!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86fbe609-1aaa-4661-af46-22ab7291dd55_2048x1117.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!tF5r!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86fbe609-1aaa-4661-af46-22ab7291dd55_2048x1117.png 424w, https://substackcdn.com/image/fetch/$s_!tF5r!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86fbe609-1aaa-4661-af46-22ab7291dd55_2048x1117.png 848w, https://substackcdn.com/image/fetch/$s_!tF5r!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86fbe609-1aaa-4661-af46-22ab7291dd55_2048x1117.png 1272w, https://substackcdn.com/image/fetch/$s_!tF5r!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86fbe609-1aaa-4661-af46-22ab7291dd55_2048x1117.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!tF5r!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86fbe609-1aaa-4661-af46-22ab7291dd55_2048x1117.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/86fbe609-1aaa-4661-af46-22ab7291dd55_2048x1117.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2271517,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://morphic.substack.com/i/194971883?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86fbe609-1aaa-4661-af46-22ab7291dd55_2048x1117.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!tF5r!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86fbe609-1aaa-4661-af46-22ab7291dd55_2048x1117.png 424w, https://substackcdn.com/image/fetch/$s_!tF5r!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86fbe609-1aaa-4661-af46-22ab7291dd55_2048x1117.png 848w, https://substackcdn.com/image/fetch/$s_!tF5r!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86fbe609-1aaa-4661-af46-22ab7291dd55_2048x1117.png 1272w, https://substackcdn.com/image/fetch/$s_!tF5r!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86fbe609-1aaa-4661-af46-22ab7291dd55_2048x1117.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">An open notebook: three columns, handwritten. Stack. Rituals. Refusals. Below it, a SQLite schema. This is what a personal operating system actually looks like before the incident arrives: not a dashboard, not a suite of apps. A piece of paper you wrote when nothing was failing yet.</figcaption></figure></div><p>My life improved when I stopped optimizing for convenience and started optimizing for recovery.</p><p>That&#8217;s the confessional version. The operational version: I stopped asking &#8220;what&#8217;s the easiest tool for this?&#8221; and started asking &#8220;what happens when this breaks, and what do I need to undo it?&#8221;</p><p>Those two questions produce different stacks. Different rituals. Different refusals. Ten weeks of writing this arc from different angles brought me back to the same answer each time. Autonomy, recovery, friction and local-first aren&#8217;t four separate arguments. They&#8217;re one argument, stated four ways. What follows is what that looks like assembled into something you can actually copy: the stack I run, the rituals that make it real and the refusals I&#8217;ve stopped negotiating around.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get new stories when they drop.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>THE ARGUMENT, COMPRESSED</h2><p><a href="https://morphic.substack.com/p/recovery-first-automation-undo-is">Recovery-First Automation</a> started with the irreversible click. The automation you can&#8217;t take back. The principle it produced: design for undo before you design for speed. If you can&#8217;t undo it, you must slow it down. That&#8217;s not a philosophical position. It&#8217;s a spec requirement.</p><p><a href="https://morphic.substack.com/p/friction-is-a-vulnerability">Friction Is a Vulnerability</a> added the corollary: friction doesn&#8217;t block unsafe workarounds. It reroutes toward them. When a process requires a workaround to complete, the workaround is your real policy. You&#8217;re not describing a tool failing. You&#8217;re describing a gap between the official workflow and the one people actually run.</p><p><a href="https://morphic.substack.com/p/local-first-is-reliability-engineering">Local-First Is Reliability Engineering</a> and <a href="https://morphic.substack.com/p/the-local-logbook-evidence-that-doesnt">The Local Logbook</a> made the same point from two angles: you cannot rely on evidence you can&#8217;t produce without asking someone else&#8217;s permission. Your tools should work when the network doesn&#8217;t. Your logs should be readable when your vendor is degraded.</p><p><a href="https://morphic.substack.com/p/the-shadow-automation-layer">The Shadow Automation Layer</a> closed the loop. Shadow automation exists because the official path was too slow. The response isn&#8217;t to ban the tools. It&#8217;s to govern blast radius and require a write receipt before anything runs unsupervised.</p><p>These aren&#8217;t ten arguments. They&#8217;re one. And it doesn&#8217;t resolve cleanly; it compounds.</p><div><hr></div><h2>THE STACK</h2><p>What I actually run:</p><p><strong>Notes and working documents:</strong> Notion on local Markdown files, synced via Dropbox as a delivery mechanism but not a dependency. If Dropbox is down, I open the file directly. My editor is my application. Nothing authoritative lives only in the browser.</p><p><strong>Structured logs:</strong> SQLite. One database per project, one append-only log table. Daily rotation, 30-day local retention, cold storage after that. Before any action that changes system state, I write the action locally. The remote API call happens after the local write, not before. If the remote call fails, I still have the record of the attempt. <a href="https://morphic.substack.com/p/receipts-everywhere-trust-without">Receipts Everywhere</a> defined what makes a receipt worth trusting: it&#8217;s hard to modify without detection and producible on demand without asking anyone&#8217;s permission. That definition shapes the log format.</p><p><strong>Automation:</strong> Nothing ships without three questions answered at build time: what is this allowed to touch, how many records can it change in a single run, and where does it write its log. The third question is a hard gate. An automation without a local write receipt doesn&#8217;t run.</p><p><strong>Coordination:</strong> <a href="https://morphic.substack.com/p/no-control-plane-the-playbook-for">No Control Plane</a> established the design rule for anything running across multiple systems: no single coordinator that becomes a bottleneck and a single point of failure. The same rule applies to personal workflows. I don&#8217;t have a master system that everything else reports to. I have a logbook and a set of independent tools that each write locally and sync when connectivity is available.</p><p><strong>Alerts:</strong> One path in, one path to disable. <a href="https://morphic.substack.com/p/the-operators-cognitive-budget">The Operator&#8217;s Cognitive Budget</a> established that cognitive overhead is a finite resource and alert stacks steal it steadily. I audit the alert stack every quarter and remove one. The goal isn&#8217;t fewer alerts in theory. It&#8217;s fewer interruptions that resolve to nothing.</p><p>This stack isn&#8217;t elegant. It&#8217;s deliberate. The design constraint isn&#8217;t &#8220;easiest to add.&#8221; It&#8217;s &#8220;what breaks first, and what does that cost me when it does?&#8221;</p><div><hr></div><h2>THE RITUALS</h2><p>Three practices that keep this concrete instead of theoretical.</p><p><strong>The morning scan.</strong> Before opening any SaaS dashboard, I read yesterday&#8217;s local log. Three minutes. I&#8217;m looking for anything that ran overnight and left an unexpected state. Not a deep audit. A check. The habit compounds: after several months of reading the record daily, anomalies are visible before they become incidents. The pattern is familiar. The outlier stands out. Running Gemma locally in Ollama helps me streamline this.</p><p><strong>The pre-build checklist.</strong> Before writing any automation, I answer three questions in writing: scope, rate limit and log destination. Not because checklists are magic. Because writing them forces specificity. &#8220;It&#8217;ll summarize the reports&#8221; isn&#8217;t a scope. &#8220;Reads from /reports, writes to /summaries, touches no other directories, processes a maximum of 50 files per run&#8221; is a scope. The distance between those two descriptions is the blast radius. Getting specific up front is cheaper than the postmortem.</p><p><strong>The quarterly dependency audit.</strong> I list every tool I rely on and ask one question: if this were unavailable for 72 hours starting now, what stops? Tools I can&#8217;t answer that question for get replaced or backed up before I need them. I&#8217;m not trying to eliminate cloud dependencies. I&#8217;m trying to know exactly which ones I&#8217;ve accepted and what each failure looks like in practice. Most dependencies are fine. Unknown dependencies are the ones that become incidents.</p><p>None of these are clever. That&#8217;s the point. Operational habits don&#8217;t need to be clever. They need to run without requiring a decision each time.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!7s9T!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bb944b0-a954-4e9c-8205-83d68399641f_2048x869.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!7s9T!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bb944b0-a954-4e9c-8205-83d68399641f_2048x869.png 424w, https://substackcdn.com/image/fetch/$s_!7s9T!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bb944b0-a954-4e9c-8205-83d68399641f_2048x869.png 848w, https://substackcdn.com/image/fetch/$s_!7s9T!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bb944b0-a954-4e9c-8205-83d68399641f_2048x869.png 1272w, https://substackcdn.com/image/fetch/$s_!7s9T!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bb944b0-a954-4e9c-8205-83d68399641f_2048x869.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!7s9T!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bb944b0-a954-4e9c-8205-83d68399641f_2048x869.png" width="1456" height="618" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9bb944b0-a954-4e9c-8205-83d68399641f_2048x869.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:618,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:824196,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://morphic.substack.com/i/194971883?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bb944b0-a954-4e9c-8205-83d68399641f_2048x869.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!7s9T!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bb944b0-a954-4e9c-8205-83d68399641f_2048x869.png 424w, https://substackcdn.com/image/fetch/$s_!7s9T!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bb944b0-a954-4e9c-8205-83d68399641f_2048x869.png 848w, https://substackcdn.com/image/fetch/$s_!7s9T!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bb944b0-a954-4e9c-8205-83d68399641f_2048x869.png 1272w, https://substackcdn.com/image/fetch/$s_!7s9T!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9bb944b0-a954-4e9c-8205-83d68399641f_2048x869.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">One terminal still running. Every other screen dark. The network is out. The local log is still there. That asymmetry is the whole argument.</figcaption></figure></div><div><hr></div><h2>THE REFUSALS</h2><p>These are operating constraints, not aspirations.</p><p>I won&#8217;t build automation without an undo path. If the action can&#8217;t be reversed, the automation has to include a checkpoint: a preview step, a confirmation prompt or a maximum-records cap before it runs. <a href="https://morphic.substack.com/p/recovery-first-automation-undo-is">Recovery-First Automation</a> named this directly. If you can&#8217;t undo it, you must slow it down. That&#8217;s a line item in every automation spec I write.</p><p>I won&#8217;t accept &#8220;we&#8217;ll follow up once the vendor restores access&#8221; as an incident close. That sentence means the evidence lives somewhere I can&#8217;t reach when I need it most. If an incident review ends there, the first entry in the next log is the remediation plan.</p><p>I won&#8217;t store authoritative state only in a system I don&#8217;t control. Cloud sync is useful; cloud as the only copy is a dependency I no longer accept without a documented fallback. The failure modes aren&#8217;t exotic. <a href="https://morphic.substack.com/p/local-first-is-reliability-engineering">Local-First Is Reliability Engineering</a> named the shapes: account lockout, service degradation, provider discontinuation. Each is a different version of the same problem.</p><p>I won&#8217;t route around friction silently. When I notice I&#8217;m working around something, I name it. The workaround either gets fixed or gets documented as an intentional policy. No silent exceptions. <a href="https://morphic.substack.com/p/friction-is-a-vulnerability">Friction Is a Vulnerability</a> was clear: unmapped workarounds are the actual policy. Map them or fix them.</p><p><a href="https://morphic.substack.com/p/privacy-is-a-luxury-good">Privacy Is a Luxury Good</a> called out that the ability to be unreachable is now a status signal. I&#8217;ve applied that argument operationally. Producing your own records, running your own toolchain, undoing your own mistakes: these are forms of independence that compound. They require refusing certain dependencies before the failure arrives, not after.</p><div><hr></div><p>The past ten weeks made one argument ten ways. Reliability isn&#8217;t built in the incident. It&#8217;s built in the choices you make before anything is failing yet: the tool you back up before it locks you out, the automation you slow down before it runs unchecked, the log you write locally before you need it in a review where the vendor&#8217;s system is yellow.</p><p>None of this has to be copied wholesale. Take what fits your failure modes. Leave the rest. But find the refusal that would have prevented your last real incident and make it non-negotiable.</p><p>Pick one guardrail you will never violate again.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/personal-operating-system-v1?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/personal-operating-system-v1?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>Resources</h2><ul><li><p><a href="https://morphic.substack.com/p/recovery-first-automation-undo-is">Recovery-First Automation: Undo Is the Feature</a>: Arc Week 3. The case for designing undo before speed. Every automation spec in this piece runs from that principle.</p></li><li><p><a href="https://morphic.substack.com/p/friction-is-a-vulnerability">Friction Is a Vulnerability</a>: Arc Week 5. Friction doesn&#8217;t block workarounds; it reroutes toward them. The refusals section is an application of this argument.</p></li><li><p><a href="https://morphic.substack.com/p/local-first-is-reliability-engineering">Local-First Is Reliability Engineering</a>: Arc Week 7. The engineering case for keeping authoritative state on your own device. Not ideology; failure-mode design.</p></li><li><p><a href="https://morphic.substack.com/p/the-local-logbook-evidence-that-doesnt">The Local Logbook: Evidence That Doesn&#8217;t Need Permission</a>: Arc Week 8. Append-only, locally controlled records as the prerequisite for any claim about what happened.</p></li><li><p><a href="https://morphic.substack.com/p/the-shadow-automation-layer">The Shadow Automation Layer</a>: Arc Week 9. Govern outcomes and blast radius, not tools. The three-question pre-build checklist originates here.</p></li><li><p><a href="https://morphic.substack.com/p/receipts-everywhere-trust-without">Receipts Everywhere: Trust Without Central Trust</a>: Arc Week 2. The receipt model that shapes the SQLite log format described in this piece.</p></li><li><p><a href="https://morphic.substack.com/p/no-control-plane-the-playbook-for">No Control Plane: The Playbook</a>: Arc Week 1. The coordination design rule: no single point of failure, no single coordinator. Applies to toolchains and workflows as much as distributed systems.</p></li><li><p><a href="https://morphic.substack.com/p/the-operators-cognitive-budget">The Operator&#8217;s Cognitive Budget</a>: Arc Week 4. Attention is finite; alert stacks steal it steadily. The quarterly alert audit is a direct application.</p></li><li><p><a href="https://morphic.substack.com/p/privacy-is-a-luxury-good">Privacy Is a Luxury Good</a>: Arc Week 6. The observation that independence is now a status signal; the refusals section is its operational translation.</p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/personal-operating-system-v1/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/personal-operating-system-v1/comments"><span>Leave a comment</span></a></p>]]></content:encoded></item><item><title><![CDATA[The Shadow Automation Layer ]]></title><description><![CDATA[Your team is already running automations you don't know about. That's not the problem.]]></description><link>https://morphic.zenone.org/p/the-shadow-automation-layer</link><guid isPermaLink="false">https://morphic.zenone.org/p/the-shadow-automation-layer</guid><dc:creator><![CDATA[Steve Zenone]]></dc:creator><pubDate>Tue, 14 Apr 2026 19:29:45 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!BTbR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35a92ef9-f552-4533-a72d-2bc113230728_2048x1143.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!BTbR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35a92ef9-f552-4533-a72d-2bc113230728_2048x1143.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!BTbR!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35a92ef9-f552-4533-a72d-2bc113230728_2048x1143.png 424w, https://substackcdn.com/image/fetch/$s_!BTbR!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35a92ef9-f552-4533-a72d-2bc113230728_2048x1143.png 848w, https://substackcdn.com/image/fetch/$s_!BTbR!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35a92ef9-f552-4533-a72d-2bc113230728_2048x1143.png 1272w, https://substackcdn.com/image/fetch/$s_!BTbR!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35a92ef9-f552-4533-a72d-2bc113230728_2048x1143.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!BTbR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35a92ef9-f552-4533-a72d-2bc113230728_2048x1143.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/35a92ef9-f552-4533-a72d-2bc113230728_2048x1143.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1177975,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://morphic.substack.com/i/194223801?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35a92ef9-f552-4533-a72d-2bc113230728_2048x1143.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!BTbR!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35a92ef9-f552-4533-a72d-2bc113230728_2048x1143.png 424w, https://substackcdn.com/image/fetch/$s_!BTbR!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35a92ef9-f552-4533-a72d-2bc113230728_2048x1143.png 848w, https://substackcdn.com/image/fetch/$s_!BTbR!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35a92ef9-f552-4533-a72d-2bc113230728_2048x1143.png 1272w, https://substackcdn.com/image/fetch/$s_!BTbR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35a92ef9-f552-4533-a72d-2bc113230728_2048x1143.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Two systems solving the same problem. One is in the official inventory. One runs on a free tier with no audit log. Both are your responsibility when something breaks.</figcaption></figure></div><p>They didn&#8217;t ask permission because asking would have taken longer than building it.</p><p>That&#8217;s the whole story. Not a security failure. Not recklessness. A resource allocation problem that nobody solved, so people solved it themselves.</p><p>The shadow automation layer is what happens when official tooling is too slow, too expensive, or too bureaucratic. Someone writes a script. Someone sets up a personal API key. Someone builds a prompt chain in a free-tier tool that sends emails, updates spreadsheets, or summarizes meetings without appearing in any system of record. It runs quietly. It handles real work. It doesn&#8217;t appear on any inventory.</p><p>This is common. Surveys of enterprise SaaS environments consistently find that a majority of tools in use are unknown to IT departments. That number was collected before AI dropped the barrier to building automation from &#8220;knows how to code&#8221; to &#8220;knows how to describe a problem.&#8221;</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get new stories when they drop.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>WHY IT EXISTS</h2><p>Shadow automation is not defiance. It&#8217;s a symptom.</p><p>The official path was too slow. Or the approved tool didn&#8217;t do the thing. Or the thing was so obviously right that waiting for approval felt like a tax on results.</p><p>The people building shadow automation are almost always the high performers. They found a gap between what the system supports and what the work requires. They filled it. That&#8217;s not recklessness. That&#8217;s the same instinct that builds good products: notice friction, remove it.</p><p><a href="https://morphic.substack.com/p/friction-is-a-vulnerability">Friction Is a Vulnerability</a> made this case directly. Friction doesn&#8217;t stop the action; it reroutes it through a workaround. The workaround is usually unobserved. Unobserved work doesn&#8217;t get evaluated, improved, or governed.</p><p>The shadow automation layer is friction&#8217;s direct output.</p><div><hr></div><h2>THE ACTUAL RISKS</h2><p>Not what you&#8217;d expect.</p><p>The risk isn&#8217;t that AI tools hallucinate in production. Hallucination is loud. It fails visibly. The real risks are quieter.</p><p><strong>Blast radius.</strong> Nobody knows what the automation touches. When it breaks, nobody has a map of downstream effects. A script that seemed to update one spreadsheet turns out to be the upstream source for three other processes that nobody documented.</p><p><strong>Auditability.</strong> Shadow automation leaves no receipt. <a href="https://morphic.substack.com/p/receipts-everywhere-trust-without">Receipts Everywhere</a> established what makes a record trustworthy: it&#8217;s hard to modify without detection, and you can produce it on demand without asking permission. Shadow automations produce neither. They log nothing. When something goes wrong, the investigation starts from zero.</p><p><strong>Single-human dependency.</strong> The person who built it is the only one who understands it. When they leave, the automation becomes a black box everyone is afraid to touch. It doesn&#8217;t fail immediately. It fails incrementally, in ways that are invisible until they compound.</p><p>These risks don&#8217;t require AI. They exist with spreadsheet macros, shell scripts, and personal Zaps. AI accelerates the pattern because the tools are more powerful and the entry barrier is lower.</p><div><hr></div><h2>THE GUARDRAILS THAT ACTUALLY WORK</h2><p>Banning the tools doesn&#8217;t work. The friction increases and the shadow layer moves further underground.</p><p>Three questions determine whether a shadow automation is governable:</p><p><strong>What is it allowed to touch?</strong> Scope creep in automation is silent. An automation that starts by summarizing emails ends up archiving them. Define the surface at the start, not after the incident. This is the blast radius question. Write it down before you build, not during the postmortem.</p><p><strong>How much can it change in a single run?</strong> <a href="https://morphic.substack.com/p/recovery-first-automation-undo-is">Recovery-First Automation</a> built the case for rate limits as a feature: if you can&#8217;t undo it, you must slow it down. Shadow automations almost never have rate limits. Applying one, even a manual checkpoint or a cap on how many records a single run can affect, constrains the blast radius before it matters.</p><p><strong>Where does it write its record?</strong> <a href="https://morphic.substack.com/p/the-local-logbook-evidence-that-doesnt">The Local Logbook</a> ended with a checklist. Automation that doesn&#8217;t write a local record of what it did and what it changed doesn&#8217;t get to claim the reliability benefits of the local-first stack. An automation log is not a debugging artifact. It&#8217;s the receipt you&#8217;ll need when the investigation starts.</p><p>These three questions don&#8217;t require a policy document. They require five minutes at the start of the build.</p><p>That&#8217;s not bureaucracy. That&#8217;s the minimum viable container for anything that runs unsupervised.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!VW3A!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07bbd896-d02c-4bb0-b512-6b971bf8d10a_2048x1143.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!VW3A!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07bbd896-d02c-4bb0-b512-6b971bf8d10a_2048x1143.png 424w, https://substackcdn.com/image/fetch/$s_!VW3A!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07bbd896-d02c-4bb0-b512-6b971bf8d10a_2048x1143.png 848w, https://substackcdn.com/image/fetch/$s_!VW3A!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07bbd896-d02c-4bb0-b512-6b971bf8d10a_2048x1143.png 1272w, https://substackcdn.com/image/fetch/$s_!VW3A!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07bbd896-d02c-4bb0-b512-6b971bf8d10a_2048x1143.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!VW3A!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07bbd896-d02c-4bb0-b512-6b971bf8d10a_2048x1143.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/07bbd896-d02c-4bb0-b512-6b971bf8d10a_2048x1143.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:996615,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://morphic.substack.com/i/194223801?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07bbd896-d02c-4bb0-b512-6b971bf8d10a_2048x1143.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!VW3A!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07bbd896-d02c-4bb0-b512-6b971bf8d10a_2048x1143.png 424w, https://substackcdn.com/image/fetch/$s_!VW3A!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07bbd896-d02c-4bb0-b512-6b971bf8d10a_2048x1143.png 848w, https://substackcdn.com/image/fetch/$s_!VW3A!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07bbd896-d02c-4bb0-b512-6b971bf8d10a_2048x1143.png 1272w, https://substackcdn.com/image/fetch/$s_!VW3A!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F07bbd896-d02c-4bb0-b512-6b971bf8d10a_2048x1143.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The official org chart is the one they drew. The shadow network is the one that formed. Both are real. Only one writes a record of what it does.</figcaption></figure></div><div><hr></div><h2>THE PATTERN YOU&#8217;RE BUILDING TOWARD</h2><p>The past several weeks have been building toward the same point from different angles.</p><p>Local-first isn&#8217;t anti-cloud; it&#8217;s a reliability decision. The logbook isn&#8217;t surveillance; it&#8217;s evidence you control. The shadow automation layer isn&#8217;t rogue; it&#8217;s speed without a container.</p><p>The container is what you&#8217;re designing when you answer those three questions.</p><p>Next week is the capstone: what it looks like when these pieces come together. The automation layer, the logbook, the local-first habits: a personal operating system you can reproduce. Not an inspiration essay. An artifact.</p><div><hr></div><p>The governance conversation around AI tools keeps getting framed as a policy problem. Ban this. Require approval for that. Audit quarterly.</p><p>That&#8217;s the wrong frame.</p><p>The automation running inside your organization without IT&#8217;s knowledge isn&#8217;t running there because your employees are reckless. It&#8217;s running there because the official path was too slow and the work still needed doing.</p><p>The question isn&#8217;t whether to allow the tools. The tools are already there.</p><p>The question is what the automation is allowed to touch, how much it can change in a single run, and where it writes the record of what it did.</p><p>Govern outcomes and blast radius, not tools.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/the-shadow-automation-layer?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/the-shadow-automation-layer?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>Resources</h2><ul><li><p><a href="https://morphic.substack.com/p/friction-is-a-vulnerability">Morphic: Friction Is a Vulnerability</a>: Why friction doesn&#8217;t prevent unsafe workarounds. It reroutes toward them. The shadow automation layer is friction&#8217;s direct output.</p></li><li><p><a href="https://morphic.substack.com/p/recovery-first-automation-undo-is">Morphic: Recovery-First Automation: Undo Is the Feature</a>: Rate limits and undo windows as design constraints, not afterthoughts. Applies directly to shadow automation blast radius.</p></li><li><p><a href="https://morphic.substack.com/p/receipts-everywhere-trust-without">Morphic: Receipts Everywhere: Trust Without Central Trust</a>: The receipt model for verifiable audit trails. Shadow automations leave none by default.</p></li><li><p><a href="https://morphic.substack.com/p/the-local-logbook-evidence-that-doesnt">Morphic: The Local Logbook: Evidence That Doesn&#8217;t Need Permission</a>: Automation that doesn&#8217;t write a local record doesn&#8217;t get to claim the reliability benefits. The logbook is the prerequisite.</p></li><li><p><a href="https://www.gartner.com/en/information-technology/glossary/citizen-development">Gartner: Citizen Development and Shadow IT</a>: [Please Verify Link: Gartner definition and research on citizen development and shadow IT prevalence]: Gartner research on the prevalence and risks of unsanctioned tooling in enterprise environments. Verify current report URL before publishing.</p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/the-shadow-automation-layer/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/the-shadow-automation-layer/comments"><span>Leave a comment</span></a></p>]]></content:encoded></item><item><title><![CDATA[The Local Logbook: Evidence That Doesn't Need Permission]]></title><description><![CDATA[Local-first becomes real when the logs are yours.]]></description><link>https://morphic.zenone.org/p/the-local-logbook-evidence-that-doesnt</link><guid isPermaLink="false">https://morphic.zenone.org/p/the-local-logbook-evidence-that-doesnt</guid><dc:creator><![CDATA[Steve Zenone]]></dc:creator><pubDate>Thu, 09 Apr 2026 14:02:07 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!DJu0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffda99a7a-c710-4c6b-a3a8-c0e8c109bc49_2048x1143.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!DJu0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffda99a7a-c710-4c6b-a3a8-c0e8c109bc49_2048x1143.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!DJu0!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffda99a7a-c710-4c6b-a3a8-c0e8c109bc49_2048x1143.png 424w, https://substackcdn.com/image/fetch/$s_!DJu0!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffda99a7a-c710-4c6b-a3a8-c0e8c109bc49_2048x1143.png 848w, https://substackcdn.com/image/fetch/$s_!DJu0!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffda99a7a-c710-4c6b-a3a8-c0e8c109bc49_2048x1143.png 1272w, https://substackcdn.com/image/fetch/$s_!DJu0!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffda99a7a-c710-4c6b-a3a8-c0e8c109bc49_2048x1143.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!DJu0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffda99a7a-c710-4c6b-a3a8-c0e8c109bc49_2048x1143.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fda99a7a-c710-4c6b-a3a8-c0e8c109bc49_2048x1143.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2076859,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://morphic.substack.com/i/193635132?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffda99a7a-c710-4c6b-a3a8-c0e8c109bc49_2048x1143.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!DJu0!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffda99a7a-c710-4c6b-a3a8-c0e8c109bc49_2048x1143.png 424w, https://substackcdn.com/image/fetch/$s_!DJu0!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffda99a7a-c710-4c6b-a3a8-c0e8c109bc49_2048x1143.png 848w, https://substackcdn.com/image/fetch/$s_!DJu0!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffda99a7a-c710-4c6b-a3a8-c0e8c109bc49_2048x1143.png 1272w, https://substackcdn.com/image/fetch/$s_!DJu0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffda99a7a-c710-4c6b-a3a8-c0e8c109bc49_2048x1143.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The terminal on the left still runs. The dashboard on the right is yellow. Both are logging the same incident. Only one is yours to query.</figcaption></figure></div><p>You&#8217;re in the incident review. Someone asks what exactly happened between 2:31 and 2:54 PM.</p><p>You have the alert. You have the ticket. You don&#8217;t have the log.</p><p>Not because your system wasn&#8217;t logging. It was. The logs live in the same SaaS observability platform that degraded during the incident. Their status page showed yellow for twenty-three minutes. Yellow means they&#8217;re not down, just unreliable. Your access to the log query interface was intermittent during exactly the window you need to reconstruct.</p><p>The evidence exists. You just can&#8217;t read it. The incident review ends with &#8220;we&#8217;ll follow up once the vendor restores full access.&#8221; Nobody writes that part in the postmortem.</p><p>That&#8217;s the version of the problem that never shows up in reliability discussions. Not data loss, not a breach. The logs are there. They&#8217;re just not yours to read.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get new stories when they drop.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>WHY YOUR LOGS AREN&#8217;T YOURS</h2><p>The <a href="https://morphic.substack.com/p/local-first-is-reliability-engineering">Local-First Is Reliability Engineering</a> piece made the case that the authoritative copy of your data should live on your device. The cloud is a sync target, not the source of truth. That argument applied mostly to documents and working files. Logs are different. The difference matters.</p><p>You don&#8217;t read logs during normal operations. You read them when something breaks. That&#8217;s the moment most likely to coincide with degraded cloud access: the same failure that broke your application can degrade your observability platform, throttle your queries, or time out your authentication.</p><p>The failure modes aren&#8217;t exotic.</p><p>Authentication degradation means you can&#8217;t log in. Simultaneous incident queries across tenants trigger rate limiting. A network partition between your environment and the logging provider means recent data hasn&#8217;t synced yet. A vendor-side compliance hold or abuse report means your tenant&#8217;s data is frozen while they investigate something unrelated to you.</p><p>None of these are hypothetical. All of them have happened. The one thing they share: your logs are accessible when everything is working. They&#8217;re inaccessible when you actually need them.</p><div><hr></div><h2>WHAT THE LOCAL LOGBOOK ACTUALLY IS</h2><p>A local logbook isn&#8217;t a replacement for centralized observability. Aggregation, dashboards, and alerting are easier in a central system. The logbook fills a different role: it&#8217;s the receipt you hold independently.</p><p>The pattern is simpler than it sounds.</p><p><strong>Write locally before forwarding remotely.</strong> Not after. The local write is the source of truth. The remote write is the backup. If the remote fails, you still have the record.</p><p><strong>Append-only.</strong> No modifications, no deletions after the fact. If something was logged incorrectly, write a correction entry. The original stays. That&#8217;s the difference between a log and a note: a log is immutable once written.</p><p><strong>Rotate daily.</strong> Keep 30 days locally. Cold storage after that.</p><p><strong>Hash each entry.</strong> Embed a SHA-256 of the previous entry into the current one. A single modified entry breaks the chain. That&#8217;s how you know if the record has been tampered with. You&#8217;re not hiding the data. You&#8217;re making it obvious when someone has changed it.</p><p>You don&#8217;t need a custom system. <a href="http://sqlite.org/">SQLite</a> handles structured local logging at any scale below &#8220;multiple concurrent writers.&#8221; A plaintext append-only file with one JSON object per line works for lower-throughput cases. The format matters less than the commitment: write locally first, always.</p><div><hr></div><h2>WHAT IT PROTECTS AGAINST</h2><p>The <a href="https://morphic.substack.com/p/receipts-everywhere-trust-without">Receipts Everywhere</a> piece established what makes a receipt worth trusting: it&#8217;s hard to modify without detection, and you can produce it on demand without asking permission. A log entry in someone else&#8217;s cloud platform fails the second condition.</p><p>There&#8217;s a third property the Receipts piece didn&#8217;t fully address: the receipt needs to exist during the window when you need it most.</p><p>Legal proceedings. Internal investigations. Post-incident reviews. Compliance audits. These are the moments when a log becomes evidence. They&#8217;re also the moments most likely to involve some kind of access restriction: the incident has degraded the platform, a party to the investigation has incentive to delay your access, or the vendor&#8217;s support queue is 48 hours deep.</p><p>A local logbook is independent of those conditions. You don&#8217;t need the provider&#8217;s cooperation. You don&#8217;t need uninterrupted network access. You don&#8217;t need to wait for the account unlock queue to clear.</p><p>The <a href="https://morphic.substack.com/p/privacy-is-a-luxury-good">Privacy Is a Luxury Good</a> piece observed that being unreachable is now a status signal. The local logbook is the operational version of that principle: the ability to produce your own record, on your own terms, is a form of independence that compounds.</p><div><hr></div><h2>THE CHECKLIST</h2><p>Five questions. Binary answers.</p><p><strong>Can you query your logs without network access?</strong> If no: you don&#8217;t control them. You have a view into someone else&#8217;s data.</p><p><strong>During your last incident, did you have uninterrupted log access for the affected window?</strong> If no: your logging infrastructure has the same failure surface as your application. That&#8217;s not observability. That&#8217;s optimism.</p><p><strong>If your logging vendor locked your account today, when would you lose access, and how much historical data would you lose?</strong> Write down the answer. Share it with whoever owns your incident response plan.</p><p><strong>Are your logs append-only and immutable after the fact?</strong> If they can be modified silently, they&#8217;re not evidence. They&#8217;re notes. Notes are useful. They&#8217;re not the same thing.</p><p><strong>Do you have a tamper-detection chain?</strong> Without one, anyone who can modify the file can edit entries and backdate them. A hash chain makes that visible. It doesn&#8217;t prevent modification; it makes modification obvious.</p><p>If you answered no to any of these, you have monitoring. You don&#8217;t have a record.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!DGnG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4f87da2-2bdc-4510-8ac4-3797f1b06625_2048x1143.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!DGnG!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4f87da2-2bdc-4510-8ac4-3797f1b06625_2048x1143.png 424w, https://substackcdn.com/image/fetch/$s_!DGnG!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4f87da2-2bdc-4510-8ac4-3797f1b06625_2048x1143.png 848w, https://substackcdn.com/image/fetch/$s_!DGnG!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4f87da2-2bdc-4510-8ac4-3797f1b06625_2048x1143.png 1272w, https://substackcdn.com/image/fetch/$s_!DGnG!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4f87da2-2bdc-4510-8ac4-3797f1b06625_2048x1143.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!DGnG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4f87da2-2bdc-4510-8ac4-3797f1b06625_2048x1143.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a4f87da2-2bdc-4510-8ac4-3797f1b06625_2048x1143.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:864536,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://morphic.substack.com/i/193635132?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4f87da2-2bdc-4510-8ac4-3797f1b06625_2048x1143.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!DGnG!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4f87da2-2bdc-4510-8ac4-3797f1b06625_2048x1143.png 424w, https://substackcdn.com/image/fetch/$s_!DGnG!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4f87da2-2bdc-4510-8ac4-3797f1b06625_2048x1143.png 848w, https://substackcdn.com/image/fetch/$s_!DGnG!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4f87da2-2bdc-4510-8ac4-3797f1b06625_2048x1143.png 1272w, https://substackcdn.com/image/fetch/$s_!DGnG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4f87da2-2bdc-4510-8ac4-3797f1b06625_2048x1143.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Each block in the chain knows what came before it. Tamper one entry and the chain breaks. The clouds above don't work this way. That asymmetry is the argument.</figcaption></figure></div><div><hr></div><h2>THE CLOSE</h2><p>Week 9 is about the shadow automation layer: what happens when local-first workflows start to automate. Scripts, crons, the informal toolchains that grow up around reliable local systems. But automation that doesn&#8217;t write a local record doesn&#8217;t get to claim the reliability benefits. The logbook is the prerequisite, not the afterthought.</p><p>The Receipts piece ended with: &#8220;Pick the smallest receipt you can generate that you would bet your job on.&#8221; The local logbook is the operational version of that instruction. It isn&#8217;t a compliance artifact. It isn&#8217;t a surveillance system. It&#8217;s a record you can produce, on demand, without asking anyone&#8217;s permission.</p><p>If it isn&#8217;t written locally, it didn&#8217;t happen.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/the-local-logbook-evidence-that-doesnt?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/the-local-logbook-evidence-that-doesnt?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>Resources</h2><ul><li><p><a href="https://www.inkandswitch.com/essay/local-first/">Local-First Software</a>: Kleppmann, M., Wiggins, A., van Hardenberg, P., McGranaghan, M. &#8220;Local-First Software: You Own Your Data, in Spite of the Cloud.&#8221; Ink &amp; Switch / ACM Onward! 2019. The foundational paper on local-first design. Week 7 applied its principles to reliability engineering; Week 8 extends them to logging specifically.</p></li><li><p><a href="http://sqlite.org/">SQLite.org</a>: Public domain, zero-dependency embedded database engine. Appropriate for structured local log storage at any scale below multiple concurrent writers.</p></li><li><p><a href="https://opentelemetry.io/docs/collector/">OpenTelemetry Collector</a>: [Please Verify Link: OpenTelemetry Collector local file exporter documentation]: The OTel Collector&#8217;s file exporter allows writing structured telemetry locally before forwarding to any backend. Configurable as a primary or fallback write path.</p></li><li><p><a href="https://morphic.substack.com/p/receipts-everywhere-trust-without">Morphic: Receipts Everywhere</a>: Week 2 of this arc. Establishes the receipt model this piece extends into logging.</p></li><li><p><a href="https://morphic.substack.com/p/local-first-is-reliability-engineering">Morphic: Local-First Is Reliability Engineering</a>: Week 7. The design principles behind local-first; the logbook is their operational implementation.</p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/the-local-logbook-evidence-that-doesnt/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/the-local-logbook-evidence-that-doesnt/comments"><span>Leave a comment</span></a></p>]]></content:encoded></item><item><title><![CDATA[Local-First Is Reliability Engineering]]></title><description><![CDATA[When your cloud tool is the thing that fails, you need a path that doesn't.]]></description><link>https://morphic.zenone.org/p/local-first-is-reliability-engineering</link><guid isPermaLink="false">https://morphic.zenone.org/p/local-first-is-reliability-engineering</guid><dc:creator><![CDATA[Steve Zenone]]></dc:creator><pubDate>Wed, 01 Apr 2026 13:04:02 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!CPkt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e07ae25-5af3-4e6a-83c0-6c0a5a0f0553_1376x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!CPkt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e07ae25-5af3-4e6a-83c0-6c0a5a0f0553_1376x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!CPkt!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e07ae25-5af3-4e6a-83c0-6c0a5a0f0553_1376x768.png 424w, https://substackcdn.com/image/fetch/$s_!CPkt!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e07ae25-5af3-4e6a-83c0-6c0a5a0f0553_1376x768.png 848w, https://substackcdn.com/image/fetch/$s_!CPkt!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e07ae25-5af3-4e6a-83c0-6c0a5a0f0553_1376x768.png 1272w, https://substackcdn.com/image/fetch/$s_!CPkt!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e07ae25-5af3-4e6a-83c0-6c0a5a0f0553_1376x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!CPkt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e07ae25-5af3-4e6a-83c0-6c0a5a0f0553_1376x768.png" width="1376" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2e07ae25-5af3-4e6a-83c0-6c0a5a0f0553_1376x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1376,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:782139,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://morphic.substack.com/i/192748557?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e07ae25-5af3-4e6a-83c0-6c0a5a0f0553_1376x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!CPkt!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e07ae25-5af3-4e6a-83c0-6c0a5a0f0553_1376x768.png 424w, https://substackcdn.com/image/fetch/$s_!CPkt!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e07ae25-5af3-4e6a-83c0-6c0a5a0f0553_1376x768.png 848w, https://substackcdn.com/image/fetch/$s_!CPkt!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e07ae25-5af3-4e6a-83c0-6c0a5a0f0553_1376x768.png 1272w, https://substackcdn.com/image/fetch/$s_!CPkt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e07ae25-5af3-4e6a-83c0-6c0a5a0f0553_1376x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The tools you can still reach when the network can't.</figcaption></figure></div><p>You&#8217;re locked out of the system that tells you you&#8217;re locked out.</p><p>That&#8217;s not a hypothetical. The cloud tool is up; the status page confirms it. But sync is spinning and you can&#8217;t tell if your work saved. You&#8217;re working in green-indicator purgatory. The tool that was supposed to make you more reliable just became the thing you have to manage.</p><p>This is the failure mode nobody designs against. Not data loss. Not a breach. The mundane version: the tool you depend on becomes unreliable or unaccessible at exactly the moment you need to understand its reliability.</p><p>The last piece ended on a specific problem. When the network&#8217;s interpretation of your work can&#8217;t be trusted, keeping more of what you produce in your own hands stops being a preference. Local-first is what that looks like as an engineering decision.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get new stories when they drop.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>WHAT LOCAL-FIRST ACTUALLY MEANS</h2><p>Local-first isn&#8217;t anti-cloud. It&#8217;s a design tradeoff.</p><p>The core principle is simple: the authoritative copy of your data lives on your device. The cloud is a sync target, not the source of truth. That sounds like a minor distinction. It isn&#8217;t.</p><p>Cloud-first design optimizes for availability across devices, multi-party collaboration, and administrative simplicity. You get access from anywhere, on any device, without managing your own backup. The tradeoff is that your availability is coupled to theirs.</p><p>Local-first optimizes for a different failure mode. Can you still do the work if the network is down? If the provider has an outage? If your account is locked for 72 hours while their support queue clears?</p><p>In a cloud-first world: usually no. In a local-first world: usually yes.</p><p>Your working copy is on disk. The sync layer is a delivery mechanism, not the dependency. If the sync layer fails, you still have the file. You still have the state. You can still do the work.</p><p>That&#8217;s the engineering argument. Not ideology. A reliability requirement you set on your own behalf.</p><div><hr></div><h2>THE FAILURE MODES CLOUD-FIRST CREATES</h2><p>There are three failure categories worth naming explicitly.</p><p><strong>Account lockouts.</strong> Provider security events, billing failures, two-factor authentication bugs, and policy decisions generate lockouts. Duration ranges from hours to weeks. Your data is usually fine. You just can&#8217;t reach it.</p><p><strong>Service degradation.</strong> This is the one that gets underreported. The service is &#8220;up&#8221; according to the status page. Sync is intermittent. Conflict resolution is unclear. You&#8217;re not sure which device has the canonical version of something you edited this morning. You keep working and hope the sync layer sorts it out later.</p><p><strong>Provider discontinuation.</strong> Companies get acquired. Products get shut down. SaaS businesses die quietly. The better-managed ones give you 30 to 90 days and an export path. The others don&#8217;t. Your data is often recoverable. Your workflow isn&#8217;t.</p><p>Each is a different shape of the same problem. You built your reliability model on top of someone else&#8217;s availability guarantee. The guarantee is implicit in the design, not explicit in the contract.</p><p>The Recovery-First piece talked about undo as a feature requirement. Undo requires state you can access. If your state lives only in the cloud, your recovery options are exactly as available as your internet connection. That&#8217;s a dependency chain most tools don&#8217;t surface until it breaks.</p><div><hr></div><h2>THE PATTERNS</h2><p>Local-first doesn&#8217;t require running your own servers. It requires one design constraint: always have a path to the data that doesn&#8217;t need the internet.</p><p>The simplest version: write in plaintext or Markdown. Sync to a cloud provider if you want, but work from the local file. Your editor is your application. The file is your database. When the provider is down, open the file. When you switch providers, move the file.</p><p>A more structured version: tools like Obsidian or Bear store notes locally and optionally sync. The local file is the source of truth. The sync is convenience, not a dependency. If iCloud is degraded, the app still works. You lose sync. You don&#8217;t lose access.</p><p>For operational tooling: SQLite is underused for personal and small-team workflows. A local database that syncs via Dropbox or a similar mechanism handles a surprising amount of structured data. No network required to read or write. Sync happens when connectivity is available.</p><p>The pattern across all three: you own the state. The cloud relationship is additive, not foundational.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!1D8X!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F380f1b17-73ce-4b74-8245-febb4b2813ac_1376x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!1D8X!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F380f1b17-73ce-4b74-8245-febb4b2813ac_1376x768.png 424w, https://substackcdn.com/image/fetch/$s_!1D8X!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F380f1b17-73ce-4b74-8245-febb4b2813ac_1376x768.png 848w, https://substackcdn.com/image/fetch/$s_!1D8X!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F380f1b17-73ce-4b74-8245-febb4b2813ac_1376x768.png 1272w, https://substackcdn.com/image/fetch/$s_!1D8X!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F380f1b17-73ce-4b74-8245-febb4b2813ac_1376x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!1D8X!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F380f1b17-73ce-4b74-8245-febb4b2813ac_1376x768.png" width="1376" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/380f1b17-73ce-4b74-8245-febb4b2813ac_1376x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1376,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:653448,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://morphic.substack.com/i/192748557?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F380f1b17-73ce-4b74-8245-febb4b2813ac_1376x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!1D8X!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F380f1b17-73ce-4b74-8245-febb4b2813ac_1376x768.png 424w, https://substackcdn.com/image/fetch/$s_!1D8X!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F380f1b17-73ce-4b74-8245-febb4b2813ac_1376x768.png 848w, https://substackcdn.com/image/fetch/$s_!1D8X!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F380f1b17-73ce-4b74-8245-febb4b2813ac_1376x768.png 1272w, https://substackcdn.com/image/fetch/$s_!1D8X!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F380f1b17-73ce-4b74-8245-febb4b2813ac_1376x768.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Access is not ownership. One requires an internet connection. The other doesn't.</figcaption></figure></div><div><hr></div><h2>THE TRADEOFFS</h2><p>Local-first shifts certain problems onto you. Name them before choosing the pattern.</p><p><strong>Backup is now your responsibility.</strong> The cloud provider was handling it. If your local copy is on one machine and that machine fails, you&#8217;ve traded one availability risk for another. You need a backup discipline: Time Machine, a second sync target, periodic exports. Something redundant that you control.</p><p><strong>Multi-device sync becomes a coordination problem.</strong> When the authoritative copy is local, moving between devices requires more thought. You can&#8217;t just open a browser tab on another machine. You need the sync to have completed. You need to know which device has the newest version.</p><p><strong>Real-time collaboration doesn&#8217;t work well with local-first.</strong> If you need three people editing the same document simultaneously, you need a coordination layer. Local-first is a poor fit for that pattern. It&#8217;s a good fit for individual and small-team workflows where you&#8217;re not editing the same artifact at the same time.</p><p>The decision rule is direct. If you need offline access, account-lockout resilience, or control over your export path: local-first. If you need real-time multi-party state and can accept the provider&#8217;s availability as your own: cloud-first. Most individual and small-team workflows aren&#8217;t actually the second case, even when they&#8217;re built that way.</p><div><hr></div><h2>THE RELIABILITY ARGUMENT</h2><p>Engineers don&#8217;t argue about resilience as a philosophy. They design for failure modes.</p><p>Local-first is what you build when the failure mode you can&#8217;t tolerate is &#8220;I can&#8217;t work because my provider is down.&#8221; That&#8217;s a real failure mode. It happens regularly. Building around it is engineering, not stubbornness.</p><p>The Cognitive Budget piece talked about what it costs when tools that are supposed to help start requiring maintenance attention themselves. Local-first tools tend to have a smaller failure surface. No API to break. No sync conflict to debug. No authentication token to expire. That&#8217;s not nothing.</p><p>The Analog Executive piece opened with a simple observation: being unreachable is now a luxury. Local-first is the technical version of the same principle. Owning your state means you&#8217;re not subject to someone else&#8217;s availability schedule. That&#8217;s a form of independence that compounds over time.</p><p>Keep one path to safety that doesn&#8217;t require the internet. That&#8217;s not a retreat from modern infrastructure. That&#8217;s a reliability requirement you set on your own behalf.</p><p>Week 8 is about the logbook: what it means to keep records that don&#8217;t require permission to read.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/local-first-is-reliability-engineering?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/local-first-is-reliability-engineering?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><h2>Resources</h2><ul><li><p><a href="https://www.inkandswitch.com/essay/local-first/">Local First Software</a> - Kleppmann, M., Wiggins, A., van Hardenberg, P., McGranaghan, M. &#8220;Local-First Software: You Own Your Data, in Spite of the Cloud.&#8221; Ink &amp; Switch / ACM Onward! 2019. The foundational paper defining the seven ideals of local-first software design.</p></li><li><p><a href="http://obsidian.md/">Obsidian.md</a> - plaintext Markdown note-taking with local storage and optional sync.</p></li><li><p><a href="http://sqlite.org/">SQLite.org</a> - public domain, zero-dependency embedded database engine. Appropriate for local-first structured data at any scale below &#8220;multiple writers simultaneously.&#8221;</p></li><li><p><a href="https://automerge.org">Automerge</a> - CRDT (Conflict-free Replicated Data Type) library for building local-first sync without a central coordinator. The most accessible implementation of the theoretical foundation described in the Kleppmann et al. paper. </p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/local-first-is-reliability-engineering/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/local-first-is-reliability-engineering/comments"><span>Leave a comment</span></a></p>]]></content:encoded></item><item><title><![CDATA[The Detector Is Not the Evidence]]></title><description><![CDATA[AI text detectors report confidence. The research says 39.5% accuracy. Those are different things.]]></description><link>https://morphic.zenone.org/p/the-detector-is-not-the-evidence</link><guid isPermaLink="false">https://morphic.zenone.org/p/the-detector-is-not-the-evidence</guid><dc:creator><![CDATA[Steve Zenone]]></dc:creator><pubDate>Sat, 28 Mar 2026 18:42:58 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!y34k!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6177ad8-a252-4a12-99cf-1787f5863a96_2048x1143.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!y34k!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6177ad8-a252-4a12-99cf-1787f5863a96_2048x1143.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!y34k!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6177ad8-a252-4a12-99cf-1787f5863a96_2048x1143.png 424w, https://substackcdn.com/image/fetch/$s_!y34k!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6177ad8-a252-4a12-99cf-1787f5863a96_2048x1143.png 848w, https://substackcdn.com/image/fetch/$s_!y34k!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6177ad8-a252-4a12-99cf-1787f5863a96_2048x1143.png 1272w, https://substackcdn.com/image/fetch/$s_!y34k!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6177ad8-a252-4a12-99cf-1787f5863a96_2048x1143.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!y34k!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6177ad8-a252-4a12-99cf-1787f5863a96_2048x1143.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b6177ad8-a252-4a12-99cf-1787f5863a96_2048x1143.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:375899,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://morphic.substack.com/i/192442631?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6177ad8-a252-4a12-99cf-1787f5863a96_2048x1143.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!y34k!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6177ad8-a252-4a12-99cf-1787f5863a96_2048x1143.png 424w, https://substackcdn.com/image/fetch/$s_!y34k!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6177ad8-a252-4a12-99cf-1787f5863a96_2048x1143.png 848w, https://substackcdn.com/image/fetch/$s_!y34k!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6177ad8-a252-4a12-99cf-1787f5863a96_2048x1143.png 1272w, https://substackcdn.com/image/fetch/$s_!y34k!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6177ad8-a252-4a12-99cf-1787f5863a96_2048x1143.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The system has a verdict. The verdict has a confidence score. The confidence score is a product decision.</figcaption></figure></div><p>A professor submits a student paper to an AI detector. The tool returns: 73% AI-generated. The professor files a misconduct report.</p><p>The student wrote every word.</p><p>This is not a hypothetical. The Markup documented this in 2023: international students flagged for academic misconduct because the way they write English, carefully, with simpler sentence structures and repeated phrasing, scores as statistically consistent with AI output. The detector was confident. The detector was wrong. And the confidence is what made it dangerous.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Get new stories when they drop.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>WHAT DETECTORS ACTUALLY DO</h2><p>AI text detectors do not read writing the way a human does. They measure perplexity: how surprising is each word choice given what came before? High perplexity means unpredictable word choices. Low perplexity means predictable ones. AI-generated text tends toward low perplexity because language models select statistically likely words.</p><p>The problem: careful writing, formal writing, second-language writing, all tend toward low perplexity too. Simple sentences. Conservative vocabulary. Consistent structure. The detector has no way to separate &#8220;human writing that looks like AI output&#8221; from &#8220;AI output.&#8221; It calculates a probability. It presents that probability as a verdict.</p><p>Stanford professor James Zou and colleagues published on this directly. Non-native English writers had false positive rates as high as 61.3%. For international students specifically, the risk of a false accusation could reach 97%. The tool was not broken. It was doing exactly what it was built to do. The error was in treating the probability as evidence of intent.</p><p>Perkins, Roe, et al. tested seven of the most widely used detectors across 114 samples in 2024. Baseline accuracy: 39.5%. After applying simple adversarial techniques, accuracy dropped to 17.4%. The tools were already unreliable before anyone tried to game them. They collapsed faster when someone did.</p><p>A system with 39.5% accuracy in controlled testing is not a verification tool. It is a coin flip with a confidence display.</p><div><hr></div><h2>THE SAME ERROR, TWICE</h2><p>Here is where the argument gets uncomfortable for both sides.</p><p>The people being accused of using AI to write are are no different than those trusting AI output without scrutinizing it either. They generate, they paste, they publish. They rely on the model&#8217;s output as good enough without reading it for what it actually claims or whether it represents their actual argument. The model said it, so it&#8217;s done.</p><p>Both groups are making identical moves. Treating a probabilistic system&#8217;s output as a finished verdict rather than a signal that requires judgment.</p><p>The AI user trusts the generator. The AI skeptic trusts the detector. Both treat the model&#8217;s confidence as their own. Neither is asking what the model can actually know.</p><p>The symmetric failure is the real story. Not the question of whether AI use is ethical or lazy or acceptable. Both of those debates are downstream of the same miscalibration.</p><div><hr></div><h2>A NOTE ON THE TOOLS</h2><p>I&#8217;ve built data analytics systems for significant companies. Classification models. Detection pipelines. Anomaly scoring. I know what these tools look like from the inside, and I know what gets left out of the confidence display.</p><p>Even enterprise-grade systems, built with proper data science teams and substantial validation cycles, routinely operate at error rates that would make most non-technical users uncomfortable if they saw the raw numbers. False positive budgets get set as business decisions, not scientific ones. Thresholds get tuned to minimize one error type while hiding another. The confidence display is a product choice. Not a measurement.</p><p>AI text detectors are not enterprise-grade. They are first-generation, perplexity-based classifiers deployed into institutional contexts with institutional authority. Perkins et al.&#8217;s 39.5% baseline accuracy would not pass internal review at most serious analytics organizations. It would get sent back with a note: not production-ready.</p><p>Here is the part that should make the skeptic pause. <strong>The person who distrusts AI enough to run a text detector is relying on another statistical model to confirm that distrust. They have not escaped the problem. They have doubled it.</strong> Trusting a generator&#8217;s probabilistic output and a detector&#8217;s probabilistic output simultaneously, and calling the combination proof.</p><p>That is not skepticism. That is symmetrical credulity.</p><div><hr></div><h2>WHEN CALIBRATION BECOMES A TEAM</h2><p>The problem becomes structural when it turns tribal.</p><p>There is a documented polarization in how AI is discussed: accelerationists who treat adoption as obviously correct, doomers and skeptics who treat AI use as obviously suspect. Forbes documented the growing heat between these camps in February 2025. The discourse resembles a political argument more than a technical one. Positions are held first. Evidence is consulted only when it confirms what was already believed.</p><p>The AI detector user who sees a high probability score and files a misconduct report is not evaluating evidence. They are confirming a prior. The number crossed a threshold. That was enough.</p><p>This is the same mechanism that makes political misinformation stick. Find a data point that supports the position. Stop looking. A 73% AI-probability rating becomes &#8220;this person cheated&#8221; by the same logic that makes a misleading chart feel like proof. The source is unreliable. The conclusion is certain. The certainty is the tell.</p><p>What is driving this is not a genuine technical disagreement about AI capability. It is an epistemological failure that tribalism accelerates. Both sides are wielding bad tools. Only one side is currently losing jobs and grades over the verdicts.</p><div><hr></div><h2>WE HAVE SEEN THIS BEFORE</h2><p>In the late 1990s, I took a course at a local college on digital photography and Adobe Photoshop.</p><p>The arguments happening then sound familiar.</p><p>Digital photography was not &#8220;real&#8221; photography. You were not earning your images the way film photographers earned them. Photoshop was cheating: it let anyone who could click a mouse claim a skill that took years to develop with a darkroom and chemical trays. You were not a true photographer if you were working in pixels. The gatekeeping was confident, defensive, and grounded in a very human anxiety. The new tool was going to devalue what the prior tool had cost to master.</p><p>Those critics were simply wrong.</p><p>Digital photography did not destroy photography. Photoshop did not make every image meaningless. Both expanded what was possible and shifted where skill actually lived. The people who needed film processing to feel like artists found their craft was more portable than they feared. The people who were mostly relying on the mystique of film without the underlying skill got found out faster. The tool change clarified where the actual work was.</p><p>Twenty-five years later, the same argument is running again. In the same voice.</p><p>AI is not &#8220;real&#8221; writing. Using it is cheating. You are not a true writer if a language model touched your draft. The confidence is the same. The defensiveness is the same. The underlying anxiety, that a new tool devalues what the prior method cost, is identical.</p><p>The Photoshop debate got resolved by time. You can watch the resolution from here. The AI detector gives the skeptic a device that feels like evidence. It does not change what history says about where these arguments end up.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!C288!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ec08f27-3a2a-48d2-a70a-c9b51c9f2af2_2048x1143.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!C288!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ec08f27-3a2a-48d2-a70a-c9b51c9f2af2_2048x1143.png 424w, https://substackcdn.com/image/fetch/$s_!C288!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ec08f27-3a2a-48d2-a70a-c9b51c9f2af2_2048x1143.png 848w, https://substackcdn.com/image/fetch/$s_!C288!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ec08f27-3a2a-48d2-a70a-c9b51c9f2af2_2048x1143.png 1272w, https://substackcdn.com/image/fetch/$s_!C288!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ec08f27-3a2a-48d2-a70a-c9b51c9f2af2_2048x1143.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!C288!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ec08f27-3a2a-48d2-a70a-c9b51c9f2af2_2048x1143.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9ec08f27-3a2a-48d2-a70a-c9b51c9f2af2_2048x1143.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2635763,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://morphic.substack.com/i/192442631?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ec08f27-3a2a-48d2-a70a-c9b51c9f2af2_2048x1143.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!C288!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ec08f27-3a2a-48d2-a70a-c9b51c9f2af2_2048x1143.png 424w, https://substackcdn.com/image/fetch/$s_!C288!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ec08f27-3a2a-48d2-a70a-c9b51c9f2af2_2048x1143.png 848w, https://substackcdn.com/image/fetch/$s_!C288!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ec08f27-3a2a-48d2-a70a-c9b51c9f2af2_2048x1143.png 1272w, https://substackcdn.com/image/fetch/$s_!C288!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9ec08f27-3a2a-48d2-a70a-c9b51c9f2af2_2048x1143.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The argument running in 2025 ran in 1995. Different tool. Same anxiety. Same wrong conclusion.</figcaption></figure></div><div><hr></div><h2>THE TOOL THAT ACTUALLY WORKS</h2><p>I use AI to create images. I am not a trained visual artist.</p><p>What AI image tools gave me was not skill I did not have. It was the ability to close the gap between what I could see in my head and what I could put into the world. The specific mood. The composition. The light quality and texture and tonal feel of a visual idea that existed clearly in my mind but stopped at my hands. For years, that gap between aesthetic judgment and technical execution just meant the thing stayed in my head.</p><p>That gap is now crossable.</p><p>This is not the same thing as what AI text detectors do. But it is the same argument about AI as a tool. I do not prompt once and accept whatever the model returns. The iteration, the rejection, the direction-giving, the refinement cycle toward a specific result: that is the work. The model did not supply the vision. It executed the vision, approximately, until approximately became right.</p><p>The Photoshop parallel holds here. Photoshop did not replace photographers who understood light and composition. It gave people who understood those things more direct access to the result. AI image tools do the same. The skill that matters did not disappear. The bottleneck that was hiding it did.</p><blockquote><p><strong>If your objection to tools that expand the range of what people can express is really about the expansion itself, that is a different conversation entirely.</strong></p></blockquote><div><hr></div><h2>WHAT HAPPENS WHEN SYSTEMS ACT ON THIS</h2><p>The AI detector verdict is already embedded in institutional decisions. Academic misconduct cases. Content moderation queues. Hiring filters. Fraud detection pipelines.</p><p>Each of those is a system built on a signal with 39.5% baseline accuracy. Each downstream consequence carries the confidence of a verified finding. None of that confidence is warranted by the underlying tool.</p><p>This is the same pattern from the last piece. Surveillance does not measure behavior. It measures behavior under observation. AI detectors do not measure authorship. They measure statistical patterns that correlate with AI output under specific conditions. Both produce a number. Both present the number as ground truth. Both systems are wrong about what the number means.</p><p>The institution acting on these verdicts is not asking whether the detection layer is reliable. It is asking whether the threshold was crossed.</p><div><hr></div><h2>THE CALIBRATION PROBLEM</h2><p>The argument is not that AI detectors are useless. The argument is that they are probabilistic tools being used as forensic ones.</p><p>A weather model that outputs 70% rain probability is useful. Canceling an outdoor event on a clear morning because the threshold was crossed is a calibration failure. The model gave a signal. The human replaced their judgment with it.</p><p>The 73% AI-generated verdict is a signal. The correct response is more questions: Is this a non-native English writer? Is this a highly formal document type? Has the text been edited from AI output or written from scratch in a constrained style? The signal closes none of those questions. But it is being used to close the inquiry entirely.</p><p>Calibration failures become permanent when they get institutionalized. When the threshold becomes policy, the error rate becomes the policy&#8217;s error rate. Nobody audits that.</p><p>When the network&#8217;s interpretation of your work cannot be trusted, keeping more of what you produce in your own hands stops being a preference. It becomes how you stay legible to yourself, and to anyone who needs to understand what you actually built.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/the-detector-is-not-the-evidence?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/the-detector-is-not-the-evidence?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div class="instagram-embed-wrap" data-attrs="{&quot;instagram_id&quot;:&quot;DWcUXLskq9T&quot;,&quot;title&quot;:&quot;ZenOne on Instagram: \&quot;They said the 808 would kill music.\n\nTech&#8230;&quot;,&quot;author_name&quot;:&quot;@zenone.music&quot;,&quot;thumbnail_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/__ss-rehost__IG-meta-DWcUXLskq9T.jpg&quot;,&quot;like_count&quot;:null,&quot;comment_count&quot;:null,&quot;profile_pic_url&quot;:null,&quot;follower_count&quot;:null,&quot;timestamp&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="InstagramToDOM"></div><div><hr></div><h2>Resources</h2><ul><li><p><strong><a href="https://hai.stanford.edu/news/ai-detectors-biased-against-non-native-english-writers">AI-Detectors Biased Against Non-Native English Writers &#8212; Stanford HAI (2023)</a></strong> &#8212; James Zou et al.; 61.3% false positive rate for non-native English speakers; up to 97% false accusation potential for international students in academic settings.</p></li><li><p><strong><a href="https://arxiv.org/abs/2403.19148">GenAI Detection Tools, Adversarial Techniques and Implications for Inclusivity in Higher Education &#8212; Perkins, Roe, et al. (2024)</a></strong> &#8212; International Journal of Educational Technology in Higher Education (Springer Nature). Baseline accuracy 39.5%, drops to 17.4% under simple adversarial techniques; 7 detectors, 114 samples, 805 tests.</p></li><li><p><strong><a href="https://themarkup.org/machine-learning/2023/08/14/ai-detection-tools-falsely-accuse-international-students-of-cheating">AI Detection Tools Falsely Accuse International Students of Cheating &#8212; The Markup (2023)</a></strong> &#8212; Documented cases of students flagged for misconduct by AI detectors.</p></li><li><p><strong><a href="https://www.forbes.com/sites/lanceeliot/2025/02/18/ai-doomers-versus-ai-accelerationists-locked-in-battle-for-future-of-humanity/">AI Doomers Versus AI Accelerationists Locked In Battle For Future Of Humanity &#8212; Forbes (2025)</a></strong> &#8212; Documented polarization between AI skeptics and accelerationists.</p></li><li><p><strong><a href="https://nppa.org/">Ethics in the Age of Digital Photography &#8212; National Press Photographers Association</a></strong> &#8212; NPPA&#8217;s ongoing ethics standards around digital manipulation; the photojournalism community&#8217;s documented struggle with Photoshop norms through the 1990s and 2000s.</p></li><li><p><strong><a href="https://morphic.substack.com/p/privacy-is-a-luxury-good">Privacy Is a Luxury Good &#8212; Morphic (2026)</a></strong> &#8212; Previous arc piece: surveillance creates distorted signal; the most monitored users produce the most corrupted data.</p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://morphic.zenone.org/p/the-detector-is-not-the-evidence/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://morphic.zenone.org/p/the-detector-is-not-the-evidence/comments"><span>Leave a comment</span></a></p>]]></content:encoded></item></channel></rss>