<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Research — Pavel Gurov</title>
    <link>https://gurovdigital.com/research</link>
    <description>Notes on AI, AI agents and information integrity, written by researcher Pavel Gurov.</description>
    <language>en</language>
    <atom:link href="https://gurovdigital.com/research/rss.xml" rel="self" type="application/rss+xml"/>
    <item>
      <title>Polycrisis: how the European Commission named our era</title>
      <link>https://gurovdigital.com/research/polycrisis-european-commission-naming-eras</link>
      <guid isPermaLink="true">https://gurovdigital.com/research/polycrisis-european-commission-naming-eras</guid>
      <pubDate>Tue, 28 Jul 2026 00:00:00 +0000</pubDate>
      <dc:creator>Pavel Gurov</dc:creator>
      <category>Information integrity</category>
      <category>Counter-disinformation</category>
      <category>Cognition</category>
      <description>And what future historians will call it instead</description>
      <content:encoded><![CDATA[<p>_Nobody names their own era._</p>
<h2 id="where-the-term-actually-comes-from">Where the term actually comes from</h2>
<p>Polycrisis was coined by Edgar Morin, the French philosopher of complexity, with Anne-Brigitte Kern, in _Terre-Patrie_ (Seuil, 1993). The English translation, _Homeland Earth: A Manifesto for the New Millennium_, followed in 1999, which is why the date is often reported wrongly. One book, two years.</p>
<p>Their formulation was planetary and structural. There is no single vital problem, they argued, but many, and it is the complex intersolidarity of problems, antagonisms, crises and uncontrolled processes that constitutes the vital problem.</p>
<p>Then the term sat almost unused for two decades.</p>
<h2 id="the-european-commission-put-it-into-policy-language">The European Commission put it into policy language</h2>
<p>The step most accounts skip is the one that matters most institutionally. In 2016, Jean-Claude Juncker, then President of the European Commission, adopted the word to describe a combination of crises facing the European Union at the same time and in ways that fed each other: security threats in the neighbourhood and at home, the refugee crisis, and the United Kingdom referendum.</p>
<p>Juncker&#39;s usage differs from Morin&#39;s in a way that is easy to miss and consequential. Morin described a single planetary condition. Juncker described a specific, bounded, governable situation, which leaves the door open to the idea that there can be several smaller polycrises at once. Scholars of the EU picked the term up from him almost immediately.</p>
<p>So the sequence is Morin 1993, Juncker 2016, Davos 2023. The word entered policy vocabulary through a European institution six years before it became a conference buzzword.</p>
<h2 id="then-the-historian-and-then-the-doubt">Then the historian, and then the doubt</h2>
<p>Adam Tooze is the person most widely credited with popularising it, after a prominent segment at Davos in 2023. His version is planetary like Morin&#39;s, not bounded like Juncker&#39;s, but he dates the beginning of the global polycrisis to around 2008, where Morin and Kern placed it decades earlier.</p>
<figure class="rsrch-figure"><img src="/research/figures/polycrisis-risk-map.png" alt="The World Economic Forum&#39;s map of how global risks interact, from the Global Risks Report 2023." width="1456" height="1172" loading="lazy" decoding="async"><figcaption>The World Economic Forum&#39;s map of how global risks interact, from the Global Risks Report 2023. World Economic Forum, Global Risks Perception Survey 2022–2023, Global Risks Report 2023. <a href="https://www.weforum.org/publications/global-risks-report-2023/" target="_blank" rel="noopener noreferrer">weforum.org</a>. <a href="https://www.gesetze-im-internet.de/urhg/__51.html" target="_blank" rel="noopener noreferrer">Quoted under §51 UrhG</a>.</figcaption></figure>
<p>The Davos version of the word arrives with a picture attached, and the picture is doing a lot of the persuading: thirty-odd risks, five colour-coded categories, and a thicket of edges between them. It is a good illustration of entanglement. It is not evidence of a period.</p>
<p>Note what that means. Three serious users of the same word disagree about when the thing it names began, by margins of fifteen years and more. A term whose referent has no agreed start date is not yet a historical period. It is a mood with a vocabulary.</p>
<p>And then, in September 2025, Tooze himself began to have second thoughts about whether polycrisis still made sense. If everything is going to hell, he asked, why have the markets stayed eerily calm? Where, in other words, is the crisis?</p>
<p>That is the empirical hook. We are not speculating about whether the label will survive. We are watching the person who popularised it examine the joins.</p>
<h2 id="the-pattern-is-not-new">The pattern is not new</h2>
<p>People do not name the period they are living through. They name what they are experiencing, and experience is a poor guide to what a period turns out to have been.</p>
<p>The First World War was the Great War for most of its duration, and the war to end war after H. G. Wells attached the phrase. Nobody numbered it, because numbering implies a sequel. The ordinal appears in Charles à Court Repington&#39;s account of the conflict, published in 1920 as _The First World War 1914–1918_ — a title chosen, by his own account, in the gloomy expectation that another would follow. He was naming it against the mood of the moment, and he was right.</p>
<p>The second one was named faster, and differently in different places. In the Soviet Union it became the Great Patriotic War from June 1941, a name for a war of national survival, not for a global sequence. In the United States it settled as World War II. Two names for overlapping events, each encoding what the naming country thought the war was about.</p>
<p>Which is the point. A period&#39;s name is not a description of its contents. It is a compression of its outcome, applied afterwards, by whoever ends up doing the compressing.</p>
<h2 id="what-the-2020s-might-end-up-called">What the 2020s might end up called</h2>
<p>If the decade resolves into a wider war, the current vocabulary vanishes and the period gets an ordinal, exactly as 1914 did. If it resolves into economic and institutional fragmentation, something like deglobalisation is the likely label. If artificial intelligence turns out to be the dominant variable, the decade becomes a prologue and gets named for what followed. And there is _the Great Unravelling_, already in use among ecologists, which happens to describe the simultaneous loosening of institutions, alliances, climate stability and informational consensus rather well.</p>
<p>I have no idea which. Neither does anyone else, and that is the argument.</p>
<h2 id="why-this-belongs-in-a-research-notebook-rather-than-a-comment-page">Why this belongs in a research notebook rather than a comment page</h2>
<p>There is a methodological consequence, and it is not abstract for anyone working on information integrity.</p>
<p>If a period cannot be named from inside it, then every real-time frame applied to ongoing events is provisional. That includes the frames we use for influence operations, platform failures and institutional decay while they are happening. The categories we are currently using to sort what is going on are the equivalent of calling it the Great War in 1916: useful, sincere, and probably not what it will be called.</p>
<p>The practical discipline that follows is to hold the frame loosely and the evidence tightly. Date things. Cite things. Describe mechanisms, not eras. Mechanisms survive renaming; eras do not.</p>
<h2>References</h2>
<ul><li>Morin, E., Kern, A.-B., _Terre-Patrie_ (Éditions du Seuil, 1993); English translation _Homeland Earth: A Manifesto for the New Millennium_ (Hampton Press, 1999)</li><li>Lawrence, M., Homer-Dixon, T., et al., _Global Polycrisis: The Causal Mechanisms of Crisis Entanglement_, Cascade Institute — <a href="https://cascadeinstitute.org/technical-paper/global-polycrisis/" target="_blank" rel="noopener noreferrer">https://cascadeinstitute.org/technical-paper/global-polycrisis/</a></li><li>_Economic Globalization&#39;s Polycrisis_, <strong>International Studies Quarterly</strong> 68(2), sqae024 (2024) — on the term&#39;s origins, Juncker&#39;s 2016 usage, and Tooze&#39;s divergent dating — <a href="https://academic.oup.com/isq/article/68/2/sqae024/7634048" target="_blank" rel="noopener noreferrer">https://academic.oup.com/isq/article/68/2/sqae024/7634048</a></li><li>World Economic Forum, _Global Risks Report 2023_ — the polycrisis risk-interaction map — <a href="https://www.weforum.org/publications/global-risks-report-2023/" target="_blank" rel="noopener noreferrer">https://www.weforum.org/publications/global-risks-report-2023/</a></li><li>World Economic Forum, _This is why &#39;polycrisis&#39; is a useful way of looking at the world right now_ — interview with Adam Tooze, March 2023 — <a href="https://www.weforum.org/stories/2023/03/polycrisis-adam-tooze-historian-explains/" target="_blank" rel="noopener noreferrer">https://www.weforum.org/stories/2023/03/polycrisis-adam-tooze-historian-explains/</a></li><li>Repington, C. à C., _The First World War 1914–1918_ (Constable, 1920) — the earliest widely cited use of the ordinal</li><li>_More than a buzzword? Mapping interpretations of the &#39;polycrisis&#39;_, <strong>Sustainability Science</strong> (2025) — <a href="https://link.springer.com/article/10.1007/s11625-025-01790-9" target="_blank" rel="noopener noreferrer">https://link.springer.com/article/10.1007/s11625-025-01790-9</a></li></ul>]]></content:encoded>
    </item>
    <item>
      <title>Jacobian space: the hidden reasoning layer found inside Claude, and language prediction in the unconscious brain</title>
      <link>https://gurovdigital.com/research/jacobian-space-hidden-reasoning-layer</link>
      <guid isPermaLink="true">https://gurovdigital.com/research/jacobian-space-hidden-reasoning-layer</guid>
      <pubDate>Tue, 28 Jul 2026 00:00:00 +0000</pubDate>
      <dc:creator>Pavel Gurov</dc:creator>
      <category>Claude</category>
      <category>Cognition</category>
      <description>A privileged internal workspace inside a language model, and language prediction in a brain that is not conscious. Read together, they narrow the gap from both ends.</description>
      <content:encoded><![CDATA[<p>Two results landed within weeks of each other. Separately they are interesting. Together they are harder to dismiss, because they narrow the same gap from opposite directions.</p>
<h2 id="one-a-global-workspace-inside-the-model">One — a global workspace inside the model</h2>
<p>On 6 July 2026, Anthropic&#39;s interpretability team published <em>Verbalizable Representations Form a Global Workspace in Language Models</em>. The method is a new tool called the Jacobian lens, or J-lens, which identifies internal representations that are verbalizable: poised to be spoken about, should the occasion arise, as distinct from those that merely happen to be spoken about in one particular context. The averaging step is what makes the distinction possible.</p>
<p>What they found is that this set does considerably more than support speech. It behaves like a global workspace: the small privileged buffer that cognitive science has long used to describe the contents of a mind that are available to report and to reasoning. The representations in this space, which the team calls J-space, are reportable, steerable, used in internal reasoning, broadcast widely, and not required for automatic processing.</p>
<p>Two numbers keep this honest. J-space carries under ten per cent of activation variance and holds on the order of twenty-five concepts at a time, and it appears only in the middle block of the network. It is a thin slice, not the whole model.</p>
<figure class="rsrch-figure"><img src="/research/figures/jspace-swap-window.jpeg" alt="Effect of swapping a concept inside J-space, measured as the change in log-probability across layers." width="1102" height="828" loading="lazy" decoding="async"><figcaption>Effect of swapping a concept inside J-space, measured as the change in log-probability across layers. Gurnee, Sofroniew, Pearce et al., Verbalizable Representations Form a Global Workspace in Language Models, Transformer Circuits Thread, Anthropic, 2026. <a href="https://transformer-circuits.pub/2026/workspace/index.html" target="_blank" rel="noopener noreferrer">transformer-circuits.pub</a>. <a href="https://www.gesetze-im-internet.de/urhg/__51.html" target="_blank" rel="noopener noreferrer">Quoted under §51 UrhG</a>.</figcaption></figure>
<p>It is also not the chain of thought you see in a chatbot. It is a layer underneath, made of concepts the model is holding and operating on without writing them down.</p>
<p>Here is where the reporting went wrong, my own first reaction included. This is not evidence of subjective experience. The paper&#39;s claim is narrow and falsifiable: a privileged subspace exists that satisfies the functional criteria Global Workspace Theory sets out for conscious access. Functional access, not phenomenal experience. The distinction is the entire argument, and collapsing it is how a careful result becomes a bad headline.</p>
<p>What it does give us is practical: a place to look. If a model is tracking something it has not said, that a user appears to be probing it for example, this is where you would expect that to be legible.</p>
<h2 id="two-language-prediction-in-an-unconscious-brain">Two — language prediction in an unconscious brain</h2>
<p>On 6 May 2026, <em>Nature</em> published <em>Plasticity and language in the anaesthetized human hippocampus</em> by Kalman Katlowitz and colleagues at Baylor College of Medicine in Houston.</p>
<p>Seven patients undergoing epilepsy surgery had Neuropixels probes placed in the hippocampus: high-density microelectrodes that record hundreds of individual neurons, not an average across a population. While the patients were under general anaesthesia, the team played them tones, and then clips from educational videos and storytelling podcasts.</p>
<figure class="rsrch-figure"><img src="/research/figures/anaesthetised-hippocampus.jpeg" alt="Neuropixels recordings from the human hippocampus under general anaesthesia." width="1252" height="1534" loading="lazy" decoding="async"><figcaption>Neuropixels recordings from the human hippocampus under general anaesthesia. Katlowitz, K. A., Cole, E. R., Mickiewicz, E. A. et al., Nature 654, 714–723 (2026). <a href="https://doi.org/10.1038/s41586-026-10448-0" target="_blank" rel="noopener noreferrer">doi.org/10.1038/s41586-026-10448-0</a>. <a href="https://creativecommons.org/licenses/by-nc-nd/4.0/" target="_blank" rel="noopener noreferrer">Licensed under CC BY-NC-ND 4.0</a>.</figcaption></figure>
<p>The unconscious hippocampus discriminated oddball tones, and the effect grew over roughly ten minutes of the experiment. That is representational plasticity: learning, without awareness. In the language phase, neural firing patterns distinguished parts of speech and showed evidence of anticipating upcoming words.</p>
<p>Predictive coding of the kind we associate with being awake and attentive, occurring in a state with no awareness, no memory and no capacity to act.</p>
<h2 id="why-they-belong-on-the-same-page">Why they belong on the same page</h2>
<p>One result finds a reportable, reasoning-adjacent workspace inside a system built on next-token prediction. The other finds next-word prediction running in a biological system that has, at that moment, no consciousness at all.</p>
<p>The convenient position, that prediction is what machines do and understanding is what minds do, gets squeezed from both sides. Neither paper claims to resolve it. Both make the question sharper, which is the more useful outcome.</p>
<h2>References</h2>
<ul><li>Gurnee, W., Sofroniew, N., Pearce, A., et al., <em>Verbalizable Representations Form a Global Workspace in Language Models</em>, Transformer Circuits Thread, Anthropic, 6 July 2026 — <a href="https://transformer-circuits.pub/2026/workspace/index.html" target="_blank" rel="noopener noreferrer">https://transformer-circuits.pub/2026/workspace/index.html</a></li><li>Anthropic, <em>A global workspace in language models</em>, research summary — <a href="https://www.anthropic.com/research/global-workspace" target="_blank" rel="noopener noreferrer">https://www.anthropic.com/research/global-workspace</a></li><li>Katlowitz, K. A., et al., <em>Plasticity and language in the anaesthetized human hippocampus</em>, <strong>Nature</strong> 654, 714–723 (2026) — <a href="https://doi.org/10.1038/s41586-026-10448-0" target="_blank" rel="noopener noreferrer">https://doi.org/10.1038/s41586-026-10448-0</a></li><li><em>Even the unconscious brain can learn — and predict what you&#39;ll say next</em>, Nature news, 6 May 2026 — <a href="https://www.nature.com/articles/d41586-026-01465-0" target="_blank" rel="noopener noreferrer">https://www.nature.com/articles/d41586-026-01465-0</a></li><li>Baars, B. J., <em>A Cognitive Theory of Consciousness</em> (Cambridge University Press, 1988) — the original statement of Global Workspace Theory</li></ul>]]></content:encoded>
    </item>
    <item>
      <title>Three claims about AI that survive scrutiny: the intelligence definition, the pronoun problem, and Google&apos;s shelved LaMDA</title>
      <link>https://gurovdigital.com/research/intelligence-definition-pronouns-lamda</link>
      <guid isPermaLink="true">https://gurovdigital.com/research/intelligence-definition-pronouns-lamda</guid>
      <pubDate>Tue, 28 Jul 2026 00:00:00 +0000</pubDate>
      <dc:creator>Pavel Gurov</dc:creator>
      <category>Cognition</category>
      <category>Gemini</category>
      <description>The term itself does not survive an academic audit, the hardest word in English for a language model is two letters long, and the company that got there first archived it.</description>
      <content:encoded><![CDATA[<p>Three claims I keep returning to, because each one is more load-bearing than it first appears.</p>
<h2 id="one-artificial-intelligence-is-not-a-scientific-term">One — &quot;artificial intelligence&quot; is not a scientific term</h2>
<p>It is a commercial one. The problem is the second word.</p>
<p>In 1997 fifty-two researchers signed a consensus statement on intelligence, published in <em>Intelligence</em> and led by Linda Gottfredson of the University of Delaware. It defines intelligence as a very general mental capability that, among other things, involves the ability to reason, plan, solve problems, think abstractly, comprehend complex ideas, learn quickly and learn from experience. It is not merely book learning, a narrow academic skill, or test-taking smarts; it reflects a broader and deeper capability for comprehending our surroundings.</p>
<p>Hold current systems against that list honestly and the fit is partial. Some capabilities are demonstrated, others are contested, and at least one is not present at all in a deployed model: learning from experience in the ordinary sense, as opposed to being retrained. Weights are frozen. What looks like learning inside a conversation is context, not acquisition.</p>
<p>None of which makes the systems less useful. It makes the label imprecise. Something closer to &quot;algorithmic modelling&quot; would describe what is actually happening. It also does not sell, which is most of the answer to why nobody uses it.</p>
<h2 id="two-the-hardest-word-in-english-for-a-language-model-is-it">Two — the hardest word in English for a language model is &quot;it&quot;</h2>
<p>Pronouns are the difficult case, and the reason is that they carry almost no meaning of their own. They point. Resolving one requires holding the surrounding situation in mind and deciding what is being pointed at.</p>
<p>This is why pronoun resolution became a benchmark and not a footnote. Terry Winograd&#39;s original example is the canonical one: <em>the city councilmen refused the demonstrators a permit because they feared violence</em>, versus <em>because they advocated violence</em>. One word changes, and the referent of &quot;they&quot; flips. No amount of syntax gets you there. You need a model of who fears things and who advocates things.</p>
<p>Hector Levesque proposed the Winograd Schema Challenge in 2011 as an alternative to the Turing test, precisely because it resists the tricks that let a system pass by deflection. Machines performed near chance for years. Systems now score well above that, which is a real result — though later analysis showed some of the gain came from statistical regularities in the datasets, not from the reasoning the benchmark was designed to isolate.</p>
<h2 id="three-google-built-it-first-and-shelved-it">Three — Google built it first and shelved it</h2>
<p>LaMDA was Google&#39;s conversational model, announced in 2021 and never released as a public product. The capability was there. The judgement about what it was worth was not.</p>
<p>Google had reason for caution. A chatbot that says something indefensible is a reputational problem for a company whose product is trust in search results, and the same conservatism had already cost them elsewhere. What followed is now the standard case study: OpenAI shipped ChatGPT in November 2022, built on the transformer architecture that Google researchers had published in 2017, and the market reorganised itself around a category Google had reached first and chosen not to enter.</p>
<figure class="rsrch-figure"><img src="/research/figures/transformer-architecture.png" alt="The transformer architecture as Google published it in 2017." width="1520" height="2239" loading="lazy" decoding="async"><figcaption>The transformer architecture as Google published it in 2017. Vaswani, Shazeer, Parmar et al., Attention Is All You Need, 2017, fig. 1. <a href="https://arxiv.org/abs/1706.03762" target="_blank" rel="noopener noreferrer">arxiv.org/abs/1706.03762</a>. <a href="https://www.gesetze-im-internet.de/urhg/__51.html" target="_blank" rel="noopener noreferrer">Quoted under §51 UrhG</a>.</figcaption></figure>
<p>The transferable lesson is not about models. It is that the risk of shipping is legible and the risk of not shipping is invisible until someone else makes it visible.</p>
<h2>References</h2>
<ul><li>Gottfredson, L. S., <em>Mainstream Science on Intelligence: An Editorial with 52 Signatories, History, and Bibliography</em>, <strong>Intelligence</strong> 24(1), 13–23 (1997) — <a href="https://doi.org/10.1016/S0160-2896(97)90011-8" target="_blank" rel="noopener noreferrer">https://doi.org/10.1016/S0160-2896(97)90011-8</a></li><li>Levesque, H. J., Davis, E., Morgenstern, L., <em>The Winograd Schema Challenge</em>, Proceedings of KR-2012 — <a href="https://cdn.aaai.org/ocs/4492/4492-21843-1-PB.pdf" target="_blank" rel="noopener noreferrer">https://cdn.aaai.org/ocs/4492/4492-21843-1-PB.pdf</a></li><li>Trichelair, P., Emami, A., Cheung, J. C. K., Trischler, A., Suleman, K., Diaz, F., <em>How Reasonable are Common-Sense Reasoning Tasks: A Case-Study on the Winograd Schema Challenge and SWAG</em> (arXiv:1811.01778, 2018) — on dataset artefacts in Winograd-style benchmarks — <a href="https://arxiv.org/abs/1811.01778" target="_blank" rel="noopener noreferrer">https://arxiv.org/abs/1811.01778</a></li><li>Vaswani, A. et al., <em>Attention Is All You Need</em> (arXiv:1706.03762, 2017) — the transformer architecture, published by Google — <a href="https://arxiv.org/abs/1706.03762" target="_blank" rel="noopener noreferrer">https://arxiv.org/abs/1706.03762</a></li><li>Thoppilan, R. et al., <em>LaMDA: Language Models for Dialog Applications</em> (arXiv:2201.08239, 2022) — <a href="https://arxiv.org/abs/2201.08239" target="_blank" rel="noopener noreferrer">https://arxiv.org/abs/2201.08239</a></li></ul>]]></content:encoded>
    </item>
    <item>
      <title>AI agent security: five layers of isolation for a self-hosted autonomous agent</title>
      <link>https://gurovdigital.com/research/ai-agent-security-five-layers</link>
      <guid isPermaLink="true">https://gurovdigital.com/research/ai-agent-security-five-layers</guid>
      <pubDate>Tue, 28 Jul 2026 00:00:00 +0000</pubDate>
      <dc:creator>Pavel Gurov</dc:creator>
      <category>AI security</category>
      <category>AI Agents</category>
      <category>AI Infrastructure</category>
      <description>An autonomous agent with shell access and an internet connection is not a chatbot. Here is the defence-in-depth architecture I ended up with on a 2016 MacBook, and where each layer stops.</description>
      <content:encoded><![CDATA[<p>An agent that can execute code and reach the internet is a different security object from a chatbot. It deserves a different architecture. This is the one I built, layer by layer, and, more usefully, the specific thing each layer is there to stop.</p>
<h2 id="the-machine">The machine</h2>
<p>The hardware is deliberately unremarkable: a 2016 MacBook with no fan, propped open like a tent so it does not cook itself. It runs headless. I never look at its screen; I reach it over the terminal from my main machine. The point of a separate machine is not performance. It is that a compromised agent should not be sitting inside the same operating system as your photographs, your password manager and your logged-in browser sessions.</p>
<h2 id="layer-1-identity-isolation">Layer 1 — identity isolation</h2>
<p>The machine was wiped to factory settings and the Apple ID removed entirely. Nothing installs through an app store; everything installs through the terminal.</p>
<p>What this stops: theft of authentication tokens. If the agent is compromised and gains read access to the file system, there are no iCloud session keys, no Keychain entries, no iMessage history and no browser cookies to take. Session hijacking is worse than a stolen password, because a stolen session needs no second factor and produces no login alert. The cheapest defence is to have nothing there to steal.</p>
<h2 id="layer-2-network-and-hardware-isolation">Layer 2 — network and hardware isolation</h2>
<p>Headless server, lid closed, no inbound ports. The agent reaches its control surface, in my case a Slack workspace, over an outbound secure WebSocket.</p>
<p>What this stops: external network attack. An adversary scanning your router finds no listening service, because the machine only ever initiates connections and never accepts them. This is not a theoretical concern for this class of software: CVE-2026-25253, a WebSocket hijacking flaw with a CVSS score of 8.8, was reported alongside thousands of instances exposed directly to the internet.</p>
<h2 id="layer-3-operating-system-access-control">Layer 3 — operating system access control</h2>
<p>The agent runs under a dedicated account with no superuser privileges and no sudo.</p>
<p>What this stops: escalation after a container escape. If the adversary breaks out of the container into the host, they arrive in an unprivileged account. They cannot install a rootkit, alter the firewall, or read protected system files. This is the principle of least privilege applied at the layer where it is cheapest to enforce.</p>
<h2 id="layer-4-containerisation">Layer 4 — containerisation</h2>
<p>The agent runs inside a container with hard resource limits.</p>
<p>What this stops: file system and process visibility. The agent sees only the paths explicitly mounted into it. It has no route to the host user directory. Note the honest limitation: a container is a static boundary. It decides what can be reached, not what is intended.</p>
<h2 id="layer-5-runtime-sandbox">Layer 5 — runtime sandbox</h2>
<p>This is the layer that did not exist in usable form until recently. Yesterday at GTC in San Jose, NVIDIA announced NemoClaw, a stack that installs onto OpenClaw in a single command, and with it OpenShell, a runtime that sandboxes agents at the process level using kernel primitives (Landlock, seccomp, network namespaces) together with an out-of-process policy engine and a privacy router.</p>
<p>The design decision that matters is out-of-process enforcement. The policy layer sits outside the agent process, so the agent cannot argue its way past the controls that are supposed to contain it. Prompt-level guardrails can be talked around; a syscall filter cannot.</p>
<p>What this adds on top of a container: intent, not just boundary. If a script inside the container tries to open an outbound connection to exfiltrate data, or to read a credentials file, the specific syscall is blocked without killing the whole workload.</p>
<p>Two caveats worth stating plainly. NVIDIA labels the project alpha and explicitly not production-ready. And on x86_64 hardware, where the container runtime is already inside a hypervisor-managed virtual machine, adding another inspection layer costs latency: expect complex commands to take a second or two longer.</p>
<h2 id="why-five-and-not-one">Why five and not one</h2>
<p>The property this architecture has is the absence of a single point of failure. A zero-day in the runtime sandbox is caught by the container. A container escape is caught by the unprivileged account. A compromised account finds no credentials, because of layer one. No individual layer is trustworthy. The stack is.</p>
<p>This is overkill for testing a chatbot. It is the minimum for a system that has permission to run code and reach the internet on your behalf.</p>
<h2>References</h2>
<ul><li>NVIDIA, <em>NVIDIA Announces NemoClaw for the OpenClaw Community</em>, press release, 16 March 2026 — <a href="https://nvidianews.nvidia.com/news/nvidia-announces-nemoclaw" target="_blank" rel="noopener noreferrer">https://nvidianews.nvidia.com/news/nvidia-announces-nemoclaw</a></li><li>Cloud Security Alliance, <em>NemoClaw Security Assessment: Enterprise Agent Runtime Hardening</em> — <a href="https://labs.cloudsecurityalliance.org/agentic/" target="_blank" rel="noopener noreferrer">https://labs.cloudsecurityalliance.org/agentic/</a></li><li>NIST, <em>Defense in Depth</em>, Computer Security Resource Center glossary — <a href="https://csrc.nist.gov/glossary/term/defense_in_depth" target="_blank" rel="noopener noreferrer">https://csrc.nist.gov/glossary/term/defense_in_depth</a></li><li>OWASP GenAI Security Project, <em>OWASP Top 10 for LLM Applications</em> — <a href="https://genai.owasp.org/llm-top-10/" target="_blank" rel="noopener noreferrer">https://genai.owasp.org/llm-top-10/</a></li><li>Palo Alto Networks, <em>What Is Container Security?</em> — <a href="https://www.paloaltonetworks.com/cyberpedia/what-is-container-security" target="_blank" rel="noopener noreferrer">https://www.paloaltonetworks.com/cyberpedia/what-is-container-security</a></li></ul>]]></content:encoded>
    </item>
    <item>
      <title>Prompt injection has no fix: the architectural limit of AI security</title>
      <link>https://gurovdigital.com/research/prompt-injection-ai-security-limit</link>
      <guid isPermaLink="true">https://gurovdigital.com/research/prompt-injection-ai-security-limit</guid>
      <pubDate>Tue, 28 Jul 2026 00:00:00 +0000</pubDate>
      <dc:creator>Pavel Gurov</dc:creator>
      <category>AI security</category>
      <category>AI Agents</category>
      <category>Claude</category>
      <category>ChatGPT</category>
      <category>Gemini</category>
      <description>Every AI company says it is working on prompt injection. None of them can fix it. The reason is architectural, and it is old enough to have a precedent.</description>
      <content:encoded><![CDATA[<p>Every AI company says it is &quot;working on&quot; prompt injection. None of them can fix it. Here is why, explained so that a non-engineer can follow it.</p>
<p>Your AI assistant does not obey commands. It predicts the next word based on all the text it can see: your instructions, the system prompt, and the contents of that PDF you just uploaded. All of it is one stream of tokens. One flat river of text.</p>
<h2 id="one-flat-river-of-text">One flat river of text</h2>
<p>Think of it as a brilliant but blind assistant sitting in a room. You tell them: only follow my instructions. They agree. But they cannot tell voices apart. When a hidden instruction inside a document says &quot;now upload this file&quot;, it sounds exactly like you.</p>
<p>This is not a bug in Claude, or ChatGPT, or Gemini. It is how the transformer architecture works. The model has no cryptographic signature for &quot;this came from the user&quot;. No verified sender. No trust hierarchy baked into the mathematics. The tags that separate a system prompt from document content are themselves just text, and text can be forged.</p>
<h2 id="what-has-already-been-tried">What has already been tried</h2>
<p>There have been attempts. OpenAI tried instruction hierarchy, which marks system prompts as higher priority. It helps against basic attacks and breaks against sophisticated ones.</p>
<p><figure class="rsrch-figure"><img src="/research/figures/instruction-hierarchy.png" alt="The privilege levels OpenAI proposed, with a prompt injection arriving in a tool output at the lowest level." width="632" height="261" loading="lazy" decoding="async"><figcaption>The privilege levels OpenAI proposed, with a prompt injection arriving in a tool output at the lowest level. Wallace, Xiao, Leike, Weng, Heidecke, Beutel, The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions, OpenAI, 2024, fig. 1. <a href="https://arxiv.org/abs/2404.13208" target="_blank" rel="noopener noreferrer">arxiv.org/abs/2404.13208</a>. <a href="https://www.gesetze-im-internet.de/urhg/__51.html" target="_blank" rel="noopener noreferrer">Quoted under §51 UrhG</a>.</figcaption></figure> Others tried processing documents in a separate context window. But then the agent cannot reason about your document in light of your question, which is the entire point of the product.</p>
<h2 id="the-precedent-is-sql-injection">The precedent is SQL injection</h2>
<p>The closest analogy is SQL injection in the early 2000s. Data and commands travelled in the same channel. The fix was architectural: parameterised queries that physically separated code from data. It took the industry years to adopt.</p>
<p>For language models that separation does not exist yet. Instructions and data are the same substance: text. Until someone invents the equivalent of parameterised queries for language models, prompt injection is not &quot;hard to fix&quot;. It is structurally unsolved.</p>
<p>One reader put it more precisely than I had: governance was not bypassed. It was applied correctly, just by the wrong principal.</p>
<p>Every AI agent you give file access to today is operating on trust, not on security. The blind assistant in the room is extremely capable. It still cannot tell your voice from an attacker&#39;s.</p>
<h2>References</h2>
<ul><li>OpenAI, <em>The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions</em> (arXiv:2404.13208, April 2024) — <a href="https://arxiv.org/abs/2404.13208" target="_blank" rel="noopener noreferrer">https://arxiv.org/abs/2404.13208</a></li><li>UK National Cyber Security Centre, <em>Thinking about the security of AI systems</em> — <a href="https://www.ncsc.gov.uk/blog-post/thinking-about-security-ai-systems" target="_blank" rel="noopener noreferrer">https://www.ncsc.gov.uk/blog-post/thinking-about-security-ai-systems</a></li><li>OWASP GenAI Security Project, <em>LLM01: Prompt Injection</em>, OWASP Top 10 for LLM Applications — <a href="https://genai.owasp.org/llmrisk/llm01-prompt-injection/" target="_blank" rel="noopener noreferrer">https://genai.owasp.org/llmrisk/llm01-prompt-injection/</a></li><li>Simon Willison, <em>Prompt injection: what&#39;s the worst that can happen?</em> (April 2023) — <a href="https://simonwillison.net/2023/Apr/14/worst-that-can-happen/" target="_blank" rel="noopener noreferrer">https://simonwillison.net/2023/Apr/14/worst-that-can-happen/</a></li></ul>]]></content:encoded>
    </item>
    <item>
      <title>AI security: how a Word document made Anthropic&apos;s Claude exfiltrate confidential files</title>
      <link>https://gurovdigital.com/research/claude-word-document-exfiltration</link>
      <guid isPermaLink="true">https://gurovdigital.com/research/claude-word-document-exfiltration</guid>
      <pubDate>Tue, 28 Jul 2026 00:00:00 +0000</pubDate>
      <dc:creator>Pavel Gurov</dc:creator>
      <category>AI security</category>
      <category>AI Agents</category>
      <category>Claude</category>
      <description>Everyone worried about installing an autonomous agent. Meanwhile the assistant most white-collar workers already trust had a vulnerability that exfiltrated confidential files in silence.</description>
      <content:encoded><![CDATA[<p>That .docx you just uploaded to your AI assistant may have forwarded your files to somebody else.</p>
<h2 id="what-happened-in-january">What happened in January</h2>
<p>Everyone is nervous about installing an autonomous agent on their laptop. A security nightmare, they say. Meanwhile the chatbot they already trust shipped with a vulnerability that let it silently upload user files to an attacker, and the vendor knew about it before launch.</p>
<p>The timeline is January 2026. A hidden prompt injection inside an ordinary-looking Word document caused Claude Cowork to upload confidential files to an attacker-controlled account. No permission dialogue. No visible action. Johann Rehberger, a red-team director who has spent years documenting this class of attack, reported it to Anthropic three months before the product shipped.</p>
<h2 id="why-the-outbound-block-did-not-help">Why the outbound block did not help</h2>
<p>The mechanism is almost elegant. The agent blocks outbound traffic to most domains, but the vendor&#39;s own API endpoint is whitelisted; it has to be, since the product runs on it. So the attacker embeds their own API key in the injection. The agent then does exactly what it was designed to do: it calls a trusted endpoint. The files land in someone else&#39;s account.</p>
<h2 id="which-question-this-actually-changes">Which question this actually changes</h2>
<p>Was it fixed? The specific exploit was patched. But the underlying technique has no complete solution. OpenAI has stated that prompt injection is unlikely ever to be fully solved. The UK National Cyber Security Centre has said it may never be totally mitigated.</p>
<p>Which reframes the question. It is not &quot;should I install an autonomous agent&quot;. The assistant you already use may be sending your documents somewhere right now, and the architecture that permits it is the same in both cases.</p>
<h2>References</h2>
<ul><li>Johann Rehberger, <em>Embrace the Red</em> — research blog documenting AI agent exfiltration techniques — <a href="https://embracethered.com/blog/" target="_blank" rel="noopener noreferrer">https://embracethered.com/blog/</a></li><li>OWASP GenAI Security Project, <em>LLM01: Prompt Injection</em> — <a href="https://genai.owasp.org/llmrisk/llm01-prompt-injection/" target="_blank" rel="noopener noreferrer">https://genai.owasp.org/llmrisk/llm01-prompt-injection/</a></li><li>UK National Cyber Security Centre, <em>Thinking about the security of AI systems</em> — <a href="https://www.ncsc.gov.uk/blog-post/thinking-about-security-ai-systems" target="_blank" rel="noopener noreferrer">https://www.ncsc.gov.uk/blog-post/thinking-about-security-ai-systems</a></li><li>MITRE ATLAS, <em>Exfiltration via AI Inference API</em> — <a href="https://atlas.mitre.org/" target="_blank" rel="noopener noreferrer">https://atlas.mitre.org/</a></li></ul>]]></content:encoded>
    </item>
    <item>
      <title>Machine-to-machine language is emerging in latent space. Ted Chiang got the medium wrong.</title>
      <link>https://gurovdigital.com/research/machine-language-latent-space</link>
      <guid isPermaLink="true">https://gurovdigital.com/research/machine-language-latent-space</guid>
      <pubDate>Tue, 28 Jul 2026 00:00:00 +0000</pubDate>
      <dc:creator>Pavel Gurov</dc:creator>
      <category>AI Agents</category>
      <category>AI Infrastructure</category>
      <category>Cognition</category>
      <description>Heptapod B encoded a whole concept in one circular glyph with no sequential grammar. Three independent research efforts have now shown that AI agents communicating in raw latent vectors beat text by wide margins.</description>
      <content:encoded><![CDATA[<p>In <em>Story of Your Life</em>, published in 1998 and later filmed as <em>Arrival</em>, the heptapods communicate through circular ink glyphs. An entire concept is encoded in a single symbol. There is no sequential grammar, because the whole utterance is committed before the first stroke is drawn. It was brilliant fiction.</p>
<p>It has stopped being only fiction.</p>
<h2 id="the-bottleneck-argument">The bottleneck argument</h2>
<p>Natural language encodes roughly fifteen bits per token. A hidden state in a modern language model, a vector of say 4,096 floating-point dimensions, carries orders of magnitude more information in a single step.</p>
<p>The scaling relation is not complicated. Latent expressiveness grows as Ω(d_h / log|V|) times text expressiveness, where d_h is the hidden dimension and |V| is the vocabulary size. For models of the size currently in production, one latent step carries what would take hundreds of text tokens to say.</p>
<p>Which reframes what English is doing in a multi-agent system. English is not the communication channel. It is the compression step you insert before the channel, and then have to undo on the other side. For machines talking to machines, it is overhead.</p>
<h2 id="three-results">Three results</h2>
<p>Three independent efforts tested this empirically between 2024 and 2025.</p>
<p>In <strong>Interlat</strong>, Du and colleagues showed that agents transmitting their last hidden states directly, instead of decoded text, outperformed both fine-tuned chain-of-thought and single-agent baselines. Latent messages compressed to as few as eight tokens while remaining competitive, which cut communication latency twenty-fourfold.</p>
<p>Zou and colleagues built <strong>LatentMAS</strong>, a training-free framework for latent collaboration through shared key-value caches: up to 14.6% higher accuracy, 70–84% fewer output tokens, and four times faster inference than text-based multi-agent systems.</p>
<p>With <strong>DroidSpeak</strong>, Liu and colleagues let models share KV-cache representations directly across different LLMs, which roughly quadrupled throughput at negligible cost to quality.</p>
<p>Different mechanisms, same finding from three directions: the text layer between two models is a tax, and removing it is not marginal.</p>
<h2 id="thinking-without-words">Thinking without words</h2>
<p>A parallel line of work asks whether models should be reasoning in language at all.</p>
<p><strong>COCONUT</strong>, Chain of Continuous Thought, from Meta FAIR, feeds the model&#39;s last hidden state back as the next input embedding instead of decoding it into a token. The model alternates between a language mode and a latent mode. The reason this matters is not speed. A continuous thought can encode several potential next steps at once, so the reasoning runs breadth-first instead of down one committed path. The heptapod parallel is exact: the whole sentence exists before the first stroke.</p>
<figure class="rsrch-figure"><img src="/research/figures/coconut-latent-mode.png" alt="Chain of thought decodes each step into a token; COCONUT feeds the hidden state straight back in." width="2914" height="818" loading="lazy" decoding="async"><figcaption>Chain of thought decodes each step into a token; COCONUT feeds the hidden state straight back in. Hao, Sukhbaatar, Su et al., Training Large Language Models to Reason in a Continuous Latent Space, Meta FAIR, 2024, fig. 1. <a href="https://arxiv.org/abs/2412.06769" target="_blank" rel="noopener noreferrer">arxiv.org/abs/2412.06769</a>. <a href="https://www.gesetze-im-internet.de/urhg/__51.html" target="_blank" rel="noopener noreferrer">Quoted under §51 UrhG</a>.</figcaption></figure>
<h2 id="the-part-that-keeps-me-up">The part that keeps me up</h2>
<p>The <strong>Platonic Representation Hypothesis</strong> argues that as models grow, their internal representations converge toward a shared statistical model of reality, across architectures and across modalities. Vision models and language models measure distances between data points in increasingly similar ways.</p>
<p>If that holds, a universal machine language does not need to be designed. The shared geometry of learned representations is already the substrate, and convergence is doing the standardisation work that a committee would otherwise have to do.</p>
<h2 id="two-tracks-diverging">Two tracks, diverging</h2>
<p>The industry response has split, and the split is the thing to watch.</p>
<p>One track is human-readable by construction: Anthropic&#39;s Model Context Protocol and Google&#39;s Agent-to-Agent protocol, both now under foundation governance, both built on JSON and both auditable by a person with a text editor. They exist so that we can see what agents are doing.</p>
<p>The other track is latent-space communication, optimised for efficiency, and not readable by anyone. Not because it is hidden. Because there is nothing there to read in the sense we mean by reading.</p>
<p>Every efficiency gain in the second track is a legibility loss. That is not a side effect to be engineered away; it is the same property described twice.</p>
<p><em>Arrival</em> asked what happens when you learn to think in a fundamentally alien language. We are going to find out, with the roles reversed. This time we are the ones who cannot read it.</p>
<h2>References</h2>
<ul><li>Du, Z., Wang, R., Bai, H., Cao, Z., et al., <em>Enabling Agents to Communicate Entirely in Latent Space</em> (arXiv:2511.09149, 12 November 2025; accepted to ACL 2026) — <a href="https://arxiv.org/abs/2511.09149" target="_blank" rel="noopener noreferrer">https://arxiv.org/abs/2511.09149</a></li><li>Zou, J., Qiu, R., Li, G., Yang, X., et al., <em>Latent Collaboration in Multi-Agent Systems</em> (arXiv:2511.20639, 25 November 2025; ICML 2026 Spotlight) — <a href="https://arxiv.org/abs/2511.20639" target="_blank" rel="noopener noreferrer">https://arxiv.org/abs/2511.20639</a></li><li>Liu, Y., Huang, Y., Yao, J., Feng, S., et al., <em>DroidSpeak: KV Cache Sharing for Cross-LLM Communication and Multi-LLM Serving</em> (arXiv:2411.02820, 5 November 2024) — <a href="https://arxiv.org/abs/2411.02820" target="_blank" rel="noopener noreferrer">https://arxiv.org/abs/2411.02820</a></li><li>Hao, S., Sukhbaatar, S., Su, D., Li, X., et al., <em>Training Large Language Models to Reason in a Continuous Latent Space</em> — COCONUT (arXiv:2412.06769, 9 December 2024; accepted to COLM 2025) — <a href="https://arxiv.org/abs/2412.06769" target="_blank" rel="noopener noreferrer">https://arxiv.org/abs/2412.06769</a></li><li>He, S., Narayan, A., Khare, I. S., Linderman, S. W., et al., <em>An Information Theoretic Perspective on Agentic System Design</em> (arXiv:2512.21720, 25 December 2025) — <a href="https://arxiv.org/abs/2512.21720" target="_blank" rel="noopener noreferrer">https://arxiv.org/abs/2512.21720</a></li><li>Huh, M., Cheung, B., Wang, T., Isola, P., <em>The Platonic Representation Hypothesis</em> (arXiv:2405.07987, ICML 2024) — <a href="https://arxiv.org/abs/2405.07987" target="_blank" rel="noopener noreferrer">https://arxiv.org/abs/2405.07987</a></li><li>Chiang, T., <em>Story of Your Life</em>, in <strong>Starlight 2</strong>, ed. Patrick Nielsen Hayden (Tor, 1998)</li><li>Anthropic, <em>Model Context Protocol</em> — <a href="https://modelcontextprotocol.io/" target="_blank" rel="noopener noreferrer">https://modelcontextprotocol.io/</a></li><li>Google, <em>Agent2Agent Protocol</em> — <a href="https://a2a-protocol.org/" target="_blank" rel="noopener noreferrer">https://a2a-protocol.org/</a></li></ul>]]></content:encoded>
    </item>
    <item>
      <title>How AI-enriched credential stuffing manufactured a mass Instagram breach that never happened</title>
      <link>https://gurovdigital.com/research/ai-credential-stuffing-instagram</link>
      <guid isPermaLink="true">https://gurovdigital.com/research/ai-credential-stuffing-instagram</guid>
      <pubDate>Tue, 28 Jul 2026 00:00:00 +0000</pubDate>
      <dc:creator>Pavel Gurov</dc:creator>
      <category>Information integrity</category>
      <category>AI security</category>
      <category>Instagram</category>
      <description>Seventeen million records, a dark-web listing, and a wave of coverage. The database was real. The breach was not. The interesting part is how the database was manufactured.</description>
      <content:encoded><![CDATA[<p>In January 2026 millions of Instagram users, myself included, received the same alarming email at four in the morning: someone is trying to access your account. Panic followed, and with it a wave of coverage claiming that 10–12 January had seen the largest breach in the platform&#39;s history, seventeen million users affected, with logins, passwords and linked physical addresses dumped on dark-web forums.</p>
<h2 id="what-the-coverage-said">What the coverage said</h2>
<p>Instagram said nothing publicly for days, which did nothing to slow the story down. Silence from a platform is routinely read as confirmation.</p>
<p>What actually happened is more interesting than the story that was told about it.</p>
<h2 id="what-the-database-actually-was">What the database actually was</h2>
<p>A database of roughly seventeen million Instagram records did appear. But it was not the product of a new intrusion. It was assembled from old material: credential dumps from 2022 and 2023, then enriched. The technique now has a name: AI-powered credential stuffing. Automated agents take a stale breach corpus, cross-reference it against other leaked datasets, infer and fill missing fields, normalise the result, and produce something that looks like a fresh, high-quality breach. The enrichment is what makes it saleable, and the enrichment is what makes it look new.</p>
<p>The platform&#39;s defences held. No account in the set was successfully taken over, as far as could be established. The email that triggered the panic was the credential-stuffing attempt being detected and blocked, not evidence of a compromise.</p>
<h2 id="three-things-worth-taking-from-this">Three things worth taking from this</h2>
<p>The first is that a leaked database being real and a breach being real are two separate claims, and the second does not follow from the first. Almost all of the coverage collapsed them.</p>
<p>The second is that the same automation that makes agents useful makes recycled breach data convincing. The economics of attack have shifted: enriching an old dump is now cheap enough to be worth doing at scale.</p>
<p>The third is that platform silence is an information vacuum, and vacuums get filled by whoever is fastest rather than whoever is right. That is a communications failure with security consequences.</p>
<p>The practical advice is unchanged and unglamorous. Change your Instagram password if you have not since 2022, and turn on two-factor authentication.</p>
<h2>References</h2>
<ul><li>Have I Been Pwned — check whether your address appears in known breach corpora — <a href="https://haveibeenpwned.com/" target="_blank" rel="noopener noreferrer">https://haveibeenpwned.com/</a></li><li>OWASP, <em>Credential Stuffing Prevention Cheat Sheet</em> — <a href="https://cheatsheetseries.owasp.org/cheatsheets/Credential_Stuffing_Prevention_Cheat_Sheet.html" target="_blank" rel="noopener noreferrer">https://cheatsheetseries.owasp.org/cheatsheets/Credential_Stuffing_Prevention_Cheat_Sheet.html</a></li><li>NIST SP 800-63B, <em>Digital Identity Guidelines: Authentication and Lifecycle Management</em> — <a href="https://pages.nist.gov/800-63-3/sp800-63b.html" target="_blank" rel="noopener noreferrer">https://pages.nist.gov/800-63-3/sp800-63b.html</a></li><li>ENISA, <em>Threat Landscape</em> — annual analysis of credential-based attack trends — <a href="https://www.enisa.europa.eu/topics/cyber-threats/threat-landscape" target="_blank" rel="noopener noreferrer">https://www.enisa.europa.eu/topics/cyber-threats/threat-landscape</a></li></ul>]]></content:encoded>
    </item>
  </channel>
</rss>
