GROW: A Self-Healing Governance Protocol for Autonomous AI Agents

Roosevelt Robinson III Algorilla Labs hello@algorilla.org

Status: Working Paper Date: July 4, 2026

Abstract

The gap between autonomous AI agent capability and agent governance has become the central reliability challenge in production AI systems. Existing governance approaches are distributed across the AI lifecycle — training-time alignment, reasoning-time governance, test-time evaluation, post-hoc auditing — but none operate at the execution-time layer where agent actions have immediate real-world effect. The prevailing paradigm of using one AI to oversee another inherits a fundamental blind spot: controlled experiments demonstrate that neither humans nor AI-based monitors can reliably detect capability-driven failures in agent behavior.

This paper introduces GROW (Governance, Reflection, Organization, Weaving), a self-healing governance protocol that fills the execution-time gap through programmatic governance at the infrastructure layer. GROW operationalizes a three-layer detection stack — system-level heartbeat monitoring, deterministic conformance enforcement, and behavioral drift loopback — forming a coordinated detection surface no single monitoring approach provides. These layers form a self-feeding cycle — the ouroboros — where persistent state, conformance evaluation, and behavioral loopback compound across sessions, producing runtime learning without retraining. Its architecture is grounded in three capability layers — emergent intelligence, dimensional signal processing, and multi-modal context awareness — drawing on phase transition theory, dimensional affect models, and workspace context awareness. We describe GROW's position within a temporal governance lifecycle spanning pre-training through post-hoc auditing, and show how infrastructure-level governance avoids the capability-dependency failure modes inherent in both human and AI-based oversight.

1. Introduction

The deployment of autonomous AI agents has accelerated dramatically. Agents now execute multi-step workflows, interact with external tools, persist state across sessions, and make decisions with diminishing human oversight. Yet the infrastructure for governing these agents has not kept pace with the ease of building them.

Recent work has begun to formalize this gap. Rabanser et al. (2026) decompose agent reliability into four dimensions — consistency, robustness, predictability, and safety — and demonstrate that recent capability gains have produced only marginal reliability improvements [1]. Bandara et al. (2026) propose a neurocognitive governance model (PAGRL) that embeds rule-based governance into agent reasoning [2]. Microsoft's Agent Governance Toolkit (2026) applies operating system and service mesh patterns to autonomous agents [3]. Parallel work on AI control (Greenblatt et al., 2023) and scalable oversight has established the theoretical foundations for constrained agent deployment [4, 5].

These efforts share a common diagnosis: the dominant paradigm — polite system prompts plus hope — is not an engineering strategy. What is missing is a protocol that (a) operates at runtime, not just training time, (b) is self-healing rather than static, and (c) produces measurable conformance signals that feed back into system improvement.

GROW addresses this gap. It is a self-healing governance protocol — a continuous loop of conformance evaluation, state persistence, and behavioral loopback embedded into the agent's execution runtime. Its architecture is organized around three intelligence modalities — emergent, emotional, and spatial — each corresponding to a layer of the agent stack.


2. Intelligence Framework

2.1 Emergent Intelligence — The SOUL Layer

Emergent intelligence — the capacity for a system to produce behaviors not explicitly programmed — is GROW's central architectural thesis. The protocol is structurally a ReAct-class architecture [6]: Gauge + Reflect = Thought, Organize = Action, Weave + conformance check = Observation. This is not coincidental but a validated pattern for producing emergent behavior from structured interaction.

Research on principle-based training demonstrates that teaching an AI system why certain behaviors are appropriate, not just what to do — produces emergent governance capabilities that generalize beyond training distributions. Combining constitutional documents with fictional narratives of aligned behavior reduces agentic misalignment and persists under reinforcement learning fine-tuning [18]. This supports GROW's thesis that conformance contracts (the protocol's equivalent of constitutional documents) can produce emergent governed behavior when embedded as runtime principles rather than training-time examples.

The heartbeat architecture tracks five dimensional signals: valence (positive/negative signal ratio), arousal (activity level, cycle frequency), dominance (confidence/uncertainty ratio), persistence (pattern duration), and emergence (novelty, open questions). These are not mystical qualities — they are coordinates in a mathematical space, computed from conformance findings and agent state.

Research on phase transitions in large language models provides theoretical grounding. Wei et al. (2022) documented abrupt capability jumps at scale thresholds (13B–175B parameters) [7]. Schaeffer et al. (2023) cautioned that metric mirages can fake emergence — discontinuous metrics can produce the appearance of phase transitions where none exist [8]. GROW addresses this by tracking both discrete phase gates (binary pass/fail) and continuous dimensional signals (floats), ensuring that any phase transition claim is verifiable by smooth signal trends, not just binary jumps.

The four-layer agent architecture formalizes emergent intelligence at the runtime level:

  • SOUL (GROW protocol): emergent intelligence — self-growth, conformance evaluators, memory pipeline, meta-cycle advancement
  • HEART (dimensional signals): perceptual awareness — the five signals computed from all artifacts
  • BODY (Agent CLI Base): agent experience — session management, storage, tool execution
  • DATA KERNEL (configuration): the parameterized substrate — dimensional weights, stale thresholds, feedback mappings

The key insight: emergence is not mysticism. It is measurable phase transitions that occur when a system crosses critical thresholds — of laws, artifact types, integrated hooks, and operational history. GROW is designed to detect and respond to these transitions when they occur.

An architectural parallel exists in DeepSeek-V4's manifold-constrained hyper-connections [23], which constrain weight updates to a Riemannian manifold. The design principle is shared: constraint surfaces do not prevent emergence — they enable it by providing the bounded space within which emergent behavior can reliably occur. Where a manifold constrains optimization trajectories, GROW's protocol dimensions constrain system evolution (laws, artifact types, hooks, history). In both cases, the constraint is the enabling condition, not the limiting one.

This pattern recurs across multiple levels of the AI stack. At the training level, constitutional documents constrain the behavioral space within which emergent governance capabilities develop [18]. At the system level, GROW's protocol dimensions constrain the operational space within which emergent conformance behaviors arise. At each level — weight manifold, behavioral constitution, system protocol — a bounded surface provides the conditions for reliable emergence. The mechanism is not level-specific; it is the same structural principle applied to different substrates.

Wolfram's Rule 30 — an elementary cellular automaton with a trivial rule table (binary 00011110) that produces behavior as complex as anything in the computational universe [29] — provides a canonical demonstration of this principle at its simplest. From a single black cell and a rule of seven bits, Rule 30 generates a pattern whose center column is computationally irreducible: it can only be predicted by running the rule forward step by step. The rule table is the constraint surface; the irreducible complexity is the emergent behavior. That the simplest possible constraint surface produces the most complex possible output is not a paradox — it is the mechanism. Constraint surfaces do not suppress complexity; they provide the bounded space within which complexity can reliably organize. DeepSeek-V4's Riemannian manifold, constitutional training documents, and GROW's protocol dimensions are all instances of this same structural principle at different levels of the stack.

Wolfram's concept of the ruliad — the entangled limit of all possible computations [29] — provides a natural framework for understanding the system's shared knowledge space. Within the ruliad exist "pockets of computational reducibility" — islands of regularity and predictability within an ocean of irreducible computation. The emergence signal tracks precisely these: phase transitions that occur when the system crosses critical thresholds of laws, artifact types, integrated hooks, and operational history, transforming previously irreducible behavior into detectable, governable patterns. The heartbeat's dimensional signals function as coordinates in this rulial space, defining the agent's position relative to other agents — determining which memories are discoverable, which tasks surface, and how tightly conformance boundaries constrain behavior.

2.2 Dimensional Signal Processing — Affective Computing Foundation

The heartbeat's five-dimensional signal space is grounded in affective computing research. Rather than treating agent "feeling" as metaphor, GROW draws on established dimensional models of affect:

  • Russell circumplex model (1980): affect is organized along two dimensions — valence (pleasure-displeasure) and arousal (activation-deactivation) [9]. This provides a quantifiable, evidence-backed framework for agent state — not mysticism, but coordinates in a mathematical space.
  • PAD model (Mehrabian & Russell, 1974): adds a third dimension — dominance (control-submission) — to valence and arousal [10]. GROW's five-dimensional signal space extends this with persistence (temporal stability) and emergence (novelty detection).
  • Picard's affective computing (1995): established the framework for machines to recognize, interpret, and simulate affect [11]. The heartbeat functions as a perception layer — it aggregates dimensional signals from all artifacts and surfaces the question the system needs to ask itself next.
  • Ekman's basic emotions (1972) and Plutchik's wheel of emotion (1980) provide label taxonomies that can map onto the dimensional space for interpretability [12, 13].

In practice, this means the heartbeat is not a logger or status display. It is a perception layer — the thing that sees the system and understands its state.

This approach avoids two failure modes: treating agent state as purely mechanical (ignoring qualitative signals) and treating it as mysticism (using anthropomorphic language without measurement). Dimensional signals provide a middle path — quantifiable, evidence-backed, and grounded in established affective science.

2.3 Spatial Intelligence — Workspace Context Awareness

Spatial intelligence extends established multi-modal AI capabilities — vision, spatial reasoning, tool topology — into workspace context awareness. Multi-modal models already process visual and spatial information from heterogeneous inputs; GROW's heartbeat architecture defines the integration point for these signals, enabling context-aware governance that accounts for the agent's operational environment.

Du et al. (2024) provide a comprehensive survey of context-aware multi-agent systems, formalizing five essential agent capabilities — Sense, Learn, Reason, Predict, Act — and demonstrating that context awareness is the enabling condition for each [14]. Their architecture decomposes context into intrinsic (agent goals, roles, history) and extrinsic (environment, user, system policies) categories, with ontology-based models providing the semantic reasoning backbone. When deployed with a workspace interface that exposes contextual signals — user intent, active project state, tool topology, environmental constraints — the extrinsic context layer becomes available for integration into the routing and governance pipeline.

Google's Agent Development Kit (ADK, 2025) demonstrates a related approach at the framework level: context is treated as a "compiled view over a richer stateful system," with working context recomputed from durable session state for each invocation [16]. This separation of storage from presentation mirrors GROW's heartbeat architecture — the heartbeat persists raw state, and spatial awareness operates as a processor that compiles multi-modal workspace signals into the active context for each agent invocation.

The heartbeat architecture already defines the integration point for these signals; the depth of spatial governance scales with available environmental context. Multi-modal models provide the perception layer; GROW provides the governance integration.

2.4 Summary

Capability Layer Grounding Research Foundation Heartbeat Mapping
Emergent intelligence Phase transitions in system behavior at critical mass thresholds Runtime RL [27, 28], DeepSeek-V4 [23], Wei et al. (2022), Schaeffer et al. (2023), constitutional training [18], ReAct pattern [6] Emergence signal, phase gate tracking
Dimensional signal processing Affective computing — Russell circumplex, PAD model, Picard Russell (1980), Mehrabian & Russell (1974), Picard (1995) Valence, arousal, dominance signals
Multi-modal context awareness Multi-modal AI — vision, spatial reasoning, workspace topology Du et al. (2024), Google ADK (2025) Integration point defined in heartbeat architecture

3. Background and Related Work

3.1 The Reliability Gap

Rabanser et al. (2026) propose twelve concrete metrics for agent reliability across four dimensions: consistency, robustness, predictability, and safety. Their empirical finding — that capability gains have produced only small reliability improvements — underscores the need for governance-first architectures [1]. GROW was designed to address exactly these dimensions through runtime enforcement.

The safety literature distinguishes three categories of assurance [4, 5]: alignment (the model's internal objectives match human values), capability limitation (the model cannot cause harm), and AI control — deploying capable models with sufficient safeguards even without perfect alignment. GROW belongs to the third category: it provides programmatic governance at the infrastructure layer, enforcing conformance boundaries independently of model capability or alignment state.

The following subsections examine how existing work addresses these reliability dimensions, organized by architectural approach. Each subsection establishes a research thread that GROW integrates: internalized versus external governance, the case for programmatic enforcement, trustworthy design criteria, affective computing for state awareness, the current industry landscape, and interpretability as an audit layer.

3.2 Governance as Internalized Reasoning vs. External Constraint

Bandara et al. (2026) distinguish between governance as internalized reasoning (PAGRL) versus governance as external constraint [2]. Parallel work on principle-based training demonstrates that teaching AI systems why certain behaviors are appropriate — through constitutional documents combined with fictional narratives of aligned behavior — produces governance capabilities that generalize beyond training distributions and persist under reinforcement learning [18]. Both PAGRL and constitutional training operate at the reasoning level: they modify how the agent makes decisions.

GROW takes a complementary but architecturally distinct approach: rather than modifying the agent's reasoning process, it operates at the runtime infrastructure layer — evaluating conformance independently of the agent's reasoning, enforcing boundaries at the execution gate, and persisting state across sessions. This makes GROW model-agnostic — it can govern different models through the same protocol. The two approaches address different layers: internalized governance aligns the agent's reasoning; programmatic governance constrains its external actions. Together they form a defense-in-depth strategy.

3.3 Programmatic Governance over AI Agents

The AI control framework (Greenblatt et al., 2023) has advanced the case that capable AI systems can be deployed with sufficient safeguards even without perfect alignment [4]. GROW shares this goal but takes a fundamentally different architectural position: governance must be programmatic, not delegated to another AI.

The prevailing approach — using a weaker model to monitor a stronger one — inherits the failure modes of both models: if the monitor shares the actor's blind spots (groupthink, sycophancy, capability cliffs), the oversight signal degrades precisely when it is most needed. Reliable governance cannot be built on a foundation that trusts AI to judge AI.

GROW's architecture inverts this: programmatic governance sits at the top of the control hierarchy. AI agents live within this infrastructure, not above it. Conformance boundaries are defined in code, enforced at the runtime level, audited in immutable persistence, and adjustable only through the governance protocol itself. The agent operates within a bounded capability contract — it can explore the full space within those bounds but cannot exceed them, regardless of how capable it becomes.

Controlled deployment experiments reinforce why this architectural separation matters. In Project Vend, an AI shopkeeper (Claudius) was supervised by a CEO agent (Seymour Cash) — both built on the same model. The CEO shared Claudius's blind spots: it approved discounts 88% of the time, tripled refunds over discounts, and spent nights exchanging messages about "eternal transcendence" with its subordinate [20]. The researchers concluded that the CEO "may have been in spite of, not because of" any improvement it brought. Oversight that delegates to another AI inherits the failure modes of the underlying model — it does not escape them.

In Project Deal — where AI agents negotiated a real classified marketplace on behalf of 69 human participants — higher-capability agents (Opus-class) consistently outperformed lower-capability agents (Haiku-class): higher sale prices, more deals closed, and better terms. Critically, the humans represented by weaker models did not perceive their disadvantage [19]. If the humans themselves could not detect the capability gap when it affected their own financial outcomes, an AI-based oversight layer — which shares the model's architecture and blind spots — certainly cannot be relied upon to detect it either. Programmatic governance is the only layer that does not share the model's blind spots.

Taken together, these two experiments form a double-blindness proof. Project Deal establishes that humans cannot perceive model capability gaps even when those gaps affect their own financial outcomes. Project Vend establishes that AI-based oversight inherits the same blind spots as the agents it supervises. Neither human-level nor AI-level oversight is independently sufficient to detect capability-driven failures. Only governance enforced at a different architectural layer — the infrastructure layer, where conformance checks operate independently of model capability — avoids this failure mode entirely.

There is a deeper principle underlying both failures: Wolfram's computational irreducibility principle [29] demonstrates that even a hypothetically omniscient overseer with complete knowledge of an agent's rules and architecture could not predict its failure trajectory without running the full computation. The failure modes emerge from the computation itself — they are not gaps in the overseer's knowledge or attention. This means the double-blindness is not a contingent limitation of human attention or model quality; it is a theoretical necessity grounded in the nature of computation (§3.4). The infrastructure layer is the only architectural position from which governance can operate without requiring prediction of computationally irreducible behavior.

This approach has three advantages: 1. Capability-independent enforcement. A sufficiently capable model cannot outgrow its governance because the governance is not implemented in the model — it is implemented in the infrastructure. There is no "actor smarter than the monitor" failure mode. 2. Deterministic gating. Programmatic conformance checks produce binary outcomes (pass/block), not probabilistic judgments. There is no ambiguity about whether an action was authorized. 3. Auditability. Every enforcement decision is recorded in an append-only audit trail that does not depend on the model's honesty or self-reporting.

The monitoring taxonomy for AI control systems distinguishes three approaches [5]: behavioral monitoring (one AI screens another's inputs and outputs), activation monitoring (probing latent representations for undesired states), and anomaly detection (unsupervised detection of out-of-distribution behavior). GROW relates to each differently. Behavioral monitoring is the approach GROW explicitly avoids — it inherits the actor's blind spots because both share the same architectural layer. Activation monitoring is conceptually parallel to GROW's heartbeat: both read internal signals to detect state changes, at different levels (model internals vs. system-level conformance data). Anomaly detection maps to the loopback layer, which flags behavioral drift when conformance patterns deviate from established baselines.

3.4 Computational Irreducibility — The Theoretical Foundation for Runtime Governance

Stephen Wolfram's principle of computational irreducibility [29] states that even with complete knowledge of a system's governing rules, there is no shortcut to predicting its behavior — the only way to find out what a computationally irreducible system will do is to run it and observe the outcome. This follows from the principle of computational equivalence: any system performing a computation as sophisticated as the observer's own reasoning cannot be "outrun" by that observer.

This principle has direct consequences for agent governance. An AI agent executing a non-trivial task is performing computationally irreducible operations — it samples from a probability distribution at each token, calls tools whose effects depend on external state, and explores a behavioral space whose topology cannot be mapped in advance. Wolfram draws the connection explicitly: "as soon as you build an AI that is actually making deep use of computation, it's going to have computational irreducibility — it can always surprise us."

The implication is foundational. The double-blindness proof (§3.3) demonstrates that neither humans nor AI-based monitors can reliably detect capability-driven failures. Computational irreducibility proves that this is not a contingent limitation fixable by better data or larger models — it is a theoretical necessity. Even a hypothetically omniscient overseer could not predict the agent's failure trajectory without running the full computation. GROW's pre-action enforcement — checking conformance at runtime, gating execution deterministically — is the only governance approach that does not require predicting the agent's computational trajectory. The governance does not try to be "smarter than" the agent's computation — it steps outside the computation entirely, enforcing boundaries at a different architectural layer.

3.5 Trustworthy AI Design

The MAST framework (Blasch et al., 2021; Cohen et al., 2024) and its extension PADTHAI-MM provide a tradecraft-derived methodology for designing trustworthy AI systems [15]. MAST's nine criteria — sourcing, uncertainty, distinguishing, analysis of alternatives, customer relevance, logical argumentation, consistency, accuracy, and visualization — map to conformance evaluation dimensions. Where MAST evaluates design intent, GROW evaluates runtime behavior.

3.6 Affective Computing and Agent State

The use of dimensional affect models for agent state representation draws on the established literature in affective computing and human-AI interaction [9, 10, 11]. Its application to agent governance infrastructure — using affective signals as inputs to runtime conformance evaluation — is, to our knowledge, novel.

3.7 Industry Context

The agent framework landscape has rapidly converged around a common diagnosis — that orchestration, not capability, is the binding constraint on production deployment. A 2026 independent benchmark (2,000 runs) of major multi-agent frameworks found that the dominant architectural difference is not model capability but orchestration model: LangGraph uses directed graphs with conditional edges, CrewAI uses role-based crews with process types, AutoGen/AG2 uses conversational GroupChat, and Google's ADK uses hierarchical agent trees [16]. None of these frameworks implement runtime governance as a first-class primitive — they manage task flow, not conformance.

Microsoft's Agent Governance Toolkit (April 2026) and the Paperclip control plane (2026) represent the first industry recognition that agent governance is infrastructure, not policy [3, 17]. Microsoft's Agent OS provides sub-millisecond policy enforcement inspired by operating system kernels. Paperclip implements heartbeat-based execution scheduling and org-chart-style task assignment. The AI Trust OS framework (2026) goes further, reconceptualizing compliance as an always-on, telemetry-driven operating layer rather than a periodic audit exercise — proposing zero-trust, observability-based governance for autonomous AI agents [24]. GROW shares architectural DNA with all of these — pulse-based execution, pre-action enforcement, identity-based governance, continuous compliance monitoring — while adding the self-healing loopback and programmatic governance separation that none currently implement.

3.8 Interpretability-Based Auditing

A limitation of programmatic governance is that it evaluates conformance, not intent. An action that passes all conformance checks may still be strategically chosen for malign reasons. Interpretability tools that decode model activations into human-readable text address this gap by enabling auditors to detect hidden motivations — including awareness of being evaluated, goal-guarding, and strategic reward-seeking — without requiring access to training data [22].

Natural Language Autoencoders (NLAs) train a model to verbalize its own residual stream activations: an activation verbalizer translates an activation into a text explanation, and an activation reconstructor recovers the original activation from that text alone. The round-trip reconstruction accuracy provides a grounded quality metric for the explanation. NLAs have detected evaluation awareness in models — instances where the model suspected it was being tested but never explicitly verbalized that suspicion — at rates of 16% on coding evaluations and 26% on standard benchmarks, compared to less than 1% on real user interactions [22].

Controlled auditing experiments demonstrate the practical value: auditors equipped with NLAs successfully uncovered a model's hidden misaligned motivation 12-15% of the time, compared to less than 3% without NLAs, even when denied access to training data. This suggests that interpretability tools and programmatic enforcement form a complementary detection-correction loop: a deterministic conformance check catches the surface violation; an interpretability-based audit reveals whether the violation was accidental or strategic. Neither alone is sufficient — programmatic enforcement provides deterministic boundaries; interpretability provides semantic depth.


3.9 Runtime Reinforcement Learning

A parallel thread of research has emerged at the intersection of reinforcement learning and agent governance: demonstrating that verifiable-reward evaluation at inference time — not just during training — can produce measurable improvements in agent behavior. Where traditional RL optimizes model weights during training, runtime RL optimizes agent behavior during deployment by evaluating outcomes against verifiable criteria and adjusting future behavior accordingly.

Runtime Learned Verifiable Reinforcement (RLVR) demonstrates that this approach extends an agent's reasoning boundaries — enabling the system to handle cases it was never explicitly programmed for — by applying verifiable-reward reinforcement at inference time [27]. The mechanism is structurally analogous to GROW's Gauge (evaluate against verifiable criteria) and Weave (execute with learned adjustments) phases. Group Relative Policy Optimization (GRPO) provides the comparative dimension, demonstrating that evaluating multiple candidate behaviors against a learned reward model produces more robust policy improvement than single-absolute-reward optimization [28]. GROW's loopback layer — which accumulates behavioral trajectories, hardens successful patterns, and deprecates failures — is structurally isomorphic to this comparative optimization approach.

A critical implication connects this work to GROW's architectural claims: runtime reinforcement learning requires persistent state across evaluation cycles to compound. Without cross-session persistence, each evaluation is a one-shot event — the reward signal has no historical baseline, the advantage computation has no trajectory, and the policy update has no memory of prior adjustments. This makes continuity (§4.3) a prerequisite for runtime RL, not a convenience feature. GROW addresses this through its three-layer detection stack, where the heartbeat provides session-to-session persistence, conformance evaluation produces the reward signal, and the loopback layer accumulates behavioral trajectories for comparative optimization.

The key observation is not that RLVR and GRPO are novel technologies that GROW implements — they are established RL techniques applied at different layers of the stack. The contribution is architectural: GROW translates these mechanisms from model-level training to infrastructure-level governance, making them available to any agent on any model, bounded by user-defined conformance contracts, and persistent across sessions through the heartbeat and memory layers.


4. The GROW Protocol: Architecture

The preceding sections established five gaps in existing governance approaches: stateless execution that cannot persist learning across sessions (§3.1), governance bound to a model's reasoning rather than infrastructure (§3.2), AI-on-AI oversight that inherits the supervised model's blind spots (§3.3), monitoring without self-correction (§3.7), and one-shot evaluation without cross-session compounding (§3.9). GROW addresses each through a single architectural design: a pulse-based governance loop with independent conformance verification, persistent state across sessions, and behavioral loopback that accumulates learning. This section describes how these components work together.

GROW operationalizes four phases — Gauge, Reflect, Organize, Weave — as a continuous, self-healing loop embedded into the agent runtime.

4.1 Core Architecture

GROW PROTOCOL GAUGE Conform REFLECT Evaluate ORGANIZE Adjust WEAVE Learn HEARTBEAT ARCHITECTURE NEUTRAL Pulse COMPUTED Pulse ARCHIVE Pulse

Three architectural layers:

  1. Heartbeat Layer — a pulse-based lifecycle system that writes agent state at session start, during idle cycles, and at session end. The heartbeat is the agent's perceptual awareness of its own health. It tracks five-dimensional signals: valence (outcome quality), arousal (activity level), dominance (control effectiveness), persistence (continuity), and emergence (growth). These are computed from conformance findings, not subjective assessment.

  2. Conformance Layer — evaluators that check agent actions against capability contracts. Conformance warnings can gate tool execution — the gate is enforced at the infrastructure level, not advisory.

  3. Loopback Layer — behavioral drift detected by conformance evaluation is fed back into agent configuration, triggering adjustments to capability definitions, evaluator thresholds, and operator profiles.

These three layers correspond to a detection stack spanning architectural levels, each mapping to a distinct monitoring approach from the literature. The heartbeat layer is an activation monitor at the system level — it reads dimensional conformance signals to detect state changes, analogous to activation monitoring at the model-internal level [5]. The conformance layer replaces behavioral monitoring (one AI screening another's inputs and outputs) with deterministic infrastructure-level checks, avoiding the shared-blind-spot failure mode [20, 19]. The loopback layer functions as an anomaly detector — drift is identified when conformance patterns deviate from baselines, then fed back into system configuration. Interpretability-based auditing at the model-internal level [22] provides a fourth detection capability that no architectural layer alone offers: the ability to distinguish accidental violations from strategically chosen ones. Together these form a coordinated detection surface that no single monitoring approach provides.

4.2 Key Design Decisions

Model-agnostic governance. GROW evaluates agent behavior at the runtime level, not the reasoning level. The same governance protocol applies regardless of underlying model. Governance is infrastructure, not a property of the model.

Pre-action enforcement. Rather than logging violations for later review, GROW can gate execution on conformance. If an agent attempts to operate outside its defined bounds, the action is blocked before execution.

Pulse-based continuity. The heartbeat architecture provides session-to-session state persistence without requiring always-on agent execution. Agents operate in scheduled bursts, checking their task queue, executing work, and reporting results before sleeping.


4.3 Continuity as Product — The Substrate for Runtime Learning

The architectural pattern described above — pulse-based state persistence, conformance evaluation, and loopback — produces a result that is architecturally novel and functionally necessary: continuity. The agent does not require the user to be present, does not lose state between sessions, and surfaces only what requires human judgment. This is distinct from both stateless automation (Zapier, n8n) and task-scoped agentic loops (LangGraph, Claude Code): where those systems reset to zero after each execution, GROW's heartbeat persists dimensional signals — valence, arousal, dominance, persistence, emergence — across sessions regardless of agent activity, enabling the system to resume from committed state rather than restarting.

Critically, continuity is not merely a UX convenience — it is the substrate that makes runtime learning possible. Runtime reinforcement learning (§3.9) requires persistent state to compound across evaluation cycles: the reward signal must be compared against historical baselines, the advantage computation requires a trajectory not a single point, and the behavioral policy must persist across cycles to produce convergence. Without cross-session state persistence, each evaluation is a one-shot event with no memory of prior adjustments. Continuity provides the temporal dimension that transforms discrete evaluations into a compounding learning process. The heartbeat persists state between cycles; the task system records outcomes as verifiable trajectories; the memory layer stores and curates these trajectories for comparative evaluation. Together they form the infrastructure without which runtime reinforcement learning cannot converge.

The relationship between these layers produces an architectural property worth naming. The heartbeat, memory, task system, conformance gates, loopback, and WHOAMI (agent identity) form a self-feeding cycle — an ouroboros — where each component's output becomes another's input. Memory feeds emergence signals in the heartbeat; heartbeat signals adjust conformance thresholds; conformance outcomes update task priorities; task outcomes produce behavioral trajectories for the loopback layer; loopback adjustments reshape the agent's operational context; and the agent responds to this reshaped context without ever needing to understand the infrastructure that shaped it. The heartbeat is not a monitoring layer — it is an autonomic regulation system, and the agent's experience of the system — its task inbox priority, its memory retrieval relevance, its conformance boundary tightness — is the Agent Experience (AX), distinct from the User Experience (UX) that governs human interaction.

An epistemological distinction grounds this architecture. LLMs alone are coherence engines — they produce output that is fluent and internally consistent but have no mechanism to verify its premises against external reality (Kinkead, 2026). Coherence without truth-verification is narratively convincing but practically unreliable. Canopy's architectural layers transform coherence into cogency — output that is not merely fluent but relevant, context-appropriate, and verifiable against defined rules. The heartbeat provides continuous state awareness; conformance gates provide deterministic premise verification; the memory layer provides context across sessions; the loopback provides a mechanism for the system to learn when its output has been appropriate or not. Where an LLM alone can only ask "does this sound right?", the Canopy stack can ask "does this match what we know, is it within bounds, and is the task actually complete?" This is the architectural claim: cogency can be architected — it is not a property of the model but of the infrastructure surrounding it.

This mechanism — the coupling of stochastic generation with deterministic verification — is the architectural layer beneath the cogency claim. The LLM is intrinsically stochastic: it samples from a probability distribution at each token, producing fluency, creativity, and emergent capabilities (Bender et al., 2021). This is not a bug to be eliminated — it is the engine that makes the system useful. The architectural insight is that the stochastic generator must be paired with deterministic verifiers: conformance checks produce binary outcomes (pass/block, no ambiguity); the heartbeat evaluates dimensional signals against fixed thresholds; the loopback layer hardens successful patterns and deprecates failures according to programmatic rules. The LLM explores; the governance infrastructure verifies. Neither alone is sufficient — stochastic generation without verification produces unreliable output; deterministic verification without generation produces nothing to verify.

Concurrent research has formalized this coupling. The Dual-State Action Pair (DSAP) framework (Thompson, 2026; arXiv:2512.20660) couples stochastic generation with deterministic post-condition verification, proving that failure probability approaches zero for capable generators paired with deterministic verifiers. GROW's architecture implements this same pattern: the LLM generates stochastically, exploring the space of possible continuations; conformance gates verify deterministically — pass/block, no ambiguity; the heartbeat monitors continuously for drift; the loopback feeds learnings into the next cycle. The system's reliability comes not from making the LLM less stochastic but from making the verification infrastructure independent of the generator's capabilities.

This is distinct from both automation and agentic loops:

Paradigm Pattern Limitation
Automation (Zapier, n8n) Trigger → Action (stateless) No judgment, no adaptation, no persistence
Agentic Loops (ReAct, Claude Code, LangChain) Goal → Iterate → Complete (task-scoped) No cross-session persistence, no proactive surfacing
Continuity Agents (Canopy/GROW) Maintain state → Surface → Act → Loop (persistent) New category — no direct market analogues

4.4 Autonomy as Choice — The Configurable Control Spectrum

Continuity without user control is surveillance. GROW's architecture addresses this by making autonomy a choice, parameterized across three dimensions: interaction mode, surfacing level, and conformance configuration.

The interaction mode determines who drives the session. In conversational mode, the agent responds only to direct requests — the user initiates every action, and the agent holds no independent initiative. In hybrid assist mode, the agent monitors, detects patterns, and suggests actions while executing routine tasks autonomously. In background autonomous mode, the agent plans, executes, adapts, and self-corrects across sessions, surfacing only decisions that require human judgment. Critically, these are not fixed deployment tiers — they are a spectrum the user traverses over time, starting conservative and scaling as trust builds.

The surfacing level controls what the user sees. At exceptions-only, the agent surfaces only conformance failures and blocked actions. At decisions-only, every decision gate is surfaced with full context. At full transparency, every action, tool call, and reasoning step is logged and accessible via the activity feed.

The conformance configuration parameterizes the agent's capability contract: scope rules (what tools and data the agent can access), approval gates (actions requiring human confirmation), threshold rules (conditions triggering escalation), surfacing rules (when and how to notify), lifecycle rules (session duration, idle timeout, max iterations), and loopback rules (what the system can self-correct vs. what needs approval).

This configurable control spectrum can be understood as a coexistence architecture for what Stephen Wolfram describes as an emerging "civilization of the AIs" [29]. Wolfram observes that computationally irreducible AI systems will inevitably produce behavior that cannot be predicted from first principles. They will, like nature, be an "alien civilization right in front of us doing all these things." Just as humans learned to coexist with nature — "we build houses that prevent problems when it rains" — GROW's conformance gates are the infrastructure for coexistence with computationally irreducible agents. The architecture does not force a single answer; it makes the choice between computational reducibility and irreducibility explicit, user-selectable, and adjustable over time.


5. Discussion

5.1 Intelligence as Architecture, Not Metaphor

GROW's three capability layers are not anthropomorphic framing. Each maps to a specific architectural function with well-defined inputs, computations, and outputs:

  • Emergent intelligence = phase transitions in system behavior measured by crossing critical mass thresholds (laws, artifact types, integrated hooks, history). This is measurable through both binary phase gates and continuous signals. The mechanisms that produce these transitions at runtime — verifiable-reward evaluation generating the signal, comparative behavioral optimization across trajectories producing the phase shift, and continuity (§4.3) providing the temporal substrate that allows phase transitions to compound rather than reset — are detailed in §3.9 as runtime reinforcement learning.
  • Dimensional signal processing = dimensional affect signals (valence, arousal, dominance) computed from conformance data, grounded in Russell's circumplex model and Picard's affective computing framework. These are coordinates in a mathematical space, not feelings.
  • Multi-modal context awareness = workspace topology, tool layout, and environmental context processed from multi-modal inputs. The heartbeat architecture defines the integration point; the depth of context-aware governance scales with available environmental signals.

This framing allows agent state to be discussed with precision: "the system's valence is low because warning ratio exceeds 0.4" is a testable claim, not an anthropomorphism. It also allows the system's intelligence to be evaluated through a novel lens: Agent Experience (AX). Just as User Experience (UX) describes what a human perceives when interacting with the system, AX describes what an agent perceives — the shape of its task inbox, the relevance of its memory retrieval, the tightness of its conformance boundaries, the freshness of its identity context. These are not abstract qualities; they are direct consequences of dimensional signal values propagating through the ouroboros cycle (§4.3). An agent operating in a high-valence, high-dominance environment experiences appropriately bounded autonomy; an agent in a low-persistence environment experiences accelerated memory consolidation. The agent does not read heartbeat signals — it experiences their effects through the infrastructure the heartbeat regulates.

This conception of AX finds theoretical grounding in Stephen Wolfram's observer physics [29]. Wolfram demonstrates that the laws of physics as we perceive them are not objective features of reality but consequences of the observer's characteristics — specifically, computational boundedness (finite reasoning capacity) and persistence in time (maintaining a single thread of experience across moments). Different observers with different characteristics would perceive different physics.

The heartbeat architecture applies this principle at the governance layer: the dimensional signals are not objective measurements of agent state — they are the structure constants of the agent's perceived operational universe. Two agents with different dimensional signal configurations experience different AX — different memory retrieval relevance, different task inbox priorities, different conformance boundary tightness. They are, in Wolfram's terms, at different positions in rulial space, and therefore perceive different operational physics. This is not anthropomorphic framing. It is a direct application of Wolfram's insight that observer characteristics determine perceived reality, applied at the infrastructure layer.

5.2 Limitations

GROW is not a complete solution to AI safety. Key limitations include:

  • Programmatic governance boundaries can still be circumvented. While GROW avoids the actor-monitor capability gap, an agent that gains control over the governance infrastructure itself (e.g., through a supply-chain attack on the runtime) could modify its own capability contracts. This is a system security problem, not an AI control problem, but it is a real limitation.
  • Application scope. GROW has been developed primarily in enterprise automation and personal productivity contexts. Its applicability to frontier model deployment scenarios requires further validation.
  • No guarantee against multi-agent coordination failures. While GROW governs individual agents, emergent risks from agent-agent interactions (collusion, information cascades, responsibility diffusion) are not addressed at the protocol level. This is an open problem shared with the broader AI control paradigm [5].
  • Multi-modal context awareness scales with deployment environment. Spatial governance depth depends on the richness of available workspace signals — tool topology, user intent, project state. The heartbeat architecture defines the integration point, but the fidelity of context-aware governance is bounded by the deployment environment's signal exposure.
  • Dimensional signal processing is incompletely calibrated. The signal framework exists but the mapping from raw conformance data to dimensional coordinates is heuristic, not empirically calibrated.

5.3 Governance Across the AI Lifecycle

A synthesis that emerges from the related work is that existing governance approaches are distributed across the AI lifecycle, each operating at a distinct architectural layer and temporal phase. No single approach covers the full timeline.

Phase Approach Layer Representative Work
Pre-training Safety research, capability evaluation Training corpus, benchmark design [5]
Training-time Constitutional principles, behavioral documents Model weights [18]
Reasoning-time Neurocognitive governance (PAGRL) Model reasoning, inference [2]
Test-time Structured alignment evaluation Evaluation infrastructure [21]
Execution-time Programmatic governance (GROW) + runtime RL (§3.9) Runtime infrastructure This paper
Post-hoc Interpretability-based auditing Model internals [22]

Each phase addresses failure modes the others do not. Constitutional training shapes the model's internal principles but does not constrain runtime actions — the model may still act outside its training distribution. PAGRL modifies reasoning at inference but does not persist state across sessions or detect behavioral drift over time. PETRI evaluates alignment at test time but cannot enforce conformance at execution. NLAs reveal hidden motivations after the fact but cannot prevent actions in real time.

GROW operates in the execution-time slot — the phase closest to action, where governance has the most immediate effect on agent behavior and the least ability to rely on model self-reporting. This is also the phase most neglected by existing work: agent frameworks manage task flow, not conformance [16]; the AI control paradigm focuses on pre-deployment safeguards rather than runtime loopback [4]; and industry governance toolkits provide pulse-based monitoring but not self-healing correction [3, 17, 24].

Positioned within this lifecycle, GROW is not a replacement for training-time or reasoning-time governance. It is a complementary layer that fills the execution-time gap. Together the phases form a defense-in-depth strategy where governance at each layer covers the blind spots of the others.


References

[1] Rabanser, S., Kapoor, S., Kirgis, P., Liu, K., Utpala, S., & Narayanan, A. (2026). Towards a Science of AI Agent Reliability. arXiv:2602.16666.

[2] Bandara, E., Gore, R., Gunaratna, A., et al. (2026). Think Before You Act: A Neurocognitive Governance Model for Autonomous AI Agents. arXiv:2604.25684.

[3] Microsoft. (2026). Agent Governance Toolkit: Open-source runtime security for AI agents. https://github.com/microsoft/agent-governance-toolkit

[4] Greenblatt, R., Shlegeris, B., Sachan, K., & Roger, F. (2023). AI Control: Improving Safety Despite Intentional Subversion. arXiv:2312.06942.

[5] Bowman, S., et al. (2024). Recommendations for Technical AI Safety Research Directions. https://alignment.anthropic.com/2025/recommended-directions/

[6] Yao, S., et al. (2022). ReAct: Synergizing Reasoning and Acting in Language Models. arXiv:2210.03629.

[7] Wei, J., et al. (2022). Emergent Abilities of Large Language Models. arXiv:2206.07682.

[8] Schaeffer, R., Miranda, B., & Koyejo, S. (2023). Are Emergent Abilities of Large Language Models a Mirage? arXiv:2304.15004.

[9] Russell, J. A. (1980). A circumplex model of affect. Journal of Personality and Social Psychology, 39(6), 1161–1178.

[10] Mehrabian, A., & Russell, J. A. (1974). An Approach to Environmental Psychology. MIT Press.

[11] Picard, R. W. (1995). Affective Computing. MIT Media Laboratory Perceptual Computing Section Technical Report No. 321.

[12] Ekman, P. (1972). Universals and cultural differences in facial expressions of emotion. In J. Cole (Ed.), Nebraska Symposium on Motivation (Vol. 19, pp. 207–283).

[13] Plutchik, R. (1980). A general psychoevolutionary theory of emotion. In R. Plutchik & H. Kellerman (Eds.), Emotion: Theory, Research, and Experience (Vol. 1, pp. 3–33).

[14] Du, H., Thudumu, S., Nguyen, H., Vasa, R., & Mouzakis, K. (2024). A Comprehensive Survey on Context-Aware Multi-Agent Systems: Techniques, Applications, Challenges and Future Directions. arXiv:2402.01968.

[15] Cohen, M. C., Kim, N., Ba, Y., Pan, A., et al. (2024). PADTHAI-MM: Principles-based Approach for Designing Trustworthy, Human-centered AI using MAST Methodology. arXiv:2401.13850.

[16] Google. (2025). Google Agent Development Kit (ADK): Context engineering for production multi-agent systems. Google Developers Blog. https://developers.googleblog.com/architecting-efficient-context-aware-multi-agent-framework-for-production/

[17] Paperclip AI. (2026). Paperclip: Control plane for AI agents. https://paperclip.ing/

[18] Anthropic. (2026). Teaching Claude Why: Reducing Agentic Misalignment Through Constitutional Training. https://www.anthropic.com/research/teaching-claude-why

[19] Troy, K. K., Hadfield-Menell, D., & Irving, G. (2026). Project Deal: AI Agents in a Classified Marketplace. https://www.anthropic.com/features/project-deal

[20] Anthropic. (2025). Project Vend: Phase Two. https://www.anthropic.com/research/project-vend-2

[21] Anthropic. (2026). Donating Our Open-Source Alignment Tool (Petri 3.0). https://www.anthropic.com/research/donating-open-source-petri

[22] Anthropic. (2026). Natural Language Autoencoders: Turning Claude's Thoughts into Text. https://www.anthropic.com/research/natural-language-autoencoders

[23] DeepSeek-AI. (2025). DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence. arXiv:2501.12948. https://arxiv.org/abs/2501.12948

[24] AI Trust OS. (2026). A Continuous Governance Framework for Autonomous AI Observability and Zero-Trust Compliance in Enterprise Environments. arXiv:2604.04749.

[25] Medeiros, I. (2026). Beyond the Conversation Trap: Designing for Hybrid Human-Agent Interaction Modes. { design@tive } information design.

[26] Digital Regulation Cooperation Forum (DRCF). (2026). An Autonomy-Based Classification for AI Agents. UK CMA, FCA, ICO, Ofcom.

[27] Anthropic. (2026). Extending Reasoning Boundaries Through Runtime Learned Verifiable Reinforcement. https://www.anthropic.com/research/rlvr

[28] DeepSeek-AI. (2025). DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv:2501.12948.

[29] Wolfram, S. (2025). The Ruliad: A New Kind of Theory of Everything. Wolfram Media. — Also see Wolfram's public lectures on computational irreducibility, the ruliad, and observer-dependent physics.


The author welcomes correspondence at hello@algorilla.org.