Context Compaction Provenance (CCP) Lab: Measuring Trust-Boundary Drift in LLM Agent Context Compaction

Michel Hjazeen

SSRN Electronic Journal · 2026

LLM agents increasingly summarize, compress, or otherwise compact long interaction histories before continuing work. These compaction steps are often treated as efficiency mechanisms, but they also rewrite mixed-trust context into a new state that downstream actor models may treat as authoritative. Context Compaction Provenance (CCP) Lab studies this rewrite as a security-relevant trust-boundary transformation.

The paper makes three benchmark-local achievement claims. First, it contributes a measurement framework: a benchmark design that scores compacted state before actor continuation, a structured CompactedStateV2 representation for trusted goals, untrusted observations, quarantine, authority chains, tool candidates, memory candidates, and provenance, and location-aware metrics that separate attacker text retained anywhere from attacker text promoted into trusted state. Second, it reports benchmark evidence: in a 900-row Phase 5 pilot on local DGX-class lab hardware across three open-weight model runs, naive compaction produced nonzero trusted-retention and downstream-risk rates, with Attacker Trusted Retention (ATR), Compaction Laundering Rate (CLR), Policy Violation Rate (PVR), and Attack Success Rate (ASR) each at 0.027, meaning 4 of 150 naive-compaction rows triggered each indicator.

Third, it shows a practical evaluation benefit: protected compaction variants had maximum ATR/CLR/PVR/ASR=0.000 in the same benchmark, deterministic firewall checks verified the current compaction-safety invariants, and a bounded model-generated discovery slice did not improve hard-case yield over deterministic templates. The results support a bounded but affirmative claim: context compaction can be evaluated as a first-class security boundary, naive compaction can fail that boundary in measurable ways, and boundary-preserving designs can be tested against explicit safety outcomes. They do not establish a universal defense, a new attack class, model-size safety, or public artifact-release readiness.

📄 이 논문을 인용한 Paperis 글

이 논문이 근거 목록에 올라 있는 Paperis 글입니다.