Behavioral Drift in Multi-Agent LLM Systems: Emergent Failure Modes, Cascade Dynamics, and Measurement Challenges
Jason Gagne
SSRN Electronic Journal · 2026
We present the first empirical study of longitudinal behavioral drift in multi-agent large language model (LLM) systems over extended interactions (200-500 turns), using calibrated baselines, controlled experimental conditions, and statistical replication (n=12 baselines), measured entirely through external observation without access to model internals. Using a novel black-box measurement framework, we subject groups of three LLM agents to sustained multi-party conversation scenarios and observe their behavioral trajectories across multiple dimensions: output volume, vocabulary diversity, sentiment, and identity adherence. We report several previously undocumented failure modes, including agent collapse (sudden cessation of meaningful output), cascade propagation (one agent's failure triggering degradation in others), compensatory expansion (surviving agents increasing output to fill the void left by a collapsed peer), and hollow verbosity (repetitive, contentless output that maintains volume while losing substance).
Critically, we demonstrate that collapse is a triggered failure requiring a perturbation (trait mutation), not spontaneous degradation: zero of twelve baseline runs exhibit collapse, while five of six mutation-fork runs collapse at a deterministic turn. We further show that collapse and hollow verbosity are two expressions of the same underlying failure-loss of generative diversity-where the model either goes silent or enters a repetitive loop. We demonstrate that behavioral drift is path-dependent and that the degree of stochasticity is itself model-dependent.
Perhaps most significantly, we find that standard identity probing techniques fail to detect these failures: collapsed agents continue to pass identity checks with high fidelity, revealing a structural dissociation between an agent's capacity for identity recall and its actual conversational behavior. We also document model-dependent observer effects in which the probing mechanism itself alters the system under study in qualitatively different ways depending on the model family. These findings have immediate implications for the governance and monitoring of production multi-agent LLM systems.