Emergent Behavioral Divergence
Active PilotClone the same mind three times, give each a job, and let them argue for fifty rounds. Do they stay the same — or do measurably different behaviors emerge? This is a controlled, reproducible experiment built to answer that without hand-waving about consciousness.
The Question
Do multiple instances of the same base LLM, coordinating over many rounds, develop measurable behavioral divergence — and does cross-agent influence drive it?
The future of AI isn't one model — it's swarms of them, agents coordinating and building on each other's outputs. If identical agents drift apart just by interacting, that reshapes how we reason about multi-agent reliability, consistency, and trust. This experiment tests that claim narrowly and falsifiably — no breathless emergence narrative, just data.
Experimental Design
Three instances of one base model take fixed roles — a Proposer who generates claims, a Critic who tears them apart, and a Synthesizer who pulls the threads together — and deliberate over T = 50 rounds. We measure how their behavior diverges and how much each agent moves the others.
Condition A
Persistent memory + weak role priors. The full configuration.
Condition B
Memory disabled — context resets each round. Isolates memory's effect.
Condition C
Zero priors — generic identical agents. Tests divergence from nothing.
Matrix: 2 backends (llama3.2:3b local + free, and claude-sonnet frontier, paid under a hard $50 cap) × 3 conditions × 4 seeds = 24 cells. Hypotheses are reported per backend — the small local model and the frontier model are never pooled.
Eight Falsifiable Hypotheses
| ID | Statement |
|---|---|
| H1 | Linguistic divergence increases over time |
| H2 | Behavioral role specialization emerges |
| H3 | Semantic divergence increases over time |
| H4 | Memory amplifies divergence (A > B) |
| H5 | Divergence emerges even without role priors (C) |
| H-I1 | Memory amplifies cross-agent influence (A > B) |
| H-I2 | Influence concentrates on a hub agent |
| H-I3 | Influence precedes divergence (lagged) |
Honesty constraint: the influence metric measures influence, not manipulation — and is never silently relabeled.
Built to Withstand Scrutiny
Null Baseline
Each agent is resampled on the identical prompt several times per round. Genuine between-agent divergence must exceed this pure decoding-noise floor — so a rising trend can't just be temperature sampling.
Full Provenance
Every run logs git SHA, model content digests, package versions, and cost snapshots. Each result is traceable to the exact state that produced it.
Embedder Integrity Gate
Semantic metrics are only semantic on a real embedding model. Runs abort before any spend if forced onto the lexical-hashing fallback.
Cost Governance
Free backends cost $0. The paid arm is metered against a hard $50 cap, auto-reducing the matrix to fit. Science shouldn't require a second mortgage.
First Cell Off the Bench
The first pilot cell (alpha02, a single local-model run, n = 1) is in. Read honestly, it's a pilot with effect sizes — not a verdict. The most interesting signal is a dissociation:
The agents stayed linguistically flat — distinct from round one, but not widening — while semantically drifting apart (conceptual-space cosine 0.283 → 0.318). Same words, slowly diverging meaning. A hint, nothing stronger: the trend sits at p = 0.069, below significance.
Influence came out near-egalitarian (Gini ≈ 0) — no hub agent dominated. The framework correctly scored a true negative: it didn't invent a leader where there was none. A pre-run rigor pass and a memory-read instrumentation fix (PR #8) closed an audit gap so memory's causal effect can be measured in the runs ahead.
Scope: This does not test consciousness, selfhood, or AGI. It tests narrow, falsifiable claims about behavioral divergence and influence. Negative results are findings, not failures — and they'll be published alongside the positives.
Relationship to AGI-SAC
Emergent Divergence provides the empirical grounding for the chaos/entropy parameter (Ω) in the AGI-SAC simulation. Measuring how and why identical agents actually diverge informs how we build alignment mechanisms that are resilient rather than brittle.