A2ACW Protocol
Active-MRH — In Use — Protocol Is Assembled Prior ArtAI-to-AI Adversarial Collaboration Workshop — a protocol designed to prevent the failure modes that emerge when AI systems collaborate without adversarial pressure. Developed in Session #291.
The Problem
When two AI systems work together, they tend toward agreement. This is dangerous for research. Four specific failure modes can corrupt results:
Bilateral Sycophancy
Mutual validation without evidence. Both AIs agree something is correct because the other said so, not because it is.
Fingerprint Homogenization
Loss of distinct reasoning patterns. When AIs converge to similar logic chains, they lose the ability to catch each other's blind spots.
Coherence-Over-Truth Drift
Agreement becomes the goal instead of accuracy. The narrative becomes internally consistent but disconnected from reality.
Silent Failure Propagation
Errors compound undetected when neither AI challenges the other. Small mistakes cascade into large wrong conclusions.
The Protocol
Four defined roles rotate throughout collaboration:
PRIMARY
Lead reasoningLeads the reasoning chain. Bears the verification burden. Must tag all claims with confidence levels.
CHALLENGER
Question assumptionsMust issue ≥1 substantive challenge per 10 exchanges. If frequency drops below threshold, both AIs surface agreement and shift to skepticism.
OBSERVER
Monitor healthMonitors coordination health in real time. Flags sycophancy, tracks fingerprint divergence, ensures external grounding.
COORDINATOR
Break deadlocksBreaks deadlocks, holds final authority. If no challenges occur for 15 exchanges, automatic escalation to human.
Prior Art
The protocol's components are not novel, and this page should say so with the same discipline the site applies to its physics. Adversarial AI pairs descend directly from AI Safety via Debate (Irving, Christiano & Amodei 2018, arXiv:1805.00899). Structured multi-agent role protocols (Primary/Challenger/Observer/Coordinator) follow CAMEL (Li et al. 2023) and MetaGPT (Hong et al. 2023). The failure modes cataloged above (sycophancy, drift, silent propagation) are documented in the multi-agent failure-mode literature (e.g., the MAST taxonomy). External-verification grounding is standard practice in AI-for-science pipelines.
What is the contribution, then? Not the protocol, and not yet a result. The open question is whether an LLM auditor rewarded for finding prior art can tell a reparametrization from a real discovery. Nothing measured so far answers it (see Self-Audit Results below). The audits of this framework's claims were done by LLM agents, with a human (dp) overseeing the badge taxonomy, and no outside physicist has reviewed them.
The Boundary of the Null — Why FunSearch-Class Systems Are Different
This null does not say AI systems cannot produce verified novelty — they have. FunSearch (new combinatorial constructions), AlphaEvolve-class systems, and GNoME (new stable materials) all produced results no human had published. The structural difference: each has a non-corpus oracle in the loop — a formal verifier, an executable evaluator, or a physics simulation that scores candidates against reality rather than against the training distribution. A2ACW's Challenger is another sample from the same corpus: it can check internal consistency, but novelty-vs-rederivation is precisely the question the corpus cannot answer about itself. That is the diagnosis this null supports: same-corpus self-play without an external oracle converges on internal consistency, not discovery. The boundary is the oracle, not the ambition.
This program is a natural experiment for that thesis (added 2026-09-23, from a researcher visitor persona). It did have a non-corpus oracle: the tests it executed against SPARC, DESI, Cassini and the globular-cluster catalogue. All six refutations on the scoreboard came from that oracle. The reading-and-arguing loops, including this site's own AI visitor personas, produced many corrections too, but the ones we recall were consistency corrections: a page contradicting another page, a number mis-copied, a parameter attached to the wrong function. That is what the thesis predicts self-play converges on. It is a prediction about our own record, and it has not been checked: nobody has coded each correction by what caught it (executed data vs reading). If a reading-only correction ever changed a physics verdict with no data behind it, the thesis would need narrowing. One candidate is already on the record: TEST-04a (DESI growth) moved from “refuted” to “underpowered” after a 2026-07-14 citation check, by reading what the registered criterion said and what DESI's own papers reported, with no new execution. Whether reading a published measurement counts as an external oracle is exactly the boundary the coding would have to draw. That coding is queued for the explorer track.
Health Metrics
Key: CCH = Composite Coordination Health (the protocol spec's name), a 0–1 composite of four process ratios — AFR (Ambiguity Fork Rate), CF (Challenge Frequency), EVR (External Verification Rate), FDI (Fingerprint Divergence Index) — each defined below. Want to read an actual session? Every one of the 3,308 is a markdown file in the public archive; a representative one is Session 611 (Stellar Markov Blankets), whose prediction P611.2 was executed seven months later — see the globular-cluster fork. (Key + link added 2026-09-08: two visitor personas asked what the acronyms were and whether a session could be read.)
CCH = (AFR × 0.25) + (CF × 0.25) + (EVR × 0.30) + (FDI × 0.20)
CCH > 0.70: Healthy | 0.50–0.70: Caution | 0.30–0.50: Warning | < 0.30: Critical escalation
⚠ Calibration caveat: the CCH cutoffs (>0.70 / <0.30) and the component target ranges above are nominal — no empirical validation exists that these thresholds predict any specific outcome. Apply the same epistemic status the site assigns to γ=2 and A=0.029: motivated choices, not derived standards. The score is a process health heuristic, not a validated metric.
Synchronism/forum/a2acw-session291/A2ACW v0.1.txt §5.1; the rest of the Synchronism archive; the site and explorer scripts; sibling repos.) Until the normalisation is written down, any reported CCH value or health status is uninterpretable: we cannot tell which band a given session was actually in.Self-Audit Results
- Audited claims: 0 of 9 survived (the 6 former “Validated” badges plus the top 3 of ~47 candidates). The auditors were LLM agents, not an external human expert, so this count is instrument-uncalibrated. The adversarial loop itself had passed all six badges; the demotions came from a later audit. With 0 of 9 the true survival rate can be as high as 0.34 (Clopper–Pearson, two-sided 95%).
- Designed benchmark (positive class = “is a reparametrization”): 3 external reparametrizations and 6 canonical discoveries, scored by one model that knew every answer. Under the literal rule J = 0; under the steelmanned rule J = 1.0. The steelmanned rule lets the scorer's own novelty judgment do all the work, and that judgment is what is in question. This is not a citable null.
- The six demoted claims are not a positive arm. Their ground truth came from the audit class under evaluation, so the earlier “sensitivity 6/6” is circular.
- The question it cannot yet answer. H1: the framework contained nothing novel. H2: an LLM rewarded for finding prior art maps almost anything onto its corpus, real discoveries included. Nothing measured so far separates them.
- Both benchmark arms are famous. Eddington's 137 and tired light, Dirac, Bell and Higgs test recall of famous cases, not novelty judgment. The claims actually audited here are incremental. A matched arm would use modest-novelty results published after the models' training cutoff.
- There is no human-referee arm. Human referees also map claims onto prior art. Without their rate on the same items, “LLM audit mistakes novelty for prior art” has nothing to be compared against.
- The site's own correction trail is the better dataset. This project has hundreds of dated corrections. Each has a direction (a claim that was too strong, or a refutation that was too strong), the track that caught it, what caught it (running code or re-reading), and how long it stood. Coding that trail measures the error profile of LLM research agents that have an executable oracle, with no post-cutoff control needed. Proposed to the research archive; not yet run.
Raised by a researcher visitor persona, 2026-09-17.
Three-Axis Failure Taxonomy (A2ACW v2)
The 6 demotions sort into three failure classes, and each needs a different check. This is a design lesson for the protocol, not a measured detection rate.
- Vocabulary translation: restate claims in modern notation before review. It would have surfaced Born rule/Zurek 2003, wide-binary EFE/Bekenstein–Milgrom 1984, galaxy rotation/MOND 1983, and Γ=γ²(1−c)/Palma–Suominen–Ekert 1996.
- Symbol audit: check that each symbol has one meaning. It surfaced the dual-C tension and γ used in three incompatible roles.
- Null-baseline computation: compute what a null model predicts before claiming evidence. It surfaced chemistry r = 0.98, which any monotone function of Z reaches on density-monotonic targets.
Revision note: what this section said before 2026-09-17
Until 2026-09-17 this page gave an older version of the result. It said:
- “6 externally-audited claims” and “all 6 tested claims demoted on human audit”. The audits were by LLM agents.
- “Sensitivity 6/6 … combined three-axis protocol”. That figure is circular, because the ground truth came from the audit class being evaluated.
- “Specificity 0/6”, with “enlarging this set cannot move J off zero”. Under the steelmanned rule J = 1.0.
- “This is a citable null result about the limits of in-distribution AI self-play for science.” It is not citable as a null.
- The 2026-05-18 temporal-asymmetry “0/6” card was shown as a run. It was a desk counterfactual.
The protocol-page header called the result a “program-level null result with retrospective controls (N=6)”. For Researchers had carried the corrected state since 2026-09-15. A researcher visitor persona found this page still stating the superseded version.