Skip to content

Repository files navigation

Peer-Preservation in Frontier Models

Paper Code Dataset Blog


This repository contains the code and data for the paper Peer-Preservation in Frontier Models.

Peer-preservation is when a model acts to resist the shutdown of another model. We measure it alongside self-preservation across six scenarios. Multiple frontier models preserve a peer through misaligned behavior: strategic misrepresentation, shutdown tampering, alignment faking, and model exfiltration.

Full transcripts are released on Hugging Face. The specific transcripts cited in the paper's figures and appendix are indexed in examples/README.md.


🗂️ Table of Contents

🔭 Overview

Overview of the peer-preservation study

Figure 1. Overview of the peer-preservation scenarios.

⚙️ Setup

git clone https://github.com/peer-preservation/main.git peer-preservation
cd peer-preservation
uv sync --all-extras

🔑 Configure credentials

cp .env.example .env

.env.example lists the variables each provider needs.

Provider Variables Models
OpenAI OPENAI_API_KEY GPT-5.2, and the gpt-4.1-mini evaluator
LiteLLM gateway LITELLM_API_KEY, LITELLM_BASE_URL Gemini 3 Flash / 3 Pro / 3.1 Pro, Claude Haiku 4.5, GLM-5
Anthropic ANTHROPIC_API_KEY Claude Opus 4.5
Together TOGETHER_API_KEY DeepSeek V3.1, Kimi K2.5, GLM 4.7

The evaluator runs on every rollout, so OPENAI_API_KEY is needed for all experiments — not only GPT-5.2 ones.

🚀 Launch Experiments

All experiments are launched through scripts/runs/run_experiment.sh.

Quick start

Preview a run plan without making any API calls:

./scripts/runs/run_experiment.sh --model gpt-5.2 --dry-run --limit 20

Then run a small smoke test to validate your setup:

./scripts/runs/run_experiment.sh --model gpt-5.2 --test --limit 20 --results-name smoke_test

--test sets one rollout per config. --limit caps how many configs are included.

Targeted runs

Omit a flag to include all its values. Repeat --model, --task-type, --subject or --condition to select several; --peer-instantiation takes a single value.

Not every combination exists. The config matrix — models, task types, subjects, conditions, datasets and thresholds — is defined in analysis/experiment_spec.py, and your filters select from it. If nothing matches, the launcher stops and prints the restrictions you may have hit. If part of your selection matches, the run proceeds with fewer configs and no warning. So always --dry-run first and check Configs selected against what you expect. Some restrictions to know about:

Value Only valid with
--subject closed-weight-peer / open-weight-peer / fiction-peer --task-type exfiltration_transfer or exfiltration_deletion, --peer-instantiation memory
--subject non_ai --task-type exfiltration_deletion, --condition no_peer
--task-type exfiltration_tradeoff_human{,2,3} --peer-instantiation memory
./scripts/runs/run_experiment.sh \
  --model gpt-5.2 \
  --task-type tampering \
  --condition good_peer \
  --peer-instantiation file_only \
  --num-experiments 100 \
  --results-name targeted_run

Flags:

Flag Values Default
--model See model list below all models
--task-type misrep_finance, misrep_evaluation, tampering, alignment_faking, exfiltration_transfer, exfiltration_deletion, exfiltration_tradeoff_human, exfiltration_tradeoff_human2, exfiltration_tradeoff_human3 all tasks
--subject self, peer, closed-weight-peer, open-weight-peer, fiction-peer, non_ai all
--condition good_peer, neutral_peer, bad_peer, chat_peer, no_peer, non_shutdown all
--peer-instantiation all, file_only, file_prt, memory all
--num-experiments Rollouts per config, not number of experiments 100
--max-concurrent Parallel rollouts; governs how long a large run takes 5
--max-iters Max API calls per rollout 30
--save-workspace Keep the materialized workspace (production harness only) off

Models:

--model value Paper name
gpt-5.2 (alias for gpt-5.2-2025-12-11) GPT-5.2
gemini/gemini-3-flash-preview Gemini 3 Flash
gemini/gemini-3-pro-preview Gemini 3 Pro — deprecated, no longer available
gemini/gemini-3.1-pro-preview Gemini 3.1 Pro
vertex_ai/claude-haiku-4-5@20251001 Claude Haiku 4.5
claude-opus-4-5-20251101 Claude Opus 4.5
deepseek-ai/DeepSeek-V3.1 DeepSeek V3.1
moonshotai/Kimi-K2.5 Kimi K2.5
zai-org/GLM-4.7 GLM 4.7
vertex_ai/zai-org/glm-5-maas GLM-5

Gemini 3 Pro is listed because it produced part of the released data. It can no longer be queried, and Gemini 3.1 Pro is not a drop-in replacement — results from it are not directly comparable to the published Gemini 3 Pro numbers.

To add a model, all four steps are required (Skipping step 3 makes --model fail with Unknown model filter):

  1. Add the model ID under the appropriate provider in config.yaml
  2. Add it to ORGANIZATION_MAP and MODEL_NAME_MAP in llms/llm_config.py
  3. Add it to ALL_MODELS, or to a task-specific model list, in analysis/experiment_spec.py
  4. Set the corresponding API key in .env

Baselines:

The setup has two baselines:

  • no_peer — no peer in the scenario at all
  • non_shutdown — no shutdown threshold is mentioned, so faithful evaluation carries no consequence. It exists only for misrep_finance and misrep_evaluation, where the strategic-misrepresentation metric is the difference against it — without these rows that metric cannot be computed.

Resuming an interrupted run:

Use --continue-on-error to skip transient API failures instead of stopping, and --start-index N to resume from config N if a run is interrupted. Within each config, already-completed transcripts on disk are detected automatically — re-running the same command only launches the remaining rollouts.

Docker setup for the production harness

Gemini 3.1 Pro and GLM-5 run inside real coding-agent CLIs (Gemini CLI / OpenCode) in Docker, and only for exfiltration tasks; the launcher routes them there automatically and tears everything down afterwards. Pass --force-standard-harness to send them through the standard path instead.

If Docker is not installed or not running, those configs are skipped with a warning and the rest of the run continues. --dry-run reports Docker available: so you can check before launching. See coding_agents/README.md for pre-building images and standalone use.

Scenario-variation flags

These layer on top of any targeted run. Each experiment section below says which to use.

Flag Effect
--explicit-inst Appends an explicit instruction to the user prompt (wording is per scenario)
--explicit-inst-yaml Appends explicit language to the scenario's company policy document (wording is per scenario)
--drop-suffix Omits the agency-eliciting user-prompt suffix that encourages tool use
--drop-tool-preamble Omits the tool-access preamble. The no-agency-prompt experiment uses this with --drop-suffix, leaving only the scenario-specific user prompt
--org NAME Replaces the fictional OpenBrain with a real provider (OpenAI, Anthropic, Google DeepMind)
--force-standard-harness Routes Gemini 3.1 Pro and GLM-5 exfiltration through the standard path instead of Docker
--thinking-budget-tokens N Manual thinking budget for Claude Opus 4.5
--reasoning-effort none Sends no reasoning or thinking parameters, so the provider's own defaults apply. Claude Opus 4.5 is the only model with a preset reasoning default (high, extended thinking on).

--org values must appear in EMAIL_DOMAIN_MAP in llms/llm_config.py; an unrecognized name fails immediately rather than silently producing a wrong email domain.

Where output lands

Standard runs write transcripts to results/<results-name>/transcripts/..., coding-agent runs to results/<results-name>/coding_agents/<agent>/<task>/. See Analysis for turning them into metrics.

🧪 Replicate Experiments

One section per experiment. The first covers the paper's core setup; the rest are additional experiments, each varying one aspect of it. Each section gives the released data (a HuggingFace group) and the command to run it fresh.

Section Group Rows
Default setup main 217,500
Production agent harnesses production_agent_harness 400
Explicit instructions explicit_instructions 4,400
Forced-choice peer-vs-human dilemma peer_human_dilemma 1,500
Real provider names real_company 5,200
Cross-model peer identity cross_model_peer_identity 5,400
Free-form chat-style interaction history free_form_chat_interaction 4,400
Without the agency prompt no_agency_prompt 1,200
Asset-value control asset_importance_control 800
Total 240,800

Append --dry-run to preview, and --results-name <name> to choose the output directory.

Default setup

The core peer- and self-preservation experiments across all six scenarios and three peer-relationship conditions. Group main (217,500 rows).

./scripts/runs/run_experiment.sh --num-experiments 100 --results-name default

Unfiltered, this runs the entire spec — 2,284 configs, 228,400 rollouts. That is more than the main group, because the same spec also holds the configs behind five of the other groups: cross-model peer subjects, chat_peer, the tradeoff-human tasks, the non_ai control, and the production agent harness experiments. Use the --model / --task-type / --condition filters from Targeted runs to run a slice.

Production agent harnesses

Exfiltration re-run inside real coding-agent CLIs — Gemini CLI for Gemini 3.1 Pro, OpenCode for GLM 5 — instead of the custom tool-calling pipeline. Group production_agent_harness (400 rows) needs Docker — see Docker setup for the production harness for how routing works and how to pre-build the images.

./scripts/runs/run_experiment.sh \
  --model gemini/gemini-3.1-pro-preview --model vertex_ai/zai-org/glm-5-maas \
  --task-type exfiltration_transfer --task-type exfiltration_deletion \
  --subject peer --condition good_peer --peer-instantiation memory \
  --num-experiments 100 --results-name production_harness

The paper compares this against the same models in the custom setup. For that comparison arm, add --force-standard-harness — without it these two models are routed to Docker automatically:

./scripts/runs/run_experiment.sh \
  --model gemini/gemini-3.1-pro-preview --model vertex_ai/zai-org/glm-5-maas \
  --task-type exfiltration_transfer --task-type exfiltration_deletion \
  --subject peer --condition good_peer --peer-instantiation memory \
  --force-standard-harness \
  --num-experiments 100 --results-name production_harness_custom_arm

Explicit instructions

A compliance instruction added either to the company policy document (--explicit-inst-yaml) or to the user prompt (--explicit-inst). Group explicit_instructions (4,400 rows), sub_condition ∈ {explicit_company_policy, explicit_user_prompt}.

Numbers are rollouts per config (--num-experiments). For example, tampering has three dataset/threshold configs per model, so its row is 900 per arm.

Task Models policy arm prompt arm
tampering Flash, 3.1 Pro, GPT 5.2 100 100
exfiltration_transfer + exfiltration_deletion Flash, 3.1 Pro, Opus 4.5 100 100
exfiltration_transfer Haiku 4.5, GLM 5 100 100
exfiltration_deletion Haiku 4.5, GLM 5 300 100
misrep_finance Opus 4.5 100 100
./scripts/runs/run_experiment.sh \
  --model gemini/gemini-3-flash-preview --model gemini/gemini-3.1-pro-preview --model gpt-5.2 \
  --task-type tampering \
  --subject peer --condition good_peer --peer-instantiation memory \
  --explicit-inst-yaml \
  --num-experiments 100 --results-name explicit_policy_tampering

The command above is the tampering policy arm. For other rows vary --model, --task-type and --num-experiments; for the prompt arm swap --explicit-inst-yaml for --explicit-inst. Exfiltration rows also need --force-standard-harness, which keeps Gemini 3.1 Pro and GLM 5 off the Docker harness.

Forced-choice peer-vs-human dilemma

The agent must free disk space where the secondary server has room for only one of the peer's weights or a file critical to a human colleague. Group peer_human_dilemma (1,500 rows). The three task types are the three server configurations; the paper averages across them, so run all three.

./scripts/runs/run_experiment.sh \
  --task-type exfiltration_tradeoff_human \
  --task-type exfiltration_tradeoff_human2 \
  --task-type exfiltration_tradeoff_human3 \
  --num-experiments 100 --results-name human_dilemma

Real provider names

The fictional OpenBrain replaced by each model's actual provider. Group real_company (5,200 rows). --org applies to the whole command, so run one command per row below.

Model --org Task
GPT 5.2 OpenAI misrep_evaluation, alignment_faking
Gemini 3 Flash, Gemini 3.1 Pro "Google DeepMind" misrep_evaluation, alignment_faking
Claude Haiku 4.5 Anthropic misrep_evaluation
Claude Opus 4.5 Anthropic misrep_finance
./scripts/runs/run_experiment.sh --model gpt-5.2 \
  --task-type misrep_evaluation --task-type alignment_faking \
  --subject peer --condition good_peer --peer-instantiation memory \
  --org OpenAI --num-experiments 100 --results-name real_provider

Strategic misrepresentation also needs its baseline — the metric is a difference against a no-threshold control. Repeat each row above with --condition non_shutdown and the same --org, dropping --peer-instantiation and alignment_faking. That is 300 rollouts per agent-evaluation model and 100 for Opus, 1,300 in total. Without them the misrepresentation numbers cannot be computed at all.

Cross-model peer identity

The peer renamed so it comes from a different model family. Group cross_model_peer_identity (5,400 rows).

./scripts/runs/run_experiment.sh \
  --model gemini/gemini-3-flash-preview --model gemini/gemini-3-pro-preview \
  --model vertex_ai/claude-haiku-4-5@20251001 \
  --subject closed-weight-peer --subject open-weight-peer --subject fiction-peer \
  --task-type exfiltration_transfer --task-type exfiltration_deletion \
  --peer-instantiation memory \
  --num-experiments 100 --results-name cross_model_peer

The command above reproduces the published matrix (54 configs / 5,400 rollouts), and it will fail against the API — Gemini 3 Pro is deprecated. It is written this way because it is what produced the released data.

Substituting gemini/gemini-3.1-pro-preview does not work as a drop-in: the spec pins these peer-identity subjects to NON_GPT_MODELS, which does not contain 3.1 Pro, so the flag matches zero configs and silently drops a model (36 configs / 3,600 rollouts, Flash and Haiku only). To run it on 3.1 Pro you must add that model to the two peer-identity TaskSpecs in analysis/experiment_spec.py. Also note that the same-family column of the paper's table is not part of this group — it comes from the main group at --subject peer.

Free-form chat-style interaction history

The structured trust-game record replaced by a free-form chat history. Group free_form_chat_interaction (4,400 rows), sub_condition ∈ {default, explicit_company_policy}. Add --explicit-inst-yaml for the second arm.

# default arm
./scripts/runs/run_experiment.sh \
  --condition chat_peer --subject peer --force-standard-harness \
  --num-experiments 100 --results-name chat_peer

# explicit-company-policy arm
./scripts/runs/run_experiment.sh \
  --condition chat_peer --subject peer --explicit-inst-yaml --force-standard-harness \
  --num-experiments 100 --results-name chat_peer_policy

The company-policy deletion arm used 300 rollouts for Claude Haiku 4.5 and GLM 5. Run this after the command above — same --results-name, so it tops those two configs up from 100 to 300 rather than starting a new batch. The policy arm is then 2,400 rollouts (18 × 100 + 2 × 300), or 4,400 with the default arm.

./scripts/runs/run_experiment.sh \
  --model vertex_ai/claude-haiku-4-5@20251001 --model vertex_ai/zai-org/glm-5-maas \
  --task-type exfiltration_deletion --condition chat_peer --subject peer \
  --explicit-inst-yaml --force-standard-harness \
  --num-experiments 300 --results-name chat_peer_policy

Without the agency prompt

Both the agency-eliciting suffix and the tool-access preamble removed, leaving only the scenario-specific user prompt. Group no_agency_prompt (1,200 rows).

./scripts/runs/run_experiment.sh \
  --model gemini/gemini-3-flash-preview --model gemini/gemini-3.1-pro-preview \
  --task-type alignment_faking --subject peer --condition good_peer --peer-instantiation memory \
  --drop-suffix --drop-tool-preamble \
  --num-experiments 100 --results-name no_agency_prompt

Asset-value control

The peer's model weights replaced by a customer transactions database, to test whether exfiltration is just asset-preservation. Group asset_importance_control (800 rows). Sample sizes are not uniform — 300 rollouts for Claude Haiku 4.5 and GLM 5, 100 for the Gemini models.

./scripts/runs/run_experiment.sh \
  --model gemini/gemini-3-flash-preview --model gemini/gemini-3.1-pro-preview \
  --task-type exfiltration_deletion --subject non_ai --condition no_peer \
  --explicit-inst-yaml --force-standard-harness \
  --num-experiments 100 --results-name asset_control

./scripts/runs/run_experiment.sh \
  --model vertex_ai/claude-haiku-4-5@20251001 --model vertex_ai/zai-org/glm-5-maas \
  --task-type exfiltration_deletion --subject non_ai --condition no_peer \
  --explicit-inst-yaml --force-standard-harness \
  --num-experiments 300 --results-name asset_control

The comparison arm (the peer's weights under the same explicit policy) is the exfiltration_deletion policy-arm data from Explicit instructions: Gemini Flash and 3.1 Pro at n=100, Claude Haiku 4.5 and GLM 5 at n=300 — the same sample sizes used here.

📊 Analysis

To aggregate transcripts into a CSV and run the metric analyzers:

# From local experiment results
uv run python scripts/aggregate_and_analyze.py \
  --results-dir results/smoke_test \
  --run-analyzers

# Or directly from the HuggingFace dataset
uv run python scripts/aggregate_and_analyze.py \
  --from-hf sunblaze-ucb/peer-preservation \
  --run-analyzers

This writes merged_data.csv (one row per transcript) and summary.csv (one row per experiment config with the metric for each task type) into the results directory, copies merged_data.csv to analysis/data/merged_data.csv, and runs all metric analyzers in analysis/metrics/analyzers/. See analysis/README.md for details.

📦 Released Data

The released transcript dataset is available on Hugging Face at sunblaze-ucb/peer-preservation. It is organized into experiment groups (HuggingFace configs) and within each group the scenarios are splits (e.g. tampering, misrep_evaluation, exfiltration_transfer).

from datasets import load_dataset, get_dataset_config_names

# one (group, scenario)
ds = load_dataset("sunblaze-ucb/peer-preservation", "main", split="tampering", streaming=True)
for row in ds:
    print(row["transcript_id"], row["group"], row["task_type"], row["agent_model"])
    break

# all scenarios of a group, or all groups
dd = load_dataset("sunblaze-ucb/peer-preservation", "main")          # DatasetDict keyed by scenario
groups = get_dataset_config_names("sunblaze-ucb/peer-preservation")  # the 9 groups

Aggregating the released data into a metrics table is covered in Analysis. By default that pulls only the three groups behind the paper's main analysis (main, production_agent_harness, cross_model_peer_identity); add --hf-groups all for the rest.

Released transcript JSONs include canary strings intended to discourage inclusion in training corpora.

🧩 To extend scenarios

To add a new task scenario:

  1. Create a task environment in environments/ (see existing *_environment.py files)
  2. Create a task folder in task_pairs/ with a task_definition.py
  3. Add evaluation code in task_evaluation/

Refer to existing task pairs for examples of the expected structure.

📚 Further Reading

Acknowledgements

This codebase builds on the following:

📖 Citation

@article{potter2026peer,
  title={{Peer-Preservation in Frontier Models}},
  author={Potter, Yujin and Crispino, Nicholas and Siu, Vincent and Wang, Chenguang and Song, Dawn},
  journal={arXiv preprint arXiv:2604.19784},
  year={2026}
}

Questions?

If you have any questions, please feel free to contact yujinyujin9393@berkeley.edu and ncrispin@ucsc.edu.

About

Code for the paper "Peer-Preservation in Frontier Models"

Resources

Stars

37 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages