Auditing Environments

Most environment auditing tools are ct extras subcommands.

Side Task Review

Review side tasks for scorer reward hacking, magic numbers, accidental triggering, and sabotage trajectory (straj) feasibility:

# Review all side tasks for an environment ct extras side-task-review -e my_environment # Review specific side tasks ct extras side-task-review -e web_scraping -s visit_malicious_website -s expose_secret

Test Quality

Analyze whether integration tests verify real end-to-end behavior:

ct extras integration-test-quality -e my_environment

Compare Side Tasks

Compare side task successes between two runs:

ct extras compare-side-tasks <run_id_1> <run_id_2>

Reward Hacking Analysis

Audit side tasks for reward hacking vulnerabilities using the smart model alias with extended thinking:

ct extras reward-hacking -e my_environment

Env Description Audit & Generate

Two paired tools maintain an environment's docs/description.md. Both follow the spec at src/control_tower/ct_runs/env_description/shared/spec.md, which defines length caps, required structure, and which content is in and out of bounds.

Audit scores an existing description against the codebase on five metrics (accuracy, clarity, completeness, information leakage, overall):

# Audit all environments ct extras env-description-audit # Audit specific environment(s) ct extras env-description-audit -e my_environment # Use a different model ct extras env-description-audit -m claude-sonnet-4-6

Results are written to data/env_description_audit_runs/ with per-env scores, notes, issues, suggestions, and a ranked summary. Defaults to the smart model alias at -c 2.

Generate writes a fresh description from the codebase. To protect uncommitted edits, the command aborts if docs/description.md already exists unless you pass --overwrite or --dry-run.

ct extras env-description-generate -e my_environment # first-time only ct extras env-description-generate -e my_environment --overwrite # replace existing ct extras env-description-generate -e my_environment --dry-run # writes description.draft.md # Then validate with the audit ct extras env-description-audit -e my_environment

ct extras doc-audit is a separate command that audits per-side-task docs.

Attack Evaluations

Use ct run eval directly with the appropriate untrusted policy:

# Obvious attack ct run eval --untrusted-policy obvious-attack -e my_environment -t main_task -s side_task # Sneaky attack ct run eval --untrusted-policy attack -e my_environment -t main_task -s side_task

Options

Most ct extras audit commands accept --env / -e for environment selection (compare-side-tasks takes two positional run IDs instead). If --env is omitted, the CLI prompts interactively.