This is the repo for implementing method part in the paper Simulating Organized Group Behavior: New Framework, Benchmark, and Analysis
# Create and activate conda environment
conda create -n group-sim python=3.10 -y
conda activate group-sim
# Install remaining dependencies
pip install -r requirements.txtCreate a .env file in the project root with your OpenAI API key:
OPENAI_API_KEY=sk-...
The full pipeline has 4 stages, orchestrated by run_pipeline.sh:
Build CDT --> Adapt CDT --> Inference --> Evaluation
(01) (02) (03) (04)
Constructs a Codified Decision Tree from training scene-action pairs. Uses dual-embedding clustering to discover behavioral patterns, LLM-generated hypotheses as gate/statement candidates, and validation to filter them into a tree structure. Generates multiple candidates and selects the best via structural scoring + LLM voting.
python 01_base_cdt.py \
--data_path data/Apple/train.json \
--group "Apple Inc." \
--engine gpt-4.1 \
--output_dir packages \
--num_candidates 3Adapts a base CDT to each temporal phase. First computes gate traversal and a coverage matrix over all training events (cached), then for each phase: makes rule-based keep/delete/modify/add decisions based on coverage ratios, generates replacement statements via LLM, re-validates coverage, and outputs the adapted CDT.
python 02_adapt_cdt.py \
--base_cdt_pkl packages/Apple.cdt.v3.1.package.relation.pkl \
--train_data data/Apple/all_train.json \
--data_dir data/Apple/adapter_phase \
--group "Apple Inc." \
--output_dir adapted/Apple \
--phases phase_1 phase_2 phase_3Generates predictions for test events using a local LM (e.g. Qwen). Supports multiple modes: base (scene only), profile (scene + Wikipedia profile), base_cdt (scene + base CDT traversal), adapted_cdt (scene + adapted CDT traversal). Gate traversal uses either a local DeBERTa classifier or OpenAI API.
python 03_run_inference_tree.py \
--adapter_dir data/Apple/adapter_phase \
--adapted_cdt_dir adapted/Apple \
--cdt_pkl packages/Apple.cdt.v3.1.package.relation.pkl \
--group "Apple Inc." --group_id Apple \
--output_dir outputs/Apple \
--model Qwen/Qwen2.5-7B-Instruct \
--modes base,base_cdt,adapted_cdtUnified evaluation with two scoring modes (can run both at once):
- NLI: LLM judge scores prediction vs ground truth as entails (100) / neutral (50) / contradicts (0).
- 4dim: Scores alignment on 4 strategic dimensions (initiative, scope, magnitude, horizon) as match (1) / mismatch (0).
# Run both NLI and 4-dim evaluation
python 04_eval_unified.py \
--input outputs/Apple/predictions_qwen2.5_7b_instruct_base_cdt_Apple.json \
--output outputs/Apple/scored.json \
--mode both
# NLI only
python 04_eval_unified.py --input preds.json --output scored.json --mode nli
# 4-dim only, subset of dimensions
python 04_eval_unified.py --input preds.json --output scored.json --mode 4dim --dims initiative,scopeCommon code used across the pipeline: CDT node class, pickle loading, tree serialization/verbalization, gate traversal, coverage checking, and JSON extraction from LLM responses.
Runs all 4 stages sequentially. Edit the CONFIG section at the top to set your group, paths, and model choices.
# Run full pipeline
bash run_pipeline.sh
# Skip CDT build (reuse existing .pkl)
bash run_pipeline.sh --skip-build
# Skip adaptation
bash run_pipeline.sh --skip-adapt
# Only run evaluation on existing predictions
bash run_pipeline.sh --only-evalInput JSON files are expected to contain records with these fields:
| Field | Description |
|---|---|
history |
Scene/context text |
rewritten_action |
Ground-truth action |
group |
Group/entity name |
event_id |
Unique event identifier |
question |
Prediction prompt (for inference) |