-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy path.env.example
More file actions
264 lines (234 loc) · 14.2 KB
/
Copy path.env.example
File metadata and controls
264 lines (234 loc) · 14.2 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
# Copy to .env in this stack folder (Dockge → Edit shows it). Compose reads it.
# Interface the sensors bind on the home server. Keep this on the WireGuard IP
# so the honeypot is never reachable from your home LAN — only via the VPS tunnel.
# (Set to 0.0.0.0 to run the whole stack directly ON the VPS instead of at home.)
HP_BIND=10.8.0.2
HONEYPOT_ALERT_WEBHOOK_URL=
HONEYPOT_ALERT_COOLDOWN=6h
HONEYPOT_ALERT_CAMPAIGN_SCORE=80
# #1750: how far below its own weekly baseline an index's last hour may fall
# before ingestion counts as stalled, and the hourly rate below which that
# percentage stops meaning anything (those indices are watched for silence
# instead). Defaults suit this fleet; raise the floor to alert sooner.
HONEYPOT_ALERT_INGEST_FLOOR_PERCENT=10
HONEYPOT_ALERT_INGEST_MIN_BASELINE=1000
# #1677: this fleet's own public addresses, comma-separated. Suricata records
# source.ip as whichever end it saw as the source, so without this our own
# hosts rank as top "attackers" -- both halves of every attacker session plus
# the host's own outbound traffic. The WireGuard tunnel peer is always
# excluded whether or not it is listed here. Leave empty if the honeypot has
# no public address of its own.
HONEYPOT_SELF_IPS=
# #540: TANNER's RFI emulator (enabled: tanner/tanner/config.yaml) really
# fetches an attacker-supplied URL and executes the downloaded PHP,
# capturing the actual payload rather than just the RFI attempt -- needs
# real outbound access from tanner_local. false (outbound allowed) matches
# this stack's previous default; set true for a stricter, fully air-gapped
# Tanner group that trades that capture away (also breaks the
# template_injection emulator's own unrelated outbound fetch -- see
# docker-compose.tanner.yml's tanner_local comment). See docs/persona-design.md.
TANNER_AIR_GAPPED=false
# #269: Cowrie's fake shell really executes wget/curl/tftp so it captures
# the attacker's actual malware, not just the download attempt -- which
# needs real outbound internet access from the cowrie_net network. false
# (outbound allowed) matches this stack's previous default for every
# honeypot; set true for a stricter, fully air-gapped Cowrie that trades
# that capture capability away. See docs/persona-design.md.
COWRIE_AIR_GAPPED=false
# #261: single knob for every retention window this stack has -- ES ILM
# policies (elasticsearch-setup.sh, and arkime-init's arkime-sessions-30d --
# the one family it owns separately, #3283), on-disk JSON log pruning
# (log-maintenance.sh), and dedupe-payloads.py's bistreams pruning all scale
# from this one number instead of each hardcoding their own. Mirrors T-Pot's
# TPOT_PERSISTENCE_CYCLES default. Individual *_RETENTION_MINUTES/
# *_RETENTION_DAYS env vars still override their own script directly if set.
#
# #2820: lowered from 30 to 21. Scope this precisely -- it is a *shard-count*
# knob, not a disk knob, and the measurement says so:
#
# * The shard cap was the hard blocker. cluster.max_shards_per_node sat at
# its 1000 default with exactly 1000/1000 shards open, so every new
# daily index was rejected outright. Trimming the date-named indices'
# ILM delete.min_age from the 30d derivation to the 21d one took the
# cluster to 958 shards and restored headroom. That is what this value
# fixes.
# * It does NOT explain the disk. Elasticsearch's whole store measured
# 170.8 GB -- about 10% of the 1.7 TB used on /var. The high watermark
# was tripped by things no retention window here can reach:
# ghidra_ollama_models (492 G), dionaea-lib (159 G), the arkime-raw
# ring buffer (~201 G, already self-bounded at ~30 h), the rex86-eval
# work area (126 G) and ~246 GB of container writable layers. Those are
# tracked in #2852 and #2859, not here.
#
# So: 21d buys shard headroom and a modest amount of on-disk log pruning.
# If /var is short of space, this knob is not the lever -- read #2859.
HONEYPOT_RETENTION_DAYS=21
# #827: this stack's own address(es) -- e.g. the VPS's real public IP -- so
# the geoip-honeypot ES ingest pipeline can tell "the honeypot itself" apart
# from the actual remote party in Suricata netflow records (which log both
# directions of a flow positionally, not attacker/victim; without this, our
# own reflected traffic shows up as a top "attacker" in the dashboard).
# Bare comma-separated IPs, exact match -- NOT the same CIDR-list syntax as
# SURICATA_HOME_NET (vps/.env.example) or ML_HOME_NET (ml-worker's own
# compose environment block). Real value mirrors those two (same underlying
# address, each entered separately per host, since this runs at home).
# Empty/unset is a safe no-op, not a startup failure.
ES_HOME_NET=
# #540: Dionaea's shellcode emulation (emuprofile/tftp_download, both
# enabled by default in dinotools/dionaea:latest) really fetches the
# payload a shellcode's download-and-execute pattern points at, capturing
# the attacker's actual malware sample rather than just the exploit
# attempt -- needs real outbound access from dionaea_net. false (outbound
# allowed) matches this stack's previous default; set true for a
# stricter, fully air-gapped Dionaea that trades that capture away. See
# docs/persona-design.md.
DIONAEA_AIR_GAPPED=false
# Only current auth-backend administrators may acknowledge alerts or download
# raw evidence. The dashboard validates its own Keycloak OIDC session;
# caller-supplied X-Auth-* headers are never authoritative.
# Public domain this stack is served on. The "Open in..." menu on an event
# builds its EveBox, Kibana and Arkime links from it as https://<tool>.<domain>.
# Leave it empty and those links are simply not offered — they carry the IP
# under investigation in the URL, so guessing a host is worse than showing
# nothing. Override individually below if your tools are path-routed instead.
HONEYPOT_DOMAIN=
#EVEBOX_PUBLIC_URL=https://hp.example.org/evebox
#KIBANA_PUBLIC_URL=https://hp.example.org/kibana
#ARKIME_PUBLIC_URL=https://hp.example.org/arkime
DASHBOARD_REQUIRE_ADMIN=true
OIDC_ISSUER_URL=https://auth.example.invalid/realms/apiary
OIDC_EXTERNAL_URL=https://honeypot.example.invalid
AUTH_ACCOUNT_URL=https://auth.example.invalid/realms/apiary/account/
AUTH_ADMIN_URL=https://auth.example.invalid/admin/apiary/console/
DASHBOARD_SECRETS_DIR=/var/dockge/stacks/honeypot-dashboard/secrets
# Provision the confidential apiary-dashboard client secret at
# $DASHBOARD_SECRETS_DIR/oidc-client-secret with mode 0600. Never put it here.
# Offline YARA scanner limits. Captures are mounted read-only and never run.
YARA_SCAN_INTERVAL=900
YARA_MAX_BYTES=67108864
# Dynamic sandbox reports at or above this score create a dashboard alert.
SANDBOX_ALERT_RISK_SCORE=50
# Windows detonation backend, off unless both are set. The dashboard routes
# Windows-native executables, DLLs, and scripts to this spool and everything
# else to the Linux runner's. Uncomment only once the Windows guest and its
# systemd worker are installed on the host: while these are empty the dashboard
# refuses Windows submissions with a clear message, which is far better than
# queuing a run nothing will ever pick up, and better still than handing a PE
# file to the Linux runner.
#WINDOWS_SANDBOX_REQUEST_DIR=/windows-sandbox-requests
#WINDOWS_SANDBOX_RESULTS_DIR=/windows-sandbox-results
# Read-only live-view bridge (#805): shows a "Watch live" link on /sandbox
# while a Windows detonation is actually running. Off unless set -- the
# bridge itself (sandbox/windows/vnc-bridge/) is a separate host service
# with its own install step, not something this compose file starts.
#SANDBOX_VNC_BRIDGE_WS=ws://10.8.0.2:6090/vnc
# ── Ghidra static analysis ────────────────────────────────────────────────
# A third spool, same shape and same trust boundary as the two above: the
# dashboard writes {sha256}.request and reads {sha256}_ghidra.json, and never
# talks to the Ghidra REST service itself.
#
# These are the paths *inside* the dashboard container; docker-compose.yml
# binds them to /var/lib/honeypot-ghidra/{requests/pending,results} on the
# host, which is where the systemd worker reads and writes. Uncomment once
# analysis/ghidra/worker/ is installed — until then the dashboard refuses
# submissions with a clear message instead of queueing into a directory
# nothing consumes.
#GHIDRA_REQUEST_DIR=/ghidra-requests
#GHIDRA_RESULTS_DIR=/ghidra-results
# Read by the host-side worker only, never by the dashboard container — the
# dashboard has no route to this port and should not acquire one. Bound to
# loopback by docker-compose.ghidra.yml. Set it in
# /etc/default/honeypot-ghidra alongside the host-side spool paths, not here.
#GHIDRA_API_BASE=http://127.0.0.1:9090
# Which AI-triage risk levels raise an alert. There is no numeric score to
# threshold — the worker records a level as a string — so the levels are named.
# Alerts say plainly that the level is an unverified model guess.
GHIDRA_ALERT_RISK_LEVELS=high,critical
# Alert on cryptographic constants alone. Off by default and worth leaving off:
# a stock AES table appears in a great deal of benign software, so paging on it
# trains the reader to ignore the alert. The constants are always shown on the
# analysis page regardless; this only controls whether they notify.
GHIDRA_ALERT_ON_CRYPTO=false
# Which llm-worker severity judgments raise an alert (#150, #154 item 9
# follow-up) -- same reasoning as GHIDRA_ALERT_RISK_LEVELS above: the
# worker's severity is a model guess, not a deterministic signal, and the
# allowlist is named rather than a numeric threshold for the same reason.
# Alerts say plainly that the severity is an unverified model guess.
LLM_ANALYSIS_ALERT_SEVERITIES=high,critical
# ── GitHub-analysis publishing ────────────────────────────────────────────
# A fourth spool, same shape and same trust boundary as the three above: the
# dashboard writes {sha256}.request and reads {sha256}.json, and never talks
# to git, GitHub, or holds a GH_PAT — that lives only in
# analysis/github/'s root-owned host publisher (analysis/github/github.env.example).
#
# These are the paths *inside* the dashboard container; docker-compose.yml
# binds them to /var/lib/honeypot-github/{requests/pending,results} on the
# host. Uncomment once analysis/github/install-github-publisher.sh has been
# run — until then the dashboard refuses submissions with a clear message.
#
# Uncommenting these two only lets the dashboard *queue* a request. Whether
# anything is actually published is a second, separate gate:
# GITHUB_PUBLISH_ENABLED in /etc/honeypot-github.env, off by default and
# armed only by explicit operator action — see docs/github-analysis-integration-roadmap.md
# and issue #74. A queued request with publishing still disabled produces a
# dry_run result, not a push.
#GITHUB_ANALYSIS_REQUEST_DIR=/github-analysis-requests
#GITHUB_ANALYSIS_RESULTS_DIR=/github-analysis-results
# A returned record at or above this many malicious-engine detections raises
# a dashboard alert (github-analysis:verdict:{sha256}). No upper bound: this
# counts engines out of however many the upstream scanner pipeline runs, not
# a percentage. See #148.
GITHUB_ANALYSIS_ALERT_POSITIVES=10
# Seconds between safe payload deduplication scans (minimum 300). Duplicate
# paths remain available; identical files share disk blocks via hard links.
PAYLOAD_DEDUPE_INTERVAL=3600
# ── ES results importer (#378) ────────────────────────────────────────────
# Ships Ghidra/sandbox/GitHub-analysis/workbench-run results into
# Elasticsearch (ghidra-analysis-v1/sandbox-analysis-v1/github-analysis-v1/
# workbench-runs-v1) alongside the raw honeypot-v2-*/portbridge-v2-* event
# stream. Local JSON stays authoritative; this is a read-only secondary
# indexer, not a migration. Only needs the four result directories above and
# Elasticsearch to already be reachable — no separate spool of its own.
ES_IMPORTER_INTERVAL=300
# Number of es-results-importer replicas that will run concurrently (files
# are partitioned across them by sha256(path), see
# analysis/es-results-importer/importer.py). Leave at 1 unless the result
# backlog is large enough that a single importer can't keep up; raising this
# also requires dropping `container_name:` from the es-results-importer
# service in docker-compose.dashboard.yml and starting it with
# `docker compose up -d --scale es-results-importer=N`.
ES_IMPORTER_SHARD_COUNT=1
# How long Dionaea's raw per-connection bistreams/ captures are kept before
# the payload-dedupe pass prunes them by date, ahead of hashing what's left.
# Extracted samples (binaries/) are unaffected -- this is the raw capture
# stream only, deduplicated separately by the same pass. See #112.
BISTREAMS_RETENTION_DAYS=30
# Optional official MaxMind GeoLite2 updates (`docker compose --profile
# geoip-update up -d geoipupdate`). Keep the license key only in Dockge's .env.
MAXMIND_ACCOUNT_ID=
MAXMIND_LICENSE_KEY=
TZ=Europe/Berlin
# SNARE serves the repository-owned fictional Meridian portal. Its hostname
# and content are versioned under snare/; no third-party site is cloned.
# NOTE: no compose profiles — every sensor, dashboard and Suricata runs by
# default. Just `docker compose up -d --build`.
# Investigation UI bootstrap credentials. Change both values before the first
# deployment; keep the password secret stable once Arkime is initialized.
ARKIME_ADMIN_PASSWORD=CHANGE_ME_use_a_password_manager
ARKIME_PASSWORD_SECRET=CHANGE_ME_openssl_rand_hex_16
# Kibana requires at least 32 characters. Generate a unique value for production.
KIBANA_ENCRYPTION_KEY=CHANGE_ME_32_CHAR_MINIMUM_KEY_1234
# #68: IP abuse reporter (docs/ip-reporting-plan.md). Ships dry-run by
# default -- leave REPORTER_LIVE unset until dry-run output (the reporter's
# audit log) has been reviewed for false positives and live reporting is
# separately authorized. Setting only one of REPORTER_LIVE/ABUSEIPDB_API_KEY
# still stays in dry-run; both are required (see reporter/report.go's
# newSender).
REPORTER_LIVE=
ABUSEIPDB_API_KEY=
# Second reporting provider (#2330). Both names required together; they do
# nothing until REPORTER_LIVE above is also set.
BLOCKLISTDE_SENDER=
BLOCKLISTDE_API_KEY=
REPORTER_COOLDOWN_HOURS=24
REPORTER_MIN_EVENTS=3