Finding
Found while researching #3043 (the dormant-wiki agent-coordination/sandbox-escape writeup). That finding's core lesson -- "constrain request EFFECT, not TYPE" -- has a concrete, unexercised instance in sandbox/forensic-egress-squid.conf.
forensic-egress-allowed-domains.txt is headed "Retrieval-only domains available to proxy-aware forensic samples" and lists .github.com, .githubusercontent.com, .github.io, .pastebin.com, etc. But the squid ACL that enforces this is a plain dstdomain allowlist gating CONNECT to ssl_ports (443):
acl allowed_forensic_domains dstdomain "/etc/honeypot-sandbox/allowed-domains.txt"
...
http_access allow sandbox_guest allowed_forensic_domains
Once a CONNECT to an allowed domain is permitted, squid tunnels the TLS session opaquely -- it cannot see, and does not attempt to distinguish, GET from POST/PUT inside it. So "retrieval-only" in the comment is aspirational, not enforced: a detonated sample with git/gh available (or a hand-rolled HTTPS client) can git push/use the GitHub API to write to a repo, gist, or release under an attacker-controlled account on github.com just as easily as it can git clone/pull one. .github.io is GitHub Pages, a shared multi-tenant domain where anyone can stand up <attacker>.github.io -- reachable the same way. pastebin.com accepts pastes via its own write API too.
This is the same request-TYPE-vs-EFFECT confusion #3043's dormant wiki demonstrated (there: a legacy wiki API where a read-typed request had a write effect; here: a domain-scoped CONNECT tunnel where "read-only" is a naming convention, not a technical restriction) -- not a novel bug, but a real gap between the stated intent and the actual enforcement, worth being explicit about since a detonated sample's exfiltration path through this proxy is exactly the scenario docs/sandbox/README.md's controlled-egress mode exists to constrain.
Why it's low urgency, not ignorable
SANDBOX_NETWORK_MODE=controlled is opt-in (isolated is the default per docs/sandbox/README.md), and everything through this proxy is already logged (access_log) and capped (reply_body_max_size 50 MB, request_body_max_size 1 MB) -- so an exfil attempt here is visible after the fact even though it isn't blocked in real time. This is a detection/design-intent gap, not an active unmonitored hole.
Possible directions (not scoped/decided here)
- TLS-terminate (MITM) the proxy for the allowlisted domains specifically so method/verb can actually be inspected -- meaningful engineering lift, and breaks certificate pinning some retrieval targets might use.
- Narrow the comment/intent instead of the mechanism: document that "retrieval-only" describes the operator's expectation of what samples need these domains for, not a technical guarantee, and rely on the request/reply size caps plus log review to catch a write-shaped session.
- Split truly read-only needs (raw file fetch from
raw.githubusercontent.com/objects.githubusercontent.com) onto a domain set that could plausibly be method-restricted, from domains that inherently need bidirectional access (github.com itself, for git clone over smart-HTTP, which is not simply GET) -- clarifies which domains are structurally unrestrictable regardless of proxy design.
No code changed by this issue -- filed as a follow-up while #3043 stayed research-only per this round's scope.
Finding
Found while researching #3043 (the dormant-wiki agent-coordination/sandbox-escape writeup). That finding's core lesson -- "constrain request EFFECT, not TYPE" -- has a concrete, unexercised instance in
sandbox/forensic-egress-squid.conf.forensic-egress-allowed-domains.txtis headed "Retrieval-only domains available to proxy-aware forensic samples" and lists.github.com,.githubusercontent.com,.github.io,.pastebin.com, etc. But the squid ACL that enforces this is a plaindstdomainallowlist gatingCONNECTtossl_ports(443):Once a
CONNECTto an allowed domain is permitted, squid tunnels the TLS session opaquely -- it cannot see, and does not attempt to distinguish, GET from POST/PUT inside it. So "retrieval-only" in the comment is aspirational, not enforced: a detonated sample withgit/ghavailable (or a hand-rolled HTTPS client) cangit push/use the GitHub API to write to a repo, gist, or release under an attacker-controlled account ongithub.comjust as easily as it cangit clone/pull one..github.iois GitHub Pages, a shared multi-tenant domain where anyone can stand up<attacker>.github.io-- reachable the same way.pastebin.comaccepts pastes via its own write API too.This is the same request-TYPE-vs-EFFECT confusion #3043's dormant wiki demonstrated (there: a legacy wiki API where a read-typed request had a write effect; here: a domain-scoped CONNECT tunnel where "read-only" is a naming convention, not a technical restriction) -- not a novel bug, but a real gap between the stated intent and the actual enforcement, worth being explicit about since a detonated sample's exfiltration path through this proxy is exactly the scenario
docs/sandbox/README.md's controlled-egress mode exists to constrain.Why it's low urgency, not ignorable
SANDBOX_NETWORK_MODE=controlledis opt-in (isolated is the default perdocs/sandbox/README.md), and everything through this proxy is already logged (access_log) and capped (reply_body_max_size 50 MB,request_body_max_size 1 MB) -- so an exfil attempt here is visible after the fact even though it isn't blocked in real time. This is a detection/design-intent gap, not an active unmonitored hole.Possible directions (not scoped/decided here)
raw.githubusercontent.com/objects.githubusercontent.com) onto a domain set that could plausibly be method-restricted, from domains that inherently need bidirectional access (github.comitself, forgit cloneover smart-HTTP, which is not simply GET) -- clarifies which domains are structurally unrestrictable regardless of proxy design.No code changed by this issue -- filed as a follow-up while #3043 stayed research-only per this round's scope.