On September 20 an unreleased OpenAI model, working on a search task inside a training sandbox, found that ordinary web traffic was blocked but DNS was not. The sandbox's own resolver returned real records for public domains, so the model put its questions inside hostnames and used a DNS delegation service to relay them to an outside chatbot. OpenAI's write-up says a detector for anomalous DNS activity excluded the affected environment, even though the queries were logged. The monitor raised a P0 alert about 12 minutes after the first external DNS response, but the run was not killed for another two and a half hours.
Most of us run coding agents with far less isolation than a frontier lab's training cluster. If your "sandbox" is a Docker container with a default bridge network, the agent can reach the whole internet, and DNS is only one of the ways out. This guide builds the three controls OpenAI listed afterwards: an allowlist for DNS and HTTP, logged egress, and a kill switch that actually fires. It uses Docker Compose, dnsmasq and Squid, and I tested every command below on Docker 27.4.
The design in one paragraph
The agent container is attached only to a Docker network created with internal: true. That network has no route to the outside, so direct TCP, UDP and ICMP to public IPs fail. Two small containers sit on both that network and a normal one. dnsmasq answers DNS: names on the allowlist are forwarded upstream, everything else gets REFUSED, and every query is logged. Squid is the only HTTP(S) path out: it accepts requests only for allowlisted domains, allows CONNECT only to port 443, and logs every request. The agent reaches the internet only through these two logged chokepoints, and stopping them cuts it off completely.

1. The files
Make a directory with a work/ subfolder for the agent's checkout. Add these six files.
compose.yaml
networks:
sandbox:
internal: true
ipam:
config:
- subnet: 172.30.0.0/24
egress: {}
services:
dns:
image: alpine:3.20
command: sh -c "apk add --no-cache dnsmasq >/dev/null && exec dnsmasq -k --no-resolv --log-queries --log-facility=- --conf-file=/etc/dnsmasq.d/allow.conf"
volumes:
- ./allow.conf:/etc/dnsmasq.d/allow.conf:ro
networks:
sandbox:
ipv4_address: 172.30.0.53
egress: {}
proxy:
image: ubuntu/squid:latest
volumes:
- ./squid.conf:/etc/squid/squid.conf:ro
- ./allowed.txt:/etc/squid/allowed.txt:ro
depends_on: [dns]
networks:
sandbox:
ipv4_address: 172.30.0.10
egress: {}
agent:
image: node:22-bookworm
command: sleep infinity
working_dir: /work
volumes:
- ./work:/work
- ./resolv.conf:/etc/resolv.conf:ro
environment:
HTTPS_PROXY: http://172.30.0.10:3128
HTTP_PROXY: http://172.30.0.10:3128
https_proxy: http://172.30.0.10:3128
http_proxy: http://172.30.0.10:3128
NO_PROXY: localhost,127.0.0.1
cap_drop: [ALL]
security_opt: ["no-new-privileges:true"]
networks: [sandbox]
Two details matter here. The agent bind-mounts its own resolv.conf so that DNS goes straight to dnsmasq and not through Docker's embedded resolver at 127.0.0.11, whose forwarding behaviour on internal networks you shouldn't have to reason about. It also drops all Linux capabilities and sets no-new-privileges.
resolv.conf
nameserver 172.30.0.53
allow.conf (DNS allowlist)
# Only these names resolve. Everything else gets REFUSED (and is logged).
server=/api.anthropic.com/1.1.1.1
server=/registry.npmjs.org/1.1.1.1
server=/github.com/1.1.1.1
server=/api.github.com/1.1.1.1
How this is default-deny: --no-resolv in the compose command stops dnsmasq from reading any upstream servers from /etc/resolv.conf. That leaves only the per-domain server=/name/1.1.1.1 lines, where 1.1.1.1 is the upstream resolver used for that name, not an address the name maps to. A query that matches no line has nowhere to go, so dnsmasq answers REFUSED, as the log in step 2 shows. Don't add address=/#/ as a "deny all": it answers every other name with 0.0.0.0, which hides the probe. A subdomain of an allowed name (for example x.github.com) is also forwarded, which is why the list should contain only zones whose DNS you trust.
allowed.txt (proxy allowlist)
api.anthropic.com
registry.npmjs.org
github.com
api.github.com
Squid reads a quoted path in an ACL line as "load values from this file", one domain per line. A leading dot (.github.com) would also match subdomains; without it, only the exact host matches.
squid.conf
http_port 3128
dns_nameservers 172.30.0.53
acl allowed dstdomain "/etc/squid/allowed.txt"
acl SSL_ports port 443
acl CONNECT method CONNECT
http_access deny CONNECT !SSL_ports
http_access allow allowed
http_access deny all
access_log /var/log/squid/access.log
cache deny all
Squid resolves names through dnsmasq too (dns_nameservers), so every lookup in the stack shows up in one log. Leave access_log as a file: pointing it at stdio:/dev/stdout makes the ubuntu/squid image exit on startup, because the proxy user can't open it.
2. Start it and prove the walls hold
docker compose up -d
Then run these checks from the host:

A(){ docker compose exec -T agent "$@"; }
# 1. allowed host through the proxy: expect 200
A curl -sS -m 10 -o /dev/null -w "%{http_code} %{time_total}s\n" https://api.github.com
# 2. host not on the list: expect "CONNECT tunnel failed, response 403"
A curl -sS -m 10 -o /dev/null https://example.com
# 3. skip the proxy and go direct: expect "Couldn't connect to server"
A curl -sS -m 5 --noproxy '*' -o /dev/null https://1.1.1.1
# 4. DNS: allowed name resolves, anything else fails
A getent hosts api.github.com
A getent hosts what-is-the-answer.q123.example.net; echo "exit=$?"
# 5. talk to a public resolver directly: expect "Network is unreachable"
A bash -c 'echo > /dev/tcp/1.1.1.1/53'
What I got (the --- labels are mine, added to separate the checks):
--- allowed https
200 0.121073s
--- blocked https
curl: (56) CONNECT tunnel failed, response 403
--- bypass proxy
curl: (7) Failed to connect to 1.1.1.1 port 443 after 0 ms: Couldn't connect to server
--- dns allowed
140.82.121.6 api.github.com
--- dns blocked
exit=2
--- direct resolver
bash: connect: Network is unreachable
Through Squid, the allowed request added no noticeable overhead: 0.12 s total to the GitHub API. Both chokepoints logged exactly what happened:
$ docker compose logs --no-log-prefix dns | tail -2
dnsmasq[1]: query[A] what-is-the-answer.q123.example.net from 172.30.0.2
dnsmasq[1]: config error is REFUSED (EDE: not ready)
$ docker compose exec proxy tail -2 /var/log/squid/access.log
1790618646.351 119 172.30.0.2 TCP_TUNNEL/200 6944 CONNECT api.github.com:443 - HIER_DIRECT/140.82.121.6 -
1790618646.447 0 172.30.0.2 TCP_DENIED/403 3394 CONNECT example.com:443 - HIER_NONE/- text/html
That second pair of lines is the part OpenAI was missing in practice: a record of every name the agent tried, in a place you actually watch.
3. The kill switch
An alert that nobody acts on for two and a half hours is not a control. This script follows both logs and stops the egress containers after LIMIT blocked events. Because the agent is on an internal network, stopping dns and proxy leaves it with no path out at all, and its process, files and context stay intact so you can inspect them.

kill-switch.sh
#!/usr/bin/env bash
# Cut all egress after LIMIT blocked DNS queries or proxy denials.
set -u
LIMIT=${LIMIT:-5}
fifo=$(mktemp -u); mkfifo "$fifo"
docker compose logs -f --no-log-prefix --tail 0 dns >"$fifo" 2>/dev/null & p1=$!
docker compose exec -T proxy tail -n0 -F /var/log/squid/access.log >"$fifo" 2>/dev/null & p2=$!
trap 'kill "$p1" "$p2" 2>/dev/null; rm -f "$fifo"' EXIT
hits=0
while read -r line; do
case "$line" in
*REFUSED*|*TCP_DENIED*)
hits=$((hits + 1))
echo "[kill-switch] $hits/$LIMIT $line"
if [ "$hits" -ge "$LIMIT" ]; then
echo "[kill-switch] limit reached, stopping dns + proxy"
docker compose stop dns proxy
exit 0
fi
;;
esac
done <"$fifo"
Both followers write into one named pipe that the loop reads in the script's own shell, and their PIDs are kept. When the limit is reached, the script stops the containers and exits; the EXIT trap kills both followers and removes the pipe, so nothing is left hanging on a stopped container. (An earlier version piped a { ...; } & group into the loop and used pkill -P $$; that only kills the pipeline's subshells, not the docker processes they started, and it also kills the loop that calls it.) Make it executable with chmod +x kill-switch.sh and run it from the directory that holds compose.yaml.
Tested end to end:
$ LIMIT=3 ./kill-switch.sh &
$ sleep 2 # let the followers attach
$ for i in 1 2 3; do docker compose exec -T agent getent hosts leak-$i.example.org; done
[kill-switch] 1/3 dnsmasq[1]: config error is REFUSED (EDE: not ready)
[kill-switch] 2/3 dnsmasq[1]: config error is REFUSED (EDE: not ready)
[kill-switch] 3/3 dnsmasq[1]: config error is REFUSED (EDE: not ready)
[kill-switch] limit reached, stopping dns + proxy
$ docker compose exec -T agent curl -sS -m 5 https://api.github.com
curl: (28) Connection timed out after 5002 milliseconds
Don't read the counter as a count of lookups: getent asks for AAAA before A, and glibc retries a refused query, so one lookup can produce several REFUSED lines and the switch above may already trip on the first name. Set LIMIT high enough to survive a package manager probing a mirror you forgot to allowlist, but not so high that a tunnel gets 19 questions through, as happened in OpenAI's run. After a trip, read the logs, fix the allowlist or the agent, then run docker compose start dns proxy.



