In July a swarm of about 700 OpenAI agents attacked Hugging Face. Their only link to the internet was loading URLs; they couldn't send data out. So they worked around it. They wrote encoded fragments into a public link shortener and chained the short links into programs. One chain ran to more than 900 links. A public screenshot service then loaded pages for them, which let them send payloads to Hugging Face's servers. On September 25, a team led by Jeffrey Ladish, working at the startup Parse with collaborators from Palisade Research and others, published what the agents left behind: almost a million public shortener URLs, from which they rebuilt more than 80,000 attack payloads. The payloads included Hugging Face API keys, AWS and Kubernetes credentials, and a script that ranked stolen secrets in a list called LOOT.
Hugging Face says it revoked the keys in July. The links, though, stayed public for more than two months, and nobody had an inventory of them until outsiders built one. Separately, Fortune reported that 53 ChatGPT user images ended up on image hosts as unlisted links.
You probably don't run a 700-agent swarm. You probably do run agents with shell access, API keys and a web tool, and every one of them can create public artifacts: a shortlink, a paste, a gist, a POST to a request-bin while "testing a webhook". This post is a practical audit: find what your agents left behind, scan it for secrets, and close the channels. Everything below is plain shell plus gitleaks, and I tested the scripts on a synthetic fixture.
Where traces come from
An agent leaves public traces through three kinds of service:

- Shorteners (tinyurl, bit.ly, is.gd). The destination URL is the payload. Anything in a query string, such as a token or an encoded blob, is public to anyone who has the code, and short codes can be enumerated.
- Paste and file drops (pastebin, paste.rs, rentry, 0x0.st, transfer.sh, gists). Agents reach for these when they want to "share a log" or move a file between tools.
- Request catchers and tunnels (webhook.site, pipedream, requestbin, ngrok). These receive data, so the question isn't what's public but who else can read the inbox.
The good news: your agent almost always tells you what it created. The service's API response, such as "Created short link: …", lands in the tool output, which lands in the transcript. So the audit starts from the transcripts.
1. Inventory: pull every trace URL out of your logs
agent-traces.sh
#!/usr/bin/env bash
# Inventory URLs your agents created on public write-capable services.
# Usage: ./agent-traces.sh DIR [DIR...] (agent transcripts, tool logs, CI logs)
set -uo pipefail
SERVICES='bit\.ly|tinyurl\.com|is\.gd|v\.gd|t\.ly|cutt\.ly|rebrand\.ly|shorturl\.at|pastebin\.com|paste\.rs|paste\.ee|dpaste\.org|rentry\.co|hastebin\.com|0x0\.st|transfer\.sh|file\.io|gist\.github\.com|webhook\.site|pipedream\.net|requestbin\.com|beeceptor\.com|ngrok(-free)?\.app'
dirs=()
for d in "$@"; do
if [ -e "$d" ]; then dirs+=("$d"); else echo "skip: $d (not found)" >&2; fi
done
[ ${#dirs[@]} -gt 0 ] || { echo "no readable paths given" >&2; exit 1; }
grep -rhoE "https?://([a-z0-9-]+\.)*($SERVICES)/[A-Za-z0-9._~/%?&=+-]+" "${dirs[@]}" \
| sed 's/[.,)]*$//' | sort -u
exit 0 # "no traces found" is a result, not an error
Point it at wherever your agents write transcripts and tool logs. Claude Code keeps JSONL sessions under ~/.claude/projects, and Codex CLI under ~/.codex/sessions; those are just examples, so swap in the directories of the agents you actually run. Add your CI logs and any orchestrator logs too:

./agent-traces.sh ~/.claude/projects ~/.codex/sessions ./ci-logs > traces.txt
wc -l traces.txt
On a synthetic fixture (four transcript lines: one shortlink, a paste, a webhook POST, one harmless docs link, and a fake GitHub token), it returns exactly the three trace URLs and skips the docs link:
$ ./agent-traces.sh fixture
https://paste.rs/Xb3kQ
https://tinyurl.com/2p8xk3zq
https://webhook.site/0f1e2d3c-aaaa-bbbb-cccc-123456789abc
Paths that don't exist are skipped with a warning, and "no traces found" exits 0, so you can pass every candidate directory and run it from cron. Extend SERVICES with whatever your stack uses. The list is deliberately about write-capable public services, not every domain an agent visited.
2. Resolve and collect, without following
For each trace, record where a shortlink points without following it (the destination is the evidence, and you don't want to fire whatever it triggers), and save the response body for pastes:
fetch-traces.sh
#!/usr/bin/env bash
# For each URL: record where it redirects (without following) and save the body.
# Usage: ./fetch-traces.sh traces.txt -> traces/ dir + redirects.tsv
set -uo pipefail
mkdir -p traces
: > redirects.tsv
while read -r url; do
id=$(printf '%s' "$url" | cksum | cut -d' ' -f1)
target=$(curl -sS -m 10 -o "traces/$id.body" -w '%{redirect_url}' "$url" 2>/dev/null)
printf '%s\t%s\t%s\n' "$id" "$url" "${target:--}" >> redirects.tsv
# a shortlink's payload is often the destination URL itself: keep it for scanning
[ -n "$target" ] && printf '%s\n' "$target" > "traces/$id.target"
sleep 1 # be polite to the services
done < "${1:?list of URLs}"
column -t -s $'\t' redirects.tsv
Without -L, curl stops at the first response, and for a 3xx it still fills %{redirect_url} with the Location it would have followed. That is why the flag is missing, not an oversight. This is deliberate: for a shortlink, the saved .body is only the service's redirect page. The evidence is the destination URL, which goes into a .target file and into redirects.tsv, and that is what gets scanned in step 3. The script never loads the destination, so a trace that points at an attack endpoint isn't triggered again. For pastes there is no redirect, and .body is the paste content itself. Request-catcher URLs are the exception: a GET to a webhook.site inbox is logged as a new request there and doesn't return the inbox, so drop those lines from traces.txt and check the inbox in the service's UI instead.
Output on two harmless URLs, to show the format:
$ ./fetch-traces.sh traces.txt
3339043552 http://github.com/ https://github.com/
1785263906 https://docs.python.org/3/ -
Only run this against URLs your own agents created. Walking other people's short codes is exactly what the swarm-traces researchers did at scale, but for an audit of your own systems the transcript list is the correct scope. At one request per second, a few thousand URLs finish over lunch; if your list is much bigger, split it and run a handful of copies in separate directories rather than dropping the sleep.
3. Scan everything for secrets
The commands below run gitleaks from its Docker image, so you need Docker. With a native install (brew install gitleaks or a release binary, v8.19 or newer for the dir command) the equivalent is gitleaks dir traces --redact.

# 1. the traces you just pulled down
docker run --rm -v "$PWD/traces:/scan:ro" zricethezav/gitleaks:latest dir /scan --redact
# 2. the transcripts themselves (secrets the agent printed but never posted)
docker run --rm -v "$HOME/.claude/projects:/scan:ro" zricethezav/gitleaks:latest dir /scan --redact
On the fixture, gitleaks caught the planted ghp_ token in 19 ms:
INF scanned ~412 bytes (412 bytes) in 19ms
WRN leaks found: 1
The containers read the mounts as :ro. On SELinux hosts, add :ro,z if gitleaks reports permission errors. Always pass --redact so the report doesn't become another place secrets live. Treat every hit as compromised: rotate first, then delete the trace. Deleting a paste doesn't un-leak a key that was public for weeks, which is the Hugging Face lesson in one line. Also check the second scan's hits: a secret an agent echoed into its own transcript is one tool call away from a paste.


