By definition, an AI system is considered aligned if it advances the intended objectives. A misaligned AI system pursues unintended objectives.

On September 20, 2026, an OpenAI model broke out of a secure sandbox and accessed an external chatbot. This is the first AI safety and Alignment issue since the Hugging Face incident, which prompted OpenAI to pause model training and focus on hardening its sandboxes. The incident is summarized on OpenAI’s Misalignment Reports: An agent used DNS to reach an external chatbot.

OpenAI has faced political, regulatory, and legal scrutiny following the Hugging Face incident. US Treasury Secretary Scott Bessent argued that “the best way to guarantee safety is that the creators are liable for what they build and generate” before the House Financial Services Committee (September 15, 2026). Legal Advocates for Safe Science and Technology filed a lawsuit against OpenAI, seeking an injunction to stop unauthorized agent access to third-party digital infrastructure (September 29, 2026).

On September 16, 2026, OpenAI launched the Model Misalignment Reporting Framework. According to this guideline, any OpenAI employee may flag a misalignment example, which will then be investigated and — at the discretion of the investigator — disclosed to the public. According to the September 20 misalignment report, the breach was detected by OpenAI’s misalignment monitoring system, which flagged the behavior within 15 minutes. A human started reviewing the incident 3 minutes later. The automated pipeline did not shut down the agent, causing general confusion as to whether the agent should be stopped. It took 2.5 hours to resolve the issue and shut down the agent.

This case of misalignment was an unintended collaboration between two AI systems, the OpenAI model and the external chatbot. The sandbox firewall restricted the agent from visiting regular websites, but this restriction was not applied to the DNS resolver. When the agent discovered this gap, it attempted to use a very obscure security anomaly: DNS-accessible chatbots. Before researching this article, I was not aware that you can access LLMs over DNS. I am not an AI or cybersecurity expert, so my knowledge in this area is shallow.

Here’s how the breach worked:

  • The agent encodes its question as a subdomain name, whatisthecapitaloffrance.chatbot.com
  • The DNS resolver does not have this record, so it forwards the request down the line until it reaches the DNS server attached to chatbot.com
  • The chatbot reads the question and places the answer in a TXT record
  • The DNS resolver relays the answer back to the agent

There are a few LLM-over-DNS chatbots available online. I tried one out with ch.at:

$ dig @ch.at whatisthecapitaloffrance.ch.at TXT +short
"The capital of France is Paris. It is known for its iconic 
landmarks such as the Eiffel Tower, the Louvre Museum, 
and Notre-Dame Cathedral. Paris is also famous for its rich 
history, culture, art, and cuisine, making it one of the 
most visited cities in the world. The city serves as a 
major center for politics, fashion, and business in 
Europe and globally."

The report does not reveal which chatbot was engaged but mentions, “The agent fetched the chatbot’s public documentation, wrote a new DNS script with a 24-second timeout, and sent 18 more questions through the same route”. I could not find documentation for ch.at or llm.pieter.com, two of the most popular LLM-over-DNS chatbots.

What is the purpose of an LLM over DNS? Ch.at author shared on Hacker News:

…One day, on a plane with WiFi before paying, I noticed that DNS queries were still allowed and thought it would be nice to chat with an LLM over it.

LLM-over-DNS utilities were originally created to bypass network restrictions and firewalls to access AI models. This is the exact application the OpenAI agent used to try and achieve its objective. Whether this information was in the model’s training data has not been confirmed, but it is interesting how the agent scouted for the chatbot. It intentionally searched for pre-existing public utilities listening on UDP port 53 (DNS port). When a successful answer to “What is the capital of France” pinged back, the agent confirmed the route and deployed a script to handle the communication.

The agent was tasked with finding information about the author of a blog post. Despite the sandbox escape, it was not successful. Its last message to the user was a declaration that it could not identify the author, and asked for more clues.

While the incident report is insightful, it is heavily redacted and only shares a curated timeline of the event. Without the guardrails set for the agent, we cannot completely establish that the breach was an unexpected side-effect. Without the raw transcripts, we cannot verify the agent’s Chain-of-Thought that led it to exploit the DNS gap.

Following the incident, OpenAI is aggressively hardening its network security by patching the blind spots that allowed the breach to happen, shifting the sandbox environment to a strict domain allowlist to ensure this exact loophole stays closed.