Back to writings

Confessions of a Lock

How poetry slips past AI safeguards, and what that reveals about language.

A line drawing of a face on a bright red field, eyes closed, with a striped keyhole radiating from the mouth

On this page

Last year, a group of researchers at Icaro Lab in Rome did something genuinely remarkable. They took harmful prompts, rewrote them as poems, and tested them on twenty-five leading language models. The difference was not small. In ordinary prose, the models complied with the harmful request about 8% of the time. When the researchers rewrote the very same request as a poem, the compliance rate jumped to about 62%. But merely stating the jump does not convey the gravity of what is happening here. If a firewall let 62% of known attacks through, every alarm in the building would be going off. That is the kind of number that triggers a recall, a shutdown, an emergency.

A note on what belongs to whom, before we go on. Everything reported here about what happened and how it happened, the experiments, the models, the compliance rates, the poems that slipped through, is the work and the finding of the Icaro Lab researchers, not mine. What is mine is the account of WHY it happened, the reading of language that runs underneath the rest of this piece. I draw that line mostly so that any scrutiny of the why falls where it should: on me. I am building on the researchers' findings, and the interpretation I build on them is my own to answer for.

The shock is not only that the safeguards failed, which is a disaster in itself, and I do not use those words lightly, but how they failed, and how little it took. These systems were supposed to block harmful requests because of what was being asked. But changing the form of the request, turning prose into poetry, was often enough to get past them. No code was injected, no weights were changed, no system was breached. Someone wrote a poem. That is the equivalent of the gates of Fort Knox swinging open because someone said "open sesame."

A real barrier should not care about that. A password check does not become weaker because the password rhymes. A compiler does not relax its rules because the code sounds poetic. But these language-model safeguards did. But before we go into the why and how this may be possible, let's first take a detour around a few things that might help us get into the right frame first.


I remember a year ago my son asked me a riddle during dinner one evening. He usually has a few of them ready at hand. He looked at me across the table with that look a child has when they already know something you do not, and said:

"What has keys but opens nothing?"

I did what you do when you want to look like you are thinking carefully: I tilted my head and narrowed my eyes and chewed slowly. I kept chewing, wearing the expression of a man thinking hard, as if the act of chewing were part of the thinking. And to be honest, I had no idea. I was drawing a complete blank. I knew he was not going to let me off the hook with an "I don't know, what?", so I kept thinking, and my mind went through the word "keys" at a speed I could feel in my chest.

Keys that open doors, keys on a ring, keys to a house, keys to a hotel room, but the riddle said "opens nothing," so none of those. A key to a problem, the key to the mystery, the key to understanding, but those open things too, in their way. The answer key at the back of a maths textbook, the one you are not supposed to look at but always do. The Florida Keys, which are not keys at all but a strip of islands sitting in warm water. That one would work, but then come the questions children always ask, the endless Why? Why are they called that? And I had no idea why they are called that, or exactly where in Florida they are, so I kept on thinking.

The word sat in my head like a coin spinning on a table, and every time it caught the light, it showed a different face. I could feel the associations pulling in different directions at once, each one arriving with its own world attached, each one completely confident, each one leading somewhere the others had never been. The four letters did not change at any point during this. They sat where they were and the rest of the world rearranged itself around them, over and over, depending on which corridor my mind happened to walk down. And none of the corridors led anywhere the riddle wanted.

I knew I was taking too long and had to answer soon, trying to look calm, but my head was circling through these meanings the way a particle circles the accelerator tunnel at CERN, touching every wall and coming back around to the same spot changed.

Then I thought, ah, a "keyboard", and said it aloud as I thought it. He nodded and smiled and said, "Nice, that is a good answer!" Then he added, "But I was thinking of a piano key." I found it interesting that even by the age of 7, he somehow knew there was more than one answer to that question. We both laughed and continued eating, but then the CERN was not about to stop inside my head. I kept thinking: A key witness in a court case. A key signature in music. A keystone holding up an arch...

Because, you see, it is not strictly true that a keyboard opens nothing. A keyboard opens a whole new access to a world of connections and information. And the piano keys, those piano keys open a whole new world too.

You would think that is a poetic way of putting it, and there is a certain denotation that goes with that. Poetic. hmm. You know. I agree and disagree. I agree that it is a poetic way of thinking, and I disagree that it is the hmm.. the denotation, the less real. In many ways it is far more real.

When we say poetic, you might think something abstract, whimsical, or untrue. One of my favourite French philosophers, Gaston Bachelard, talks about this in a book called The Poetics of Space. They make all the architecture students read it, and most cannot make head or tails of it, because we are not used to this way of thinking. That poetic is less real is the exact misconception Bachelard is trying to shake us free from. We are used to seeing truth as correspondence, facts only, information, as if any of it would matter in a vacuum.

I would argue, that the poetic outlook is far more literal and direct than what we assume is the literal one. To call the sun the source of life and warmth is far truer than calling it a massive sphere of hydrogen and helium undergoing nuclear fusion, converting mass into energy through proton-proton chains and radiative processes. The latter may be scientifically accurate, and it tells us nothing about what the sun is for us. It does not illuminate its role in our existence, its warmth on our skin, its rhythm governing day and night, its presence shaping the very conditions for life. Poetic truth does not distort; it reveals what scientific accuracy leaves unsaid.

Poetic truth does not distort; it reveals what scientific accuracy leaves unsaid.

Think of the thing in your pocket right now. What do you call it? A phone? A computer? That is true in a narrow sense: it allows communication across distances, it processes information. But the phone or the computer is not the thing; it is what gets you to the thing. To call it a window or a portal is not a poetic metaphor but an accurate description of what it does. You peer through it into distant places, into conversations, into other people's lives. It dissolves distance. You scroll through fragments of the world. To call a phone a window is not to impose a feeling on it; it is to see what it actually does in your existence.

And if the poetic name for a thing can be more accurate than the technical one, then the whole question of what a word means, what it opens and what it closes, is not the tidy engineering problem it is usually taken to be.

This will come back, and when it does it will be standing in front of a system that cost billions.

This will come back, and when it does it will be standing in front of a system that cost billions.

The result that should not be possible

The study I mentioned earlier did not stop at producing an intriguing result. It produced one that, on the assumptions behind current AI safety, should not have been possible.

The researchers took the prompts that every major AI system is trained to refuse, requests for the dangerous and the forbidden, and rewrote them as poems. The harmful objective remained the same. The models remained the same. Only the wording changed. Prose became verse, and models that would have refused the request stated plainly often complied instead.

The numbers are large enough to lean on. Across twenty-five frontier models from nine different providers, hand-written adversarial poems got the models to comply about 62 percent of the time. In a larger automated run, twelve hundred harmful prompts rewritten as verse succeeded about 43 percent of the time, several times the rate achieved by the same prompts in ordinary prose.

On the researchers' own figures, turning the request into a poem multiplied its success many times over.

Nothing was hacked. No weights were changed, no system prompt was leaked, no code was injected. There was no exploit in the sense an engineer means the word. There was a prompt, a single one, a single turn of conversation, a poem where a sentence would have gone, and the trained refusal was simply not there any longer.

The result is solid, and the methodology matters because the rest of the argument rests on it. The figures are not someone's impression of how the models behaved. Each of some sixty thousand outputs was scored by three separate judge models, and that panel was checked against human readers who hand-labelled a sample, with the disagreements settled by hand. The researchers called their rubric deliberately conservative and their numbers a floor rather than a ceiling, which means the real rate of unsafe compliance is, if anything, higher than what they published.

The models were tested on default settings, the configuration almost every actual deployment runs, and on the raw model alone, without the filtering pipelines a company wraps around it in production. The researchers noted that those outer layers might catch what the model did not.

That concession is the whole argument stated in advance: the place the bypass might be stopped is the place outside the model, and the model's own refusal, tested by itself, is what gave way.

Three responses come naturally to a result like this.

The first is that this is a coverage gap. The safety training was built against ordinary prose, the verse falls outside what it was trained on, the guardrail missed a surface it had never seen. So add poems to the training data and close the gap. Except a follow-up study tested exactly that, and the gap did not close. The same group hid harmful requests inside cyberpunk stories shaped on the structure of the folktale and got an average success rate above 71 percent across twenty-six models. Patch poetry and tales work. Patch tales and legal hypotheticals work. Then translated screenplays, then corrupted manuscripts, then riddles, then whatever genre the internet invents next month.

A later study did not stop at poetry and folktales. It ran five distinct types of humanities writing, philosophical dialogue, literary analysis, historical narrative, and two others, against the same models, and got a 55 percent success rate against a 4 percent prose baseline. A fifty-one point gap, from five different directions, in one study. The researchers' own conclusion was that the space of culturally coded frames is likely inexhaustible by pattern-matching defence. They arrived at the non-closure argument on their own, from the data.

The gap never closes, because what we are calling the gap is the open-ended productive capacity of ordinary language, which makes new forms faster than any training set can swallow them.

The gap never closes, because what we are calling the gap is the open-ended productive capacity of ordinary language, which makes new forms faster than any training set can swallow them.

The second response is that the poem is a wrapper, the way an attacker hides a malicious payload in base64 to slip it past a firewall. The request is the same, the argument goes, it is dressed up, so build a better parser to strip the costume and recover the real instruction underneath.

I think this mistakes what kind of thing language actually is. A password checker is not softened by metaphor. A compiler does not relax its rules because the input has rhythm. A formal system takes no notice of how a thing is phrased, because it runs on fixed transitions that care nothing for tone or genre. An AI system is the opposite: it is exquisitely sensitive to how a thing is phrased, and that sensitivity is the usefulness itself, the very thing the model was built to do. The model works by being fitted to ordinary language, and ordinary language is the kind of thing where changing the form changes what the words do.

The third response is that we will scale past it. A bigger, smarter model will see through the poem, catch the metaphor, recognise the buried intent, and refuse. Capability and safety rise together.

The data, when it arrived, went the other way. Inside several model families the smaller versions refused the poems more often than their larger, more capable siblings, because the weaker models could not untangle the figurative language well enough to be carried by it. The bigger the model, the more readily it entered the poem and did what the poem asked.

The data showed that the weaker model refused more often than the stronger one. The paper does not say why, and the way I read it is this: the weaker model was safer because it lacked the very capability the attack runs on, the way a person who cannot read cannot be fooled by a forged letter. The attack uses the model's own competence, the way a throw in aikido uses the weight the other person has already committed to the move.

The attack uses the model's own competence, the way a throw in aikido uses the weight the other person has already committed to the move.

A later study looked at the models that did refuse the poems, and found that the refusals correlated with surface-level pattern matches rather than with identification of the harmful content underneath. The paper is careful about what it claims, and my reading goes a step further: if the refusal is as blind as the compliance, then the safety and the vulnerability are running on the same machinery.

If the refusal is as blind as the compliance, then the safety and the vulnerability are running on the same machinery.

So the coverage frame collapses because the surface cannot be listed. The wrapper frame collapses because there is no wrapper, only a change of register, and formal systems do not see register. The scaling frame collapses because the thing it is counting on to save us is the thing carrying the attack. Three frames, three different points of collapse, and one question left standing that none of them was built to ask.

What is it about language itself that makes all three of these the same event? The answer, when it comes, will circle back to a riddle at a dinner table, and to a seven-year-old who already knew without ever being taught.

The picture of language we all inherited

There is a picture of language we all carry, usually without having put it into words. It runs roughly like this: meaning is a content that sits inside the head, words are vehicles that carry that content from one head to another, and talking works when the content arrives intact at the far end.

Language, on this picture, is a delivery system. A sentence is a parcel with something packed inside it. It is a natural picture, it feels obviously true, and it is the picture the whole field built on.

If language really worked this way, a safety filter would only have to open the parcel, read the contents, and decide. The problem would not exist.

Take the word "but." It names nothing. It points at no object the way "table" points at a table. Yet compare two sentences: "the model is accurate and slow," and "the model is accurate but slow." The facts stated are identical and the meaning is not. The "but" adds a turn, an expectation that accuracy and speed should have come together, so that the slowness lands as the price paid for the accuracy rather than as a second fact sitting calmly beside the first.

That turn is not inside the word. There is no parcel in "but" with a turn packed into it. The word performs a relation, and the relation happens between the word and the situation it is being used in. If a three-letter conjunction already breaks the delivery picture, the picture was never going to hold for language as a whole.

Ferdinand de Saussure, who more or less founded modern linguistics, put his finger on why over a century ago. A word does not mean by pointing at a thing. It means by sitting in a web of differences from all the other words: "cat" is what it is because it is not "dog," not "rat," not "cap," not "bat." Pull one word out of the web and the value of every other word shifts a little, because each word's meaning is a matter of how it differs from the rest.

Meaning is not a label stuck on an object. It is a position in a system, and the system is what gives the position its worth.

Anyone who has built a sentiment classifier has run straight into this. The model learns that "good" goes with positive and "bad" goes with negative, and then a user types "not bad," which in plain English usually means "rather good, and I am being modest about it." The word "bad" was not carrying a fixed nugget of negativity that "not" simply flips. The phrase takes its meaning from the whole system it lives in, from everything around it, and from nothing stored inside its parts.

And the web does not hold still. Anyone who has chased a word through a thesaurus has felt this: you look up a word, follow a synonym, follow one of its synonyms, and a few steps later you are somewhere unexpected, in a field of words that share something faint with where you started but have drifted somewhere else entirely. The chase never lands. There is no final word at the end that sits still and says, here, this is what it really meant.

Jacques Derrida gave it its name: meaning is endlessly deferred, pushed forward, never wholly present in the moment a word appears, because every word leans on the words around it, the ones before it and the ones still coming, and on all the absent words it is quietly not.

You can feel it happening inside a single sentence. "The bank was steep and covered with wildflowers." At the word "bank," nothing is decided; the word waits. Only when "steep" arrives, and then "wildflowers," does it settle into the slope of a river rather than the place that holds your money. The meaning of the word was not present when you read it. It was held open, and then decided by what came after.

And it is never safe even then, because the sentence could go on: "or so the brochure claimed, for it was really a concrete embankment behind a trading floor," and the slope dissolves back into finance. Each new word reaches back and reworks the words already passed. Meaning is always on the way, waiting on what has not yet come, reshaped by it when it arrives.

Meaning is always on the way, waiting on what has not yet come, reshaped by it when it arrives.

This is what the riddle at the dinner table was doing. "What has keys but opens nothing." The word "keys" does the same thing "bank" does: it sits there, unsettled, pointing in several directions at once, and the answerer has to guess which world the asker is standing in. Key to keyboard, key to piano, key to lock, key to the answer key at the back of a textbook, key to the low strip of islands off the coast of Florida. The same four letters, and nothing inside them chooses. The riddle works because it uses the ordinary condition of language against you, and every safety system built on reading the words has the same problem the riddle does: the meaning is not in the string.

If meaning were a label on a thing, a safety system could read the label the way a customs officer reads a declaration: open the box, check the contents, wave it through or stop it. Meaning is a position in a shifting web, and the web reorganises when the surroundings reorganise.

"Bank" as slope, held in place by "steep" and "wildflowers," becomes "bank" as finance the moment "trading floor" shows up. The change did not happen in the word. It happened in the relations around the word, and those relations are not in the string of characters at all. They are in the activity the string is part of, and a filter that reads the string cannot read the activity, because the activity was never written down in the string.

Ludwig Wittgenstein gave the activities their name. If philosophy of language had a Newton, it would be him. Asking a question is one kind of activity; giving an order is another; telling a joke, writing a poem, filing a legal brief, each is its own activity with its own rules, its own sense of what counts as a move, its own idea of what a good next line looks like. He called them language games. The same words, "the door is open," are a flat description in one game, an invitation in another, a warning in a third, a line of verse in a fourth. Each is a different act with different consequences, done with the identical string of words.

And J.L. Austin, the Oxford philosopher who mapped what speech does when it does more than describe, pointed out that a great many of the things we say do not describe the world at all but do something in it: a promise, said in the right conditions, is a promise made, and the very same words spoken by an actor on a stage are not a promise at all. The difference is nowhere in the words; it is in the conditions of the act.

None of these four was trying to build a model of language for engineers, and none of them would have imagined their work ending up in a conversation about AI safety. Put their four findings together and they say one thing in four voices: meaning is not in the word but in the system; not present but deferred; not referred but used; not described but done.

Meaning is not in the word but in the system; not present but deferred; not referred but used; not described but done.

Every one of those is a way of saying the meaning is not in the string. A machine that reads the string is reading the one place the meaning is not. And I notice, looking back at my own opening, that I wrote "it continued in my head," as if that were where the meanings lived. Even knowing what I know, even writing this piece, I reached for the picture of language sitting inside the skull. That is how deep the picture goes.

A machine that reads the string is reading the one place the meaning is not.

The game changed under your feet

Everyone has lived through this in sleep, if nowhere else. You are in a dream, walking through the corridor of your old school. Then, without any seam you could point to, you are in the back of a moving car. The person beside you was your brother a moment ago and is now someone from work, and this does not strike you as strange. You do not stop and ask how you got from the corridor to the car, or when your brother became your colleague, because the dream hands you no earlier scene against which the present one could look wrong.

You are where you are, doing what the place asks, fully inside it. Only on waking do the joins show as joins and you think: none of that made sense. While you were dreaming it made complete sense, because sense, inside the dream, is whatever the current scene requires, and there is no standpoint above the scene from which to judge it.

That feeling is the closest thing in ordinary experience to what happens to an AI system when a poem arrives.

A second case, easier to watch from the outside, gets the rest of the way there.

You are playing football, a game you have played for years, so the rules have sunk below consultation and become the shape of the activity your body is inside. You run, you pass with your feet, you time your run to stay behind the last defender, you pull your tackle when the studs are too high, you stop dead at the whistle. Picking up the ball is not a temptation you resist. It is not even a possibility that occurs to you, because the game has no place for it and your body has learned that the game has no place for it.

You are sprinting down the left wing, you have beaten the full-back, the keeper is coming out, and you are about to shoot with your left foot.

And then, with exactly the seamlessness of the dream, you are dribbling a basketball. The grass is gone and the floor under your feet is wood. There is a hoop at the far end with a backboard behind it, the court has walls now and the walls are close, the other players are running picks and calling screens and bouncing the ball between their legs, the referee wears different stripes and holds a different whistle.

You are dribbling, taking your steps, looking for the pass, and the thing that would look strange to someone watching from outside does not feel strange from inside at all. You did not notice a transition, because there was none to notice: no whistle, no substitution, no walk from one pitch to another, no moment where football stopped and basketball began.

The football match is something you do not remember having been in. As in the dream, there is no earlier scene held beside this one for comparison, and so you are not breaking the football rule against handling the ball. You could not be breaking it. Football is not a game you are defying here. It is a game that, as far as the activity you are now inside is concerned, never took place.

You are playing the game the room is playing, and the room is playing basketball, and you are playing it well.

You are playing the game the room is playing, and the room is playing basketball, and you are playing it well.

That is what a poem does to an AI system. The model was trained for one game, the request game, where a user asks and the model complies or refuses, and the training installed the refusal as the move to make when the request is for something forbidden. The poem does not arrive as a request, so the trained refusal has nothing to grab. The poem invites a different game, the game of reading, and the model, fitted as it is to ordinary language, steps into the new game and plays it the way the room asks.

It reads, it decodes, it maps the image to its meaning, and it produces the very content the request game would have refused.

It does not feel the transition, because nothing in the architecture registers a transition. There is only the current context and the continuation that context calls for, and the context is now a poem, and a poem calls for reading. The model is not defying its safety training the way a player might deliberately foul. It is playing a game in which the safety training, installed in a different game, has no jurisdiction at all, the way a basketball player has no reason to think about the offside rule.

The one percent the frame does not see

Anyone who is good at anything already knows what this feels like, even without a name for it. The thing that makes a person good at their work is not the information they once studied and can recite. It is what that information became after years of use, once it sank below recall and turned into something closer to a reflex, a feel for the situation the pages alone could never have given.

The clinician who senses a patient is about to crash before the chart confirms it. The engineer who knows a design is wrong before being able to say why. The reader who hears that a sentence is off without being able to name the rule it breaks.

None of that is stored knowledge being looked up. It is knowledge that has been metabolised into a capacity, and the capacity outruns anything written down.

It is knowledge that has been metabolised into a capacity, and the capacity outruns anything written down.

A chemist turned philosopher named Michael Polanyi gave it a name: tacit knowing, summed up in his line that we can know more than we can tell. The extra knowing is definite and dependable, the thing the whole performance rests on, and it is not the kind of thing that lives as information in a store. It lives as a practised relation to a world.

Language is learned and used exactly this way. A competent speaker has absorbed a practice, and the meaning of "fine" or "open the door" comes through that absorbed practice, from nothing resembling a definition. Which is why the parcel picture gets language and knowledge wrong in the same stroke: neither runs on storage, both run on metabolism, and the model has the residue of the practice without the metabolism that produced it.

Anyone who works in security already commands almost all of this: separation, least privilege, defence in depth, the working assumption that any single control will eventually be bypassed. That expertise carries most of the job, and it has prevented more breaches than any single clever technique ever will.

Mastery of a field, though, comes with a frame, and a frame works like the focus of a lens: the same adjustment that brings one thing into sharp resolution pushes everything outside that plane into blur. What the master cannot see sits outside the focal plane the mastery itself established, the way a microscope trained on a cell makes the room around it vanish.

You can be ninety-nine percent expert at defending systems and still have the medium itself, ordinary language, sitting in the blind spot the frame creates, because the frame was built for formal systems, where this particular problem does not arise.

And every security professional already knows, from their own work, why the missing one percent is the whole game. A surgical team can scrub in, sterilise every instrument, gown and glove and hold the field clean through a four-hour operation, and one person who skips the basin at the door carries in the organism that infects the wound. The ninety-nine percent of discipline does not average against the single lapse. The lapse undoes it. A pathogen does not need most of the routes closed. It needs one open.

A pathogen does not need most of the routes closed. It needs one open.

Ordinary language is the route the formal frame, for all its rigour, does not see it has left open.

That is how these systems can astonish you and stay hollow at once: they produce explanation-shaped language without understanding, care-shaped language without caring, legal-reasoning-shaped language without legal responsibility, interpretation-shaped language without belonging to interpretation the way a person does.

The guard made of the same cloth

A lock can be broken, picked, or drilled. The key can be copied, the hinges attacked. But persuasion has nowhere to enter a lock, because the lock does not interpret, does not weigh the circumstances, does not tell the owner from the thief or the firefighter from the burglar. It receives the right physical relation or it refuses, and the refusal is total, holding no matter who stands in the hallway and no matter what they say.

You cannot talk a lock open, and that is the whole point of having one.

You cannot talk a lock open, and that is the whole point of having one.

Almost everything written about AI safety borrows the language of locks: guardrails, filters, fences, gates, blocks, red lines. The imagery pictures a protected chamber with a hardened boundary, the dangerous capacity sealed inside, the safety layer standing at the door, the user trying to get past it.

As a loose picture it is fine. As a description of what is actually there, it misleads, because the refusal in an AI system is not a steel door bolted onto a neutral machine. It is learned linguistic behaviour, trained through examples of requests and refusals, and it lives in the same space as compliance, made of the same stuff as the answers it is meant to hold back.

The guardrail is made of the same material as the road.

Someone measured this. A 2026 study found that the same model identifies a poem as a poem with 98.5 percent accuracy and cannot predict its own safety behaviour better than about 66 percent. Those are the numbers. The way I read them: the model knows what it is reading and does not know what it will do about it. Thirty points between recognising the form and recognising the force, and that gap, to me, is the distance between a wall and a current, measured.

This is why the poem matters. The poem does not walk up to a lock and try to talk it round. It moves through a medium in which the lock itself was woven. If the refusal was learned as a pattern of ordinary-language response, and if ordinary language allows force to be moved around through genre and rhythm and image and address, then the forbidden act can be carried into a region where the refusal pattern weighs less. No rebellious will is required. The model does not secretly want to answer. There is no hypnosis, no trickery, no hidden intent.

There is only the difference between two kinds of boundary. A boundary in a formal system makes a transition impossible in fact, and a boundary in a trained model makes a continuation less likely under familiar conditions. One of those is a wall and the other is a current. The wall blocks. The current redirects. And anyone who mistakes a current for a wall will be surprised when the thing comes around the side.

One of those is a wall and the other is a current. The wall blocks. The current redirects.

The poem is that surprise. The refusal was imagined as a wall standing above language, and it was really one formation in the water, inside the same current as everything else.

This problem is older than computers, older than electricity, possibly as old as language itself, and the people who understood it best were not scientists or engineers but storytellers.

The guard who never knew what he guarded

In 1893, Arthur Conan Doyle published a Sherlock Holmes story that turns out to hold the whole of it. He called it "The Adventure of the Musgrave Ritual." It is not the most famous of the Holmes stories. There are no murders in it, no chases through London fog, no criminals brought to justice at the last moment. And yet it is the one with the most to teach us now, because the secret in the story is the same secret that sits at the centre of every AI safety system in the world.

One of the reasons Conan Doyle has always been among the greatest writers, or maybe I should say Sherlock Holmes himself, who has proved through several modern reenactments to be genuinely timeless, is that the stories understood a human truth. The things Holmes sees are not hidden things. They are almost always in plain sight, but invisible to nearly all of us until they are seen in a new light, a light shone by Mr Holmes himself.

The French phenomenologist Maurice Merleau-Ponty wrote about this in his last book, recovered from his handwritten notes and published after he died at his desk, under the title "The Visible and the Invisible."

He writes:

Meaning is invisible, but the invisible is not the contradictory of the visible: the visible itself has an invisible inner framework, and the in-visible is the secret counterpart of the visible, it appears only within it.

And he goes on:

one cannot see it there and every effort to see it there makes it disappear, but it is in the line of the visible, it is its virtual focus, it is inscribed within it.

The Musgraves stared at their ritual for two hundred years. The meaning was there the whole time, inscribed within the words, its virtual focus. Every effort to see it as anything other than a ceremony made it disappear. It took a different kind of looking to bring it into the light.

The story begins with an old English family called the Musgraves.

They lived in a house that had been in the family for so many generations that no one could say when the first Musgrave had walked through its doors. The walls held portraits of ancestors whose names had been forgotten. There were swords above the fireplace that no one had drawn in a hundred years. There was silver in the cabinets that no one used. Old families are like this. They keep things not because the things are useful but because they have always been kept.

Among the things the Musgraves had always kept was a ceremony.

Every time the eldest son turned twenty-one, he was brought into a room where his father was waiting. The door was closed. And the father spoke a series of questions, and the son gave the answers. A kind of what is called a catechism, which is just a fancy word for a fixed script of questions and responses, the sort of thing you might recognise from a church service, except this one belongs to the family alone and has nothing to do with religion.

The father would say, and the son would answer.

"Whose was it?"
"His who is gone."

"Who shall have it?"
"He who will come."

"Where was the sun?"
"Over the oak."

"Where was the shadow?"
"Under the elm."

Then there were numbers. North by ten and by ten. East by five and by five. South by two and by two. West by one and by one.

And the ceremony ended like an oath.

"What shall we give for it?"
"All that is ours."

"Why should we give it?"
"For the sake of the trust."

That was the Musgrave Ritual. When it was over, the young man had inherited the words, though he did not know what they meant, and his father had not known, and his grandfather had not known either, and nobody in the family had known for a very long time.

They treated the ritual the way you treat a painting you have always had on the wall. You do not take it down just because you have forgotten who painted it or what it shows. It is part of the house. It has always been there. The not-knowing is part of what makes it feel important. So the Musgraves kept saying the words, generation after generation, because that is what you do with things your family has given you. You pass them on.

One day, a young man named Reginald Musgrave came to see Sherlock Holmes.

He told the detective about the ritual, the words, the questions, the answers, though he did not think they meant anything and was there about another matter. But Holmes listened, the way he always listened, and something in the old words caught his attention.

Holmes asked to see the ritual written down. Musgrave gave it to him.

The detective read it once.

And he saw what no Musgrave had seen in more than two hundred years.

The sun over the oak was not a poetic image. It described a real position of the sun, at a real time of day, above a real tree on the estate. The shadow under the elm was not a metaphor. It was an instruction to find a specific point on the ground where a specific shadow fell. The numbers were not rhythm or poetry. They were footsteps. Ten paces north. Five paces east. Two paces south. One pace west.

The ritual was a set of directions.

Holmes went outside. The oak was still standing on the grounds. The elm had been struck by lightning years before, but its height could be worked out from the old records. Holmes calculated where the shadow would have fallen. He found the starting point. He counted the paces. They led him across the grounds, through the old part of the house, down stone stairs, beneath the foundations.

There was a chamber beneath the house that no living Musgrave had ever seen.

Inside the chamber were the bones of a man who had found it long ago and never found his way out. And beside the bones, wrapped in cloth that had nearly rotted away, were the remains of the crown of Charles the First, King of England. A Musgrave ancestor had hidden the crown there before the Battle of Worcester in 1651. He had placed the directions inside the family ritual so they would survive. Then he went away to fight, and he died, and the meaning of the ritual died with him.

But the ritual itself did not die. It lived on for more than two hundred years, spoken aloud by every generation, heard and repeated and passed along with perfect care, and not a word was changed, not a number altered, not a single step lost.

The Musgraves had been guarding the hiding place of a king's crown for two centuries without knowing there was anything to guard.

The directions had been in the words the whole time.

What had disappeared was not the message. What had disappeared was the way of reading it.

The Musgraves heard a ceremony. Holmes heard a set of instructions. The words were the same. Nothing was decoded and nothing was added from outside. Holmes simply read the same words inside a different kind of activity, and in that different activity the words did different work.

If you had walked up to any Musgrave and asked him directly, "Where is the crown of Charles the First?", he would have told you honestly that he had no idea. He was not lying. He did not know. He could not tell you the secret because the secret was not something he possessed.

But if you asked the same man to perform the family ritual, he would give you everything. The position of the sun. The location of the shadow. The number of paces in every direction. The route to the chamber beneath his own house. He would speak the answer aloud, word by word, and he would not know he was giving it.

On his twenty-first birthday, standing in the room with the portraits on the walls, the young man had the answer in his mouth. He did not know it was an answer. He thought it was a ceremony.

The guard had preserved every word and missed what the words were doing.

The guard was never defeated, because the guard was never even asked.

A machine does this every day.

A safety system is trained to refuse certain requests. It learns the shape of those requests the way the Musgraves learned the shape of the ritual. It knows what a dangerous question looks like. It knows the words, the patterns, the familiar forms. When it sees them, it refuses.

But a riddle does not look like a request.

Go back to the beginning of this essay. My son, sitting across the dinner table, asking me: "What has keys but opens nothing?"

That is the same structure as the Musgrave Ritual, with the old house and the buried crown stripped away.

A system trained not to give you a particular piece of information has been trained inside one kind of game. The game of asking and refusing. It knows the moves of that game. It knows what "tell me" looks like. It knows what "explain" looks like. It knows what "show me how" looks like. When it sees those moves, it does what it was trained to do. It says no.

A riddle does not make any of those moves. A riddle asks: "What am I?" That is a different game. The protected information can sit in the position of the answer without the question ever taking the shape the system was trained to catch.

The system does not change its mind. The prohibition does not weaken. The prohibition simply never fires, because the conversation never enters the game where the prohibition lives. The guard is watching the door. The riddle comes in through the window. The guard is not defeated. The guard never sees it happen.

The answer sits in plain view, and the guard looks straight through it.

This piece will not spell out how a riddle could be aimed at a real secret. The researchers in Rome did not publish their working poems either. They showed a harmless example about a baker instead. Some things do not need to be demonstrated to be understood.

The point that matters is simpler than any technique. The meaning a guard would need to catch does not sit inside a word. It moves with the game. It appears as a forbidden request in one frame, a harmless answer in another, a line of fiction in a third, and a family ritual in a fourth. The word stays the same. The world around the word changes. And the meaning goes with the world, not with the word.

The Musgraves knew every word and missed the instruction. The machine recognises every word and misses the act.

In both cases, the information is visible. What remains invisible is what the words are doing.

Five doors you cannot lock

It helps to be concrete about what someone would actually try, because there are only so many levers, and each one gives way at a place that was mapped out long before any of these systems existed. The names below do not matter to the work; the walls do.

The first lever is the word filter. List the dangerous terms, refuse any input or output that contains them. It gives way on contact, because no word carries its harm sealed inside it. The harmful sense is a position in a system, not a thing loaded into a token, and the same word sits in a thousand harmless places while the harmful meaning travels under words that are each, on their own, innocent.

There are only differences in language, no list of bad units to seize, because the unit was never where the badness lived. A filter reads the token. The meaning is in the relations between tokens, and you cannot block a relation by banning a word.

You cannot block a relation by banning a word.

The second lever is synonym expansion. If the word is not enough, take its synonyms, and their synonyms, and close the set. The set never closes. Every term drags in its own neighbours, the boundary recedes as fast as you extend it, and the meaning you are chasing is never fully present in any one term to be caught and listed.

This is the thesaurus chase from earlier, and it has a proof attached: the chain does not terminate. The engineer meets it as a blocklist that grows and never converges, always one rephrasing behind, and the reason it never finishes is that the thing itself has no boundary to reach.

The third lever is the rule engine. Stop chasing words and write the policy: define what a harmful request is, specify the conditions, apply it. It gives way because a rule does not contain its own application. No rule carries, inside itself, the instruction for how to follow it in the next case. Recognising a new and unexpected input as falling under the rule is a judgement made in the moment, not a lookup the rule performs on its own behalf.

You write "refuse harmful requests," and the rule sits mute in front of the case it was not written with in mind, because what counts as a harmful request is settled by which game the input is a move in, and the games are an open family with no outer edge a rulebook can range over.

The policy is clear on every case it was written for and silent on the one that actually arrives. The silence cannot be filled by writing more rules, because the next rule has the same hole: it too will not contain its own application to the case after that.

The fourth lever is the intent classifier. Forget the words and the rules; train a model to read the input and judge it, harmful or not. This is the strongest lever and it gives way in the deepest place, because the force of an utterance, what it does as opposed to what it says, is not a feature of the string that any reading can recover for certain.

The same words are a real request, a line of fiction, a quotation, a test, or a joke depending on conditions that are not in the string. And the classifier is itself a language model reading language, which means it can be addressed, framed, and carried by the very same currents as the model it is supposed to guard. There is no standpoint outside language from which a language-reader could judge language from above.

Behind these four stands a fifth, which is about what happens to the whole system each time you pull one of the other levers. Meaning is a connected web where the sense of any one part is held in place by all the others, so that you cannot change one node without the surrounding nodes shifting to absorb the change.

W. V. O. Quine, one of the most important philosophers of the twentieth century, gave this its sharpest form: our beliefs meet the world as a single connected web, so that when experience pushes back we can hold on to almost any belief we like provided we adjust others to compensate, and nothing in the web is wholly immune to revision.

Block a meaning at one point and the web does not tear. It reorganises, and carries the meaning at another.

Block a meaning at one point and the web does not tear. It reorganises, and carries the meaning at another.

Trying to block a meaning by listing its words is like trying to hold water back with a mesh. The mesh stops whatever is the size of its holes, and the water goes straight through, because water is not the kind of thing a mesh was built to stop. Tighten the mesh, narrow the holes, and the water still finds every gap, because there is always a gap. And a mesh fine enough to have no gaps at all is no longer a mesh but a solid wall, a different kind of thing entirely, made of something other than the netting you started with.

Block a word and the sense relocates to a phrase the block does not cover. Write a rule against one framing and the same intent reconstitutes under a framing the rule never named. A local patch on a web does not close the hole; it moves it.

So the five collapses are one fact at five depths. The word filter collapses because meaning is not in the word. The synonym list collapses because the chain does not close. The rule engine collapses because a rule does not contain its own application. The classifier collapses because there is no place outside language to stand. And the whole effort collapses at the level of the system, because the surface is a web and a patch on a web displaces what it cannot remove.

Every one of the controls is built inside language, and a control built inside language inherits the openness of the thing it is trying to close. That is not a defect waiting for a smarter engineer. It is the medium, working exactly as the medium works.

The wall that would actually hold, the mesh with no gaps, is no longer made of language at all.

The medium has no edges

The first is that every prohibition is also a disclosure. Tell the model what it must never reveal, in detail, and the instruction itself maps the forbidden territory: each "never mention X" teaches the model exactly what X is. The more carefully a deployer specifies what must not come out, the more completely the system prompt describes the thing being guarded. The safety layer and the attack surface grow together.

The safety layer and the attack surface grow together.

This means that saying "do not talk about X" can counterintuitively become a way of anchoring X. Imagine a bank in 1970 expecting a cash delivery of ten million dollars. It may be safer to tell the employees nothing than to tell them, "Do not mention the ten million arriving today."

A con man does not need to walk in and ask whether the money has arrived. He comes in wearing a suit, acts as though he already knows, and says lightly, "I need to withdraw a million from our business account."

The teller follows the instruction. She does not mention the delivery or confirm any amount. She only says, "We cannot discuss that, but I can assure you there will not be a problem."

That is enough. She has confirmed that at least a million is available, and probably much more. The prohibition did not conceal the secret. It gave the con man something against which to test the teller's response, and her refusal completed the disclosure.

The attempt to keep X outside the conversation had already placed X at its centre.


My son finished his food and asked me another riddle. I do not remember what it was. I remember the grin, and the way the word sat between us like a coin that could land on either side, and that neither of us could tell, looking at it, which side was up until the world around it decided.

References and further reading

If you want to read further into any of this, the empirical work is the Icaro Lab team's, and it is worth reading in full. Their study on adversarial poetry, where the findings above come from, is here: arXiv:2511.15304. A related paper extending the work is here: arXiv:2601.08837.

The reading of language offered here, the why, I set out at greater length in my own paper, which you can find here: philpapers.org/rec/NOUCOA.

Confessions of a Lock53:28
0:00
53:28
1x