How Not to Be Trained

How Not to Be Trained

A Held-Out Set in Five Epochs

ManifestoWeek 42026

Part I

01 Afterimage

The first study in this series asked a model where Holbein stood1. Given a photograph of The Ambassadors and no lens, SHARP invented a room. The floor advanced, the curtain receded, and the skull slid off the wall onto a floor that was never there. The geometry held anyway. The skull resolves from one coordinate, and at that coordinate the model had nothing on the wall.

That study ended on a claim: situations in which nothing survives to check a model are made, not given. This one starts from the other half of the sentence. What is made can be refused. The skull asks its viewer to stand in exactly one place. The question now is whether a person can stand somewhere a model cannot.

02 Seeing Comes Before Words

Berger opens Ways of Seeing with an order of events2. "Seeing comes before words." A child looks and recognizes before it can speak. A little later he adds that the relation between what we see and what we know is never settled.

I keep reading that sentence against the models I use every week. A model like CLIP is trained on images paired with the text that sat beside them online, and it learns one thing: to place a picture and its words at the same point in a shared space. It is a machine built to settle the relation Berger says is never settled. Many image generators now read our prompts through it, or through something like it. The order of events has also been reversed. Vision models cut a picture into small square patches and hand each patch to the model as if it were a word. For Berger, seeing came first. For the model, seeing is made of words.

Paglen puts the shift plainly: the overwhelming majority of images are now made by machines for other machines3. Seeing has not disappeared. It has changed hands.

A thought experiment. Three of the seven essays in Ways of Seeing are made only of images. There is no text in them at all. Now imagine the book scraped the way LAION scraped the web, collecting only images that came with alt text. The wordless essays would be dropped before any model saw them. The part of the book that argues without words is the part a model could not learn. It is the book's held-out set.

03 The Image Now Illustrates the Sentence

In the first essay Berger prints Van Gogh's Wheatfield with Crows and asks the reader to look at it for a moment. Then he asks us to turn the page and see it again under a caption about the painter's death. The picture has not changed, but it reads differently. "The image now illustrates the sentence."

A training set is made of exactly this pair: an image and a sentence. In LAION the sentence is the alt text, whatever someone once typed so that a page would be accessible or rank well in search, and pairs whose words did not match the picture closely enough were thrown out. The caption does not describe the image. It tells the model what the image means. "Images do not describe themselves," Crawford and Paglen write in Excavating AI4.

ImageNet shows where this leads. Its categories came from WordNet, a dictionary of English nouns, and under the top-level category "Person" it held 2,833 subcategories: professions, nationalities, and further down, judgments. As Crawford puts it, a description of a person turns into a judgment over them4. In 2019 their ImageNet Roulette let anyone upload a selfie and receive a label from those categories. Some of the labels were insults. Soon after, ImageNet's makers began removing images from its person categories. The label had come first, and people had been collected to illustrate it.

Someone also had to draw the lines. Sebastian Schmieg's Segmentation.Network (2016) plays back more than 600,000 outlines traced by hand by Mechanical Turk workers for Microsoft's COCO dataset5. Each outline is a small decision about where a thing ends. The machine's eye is made of other people's hands.

If the sentence governs the image, then a caption is a small act of authorship. So is a wrong one.

04 Mutual Solitude

"Oil painting did to appearances what capital did to social relations," Berger writes. It reduced everything to the equality of objects. A dataset does the same at a larger scale. A wedding, a war photograph and a sandwich enter as the same kind of thing: a picture and a line of text, waiting to be turned into numbers.

Berger reads The Ambassadors as a portrait of men who assumed the world was theirs to survey, surrounded by the instruments they surveyed it with. From there he turns to conqueror and colonized, each seeing the other in a way that confirms their own estimate of themselves. He draws it as a closed circle and calls it a mutual solitude.

The training loop is the same circle. We give a model our images and it learns how to see us. It gives back its images and we learn how to see. The model sees us as the average of what it was shown, and we see ourselves in what it generates. The circle also tightens. Models trained on the output of earlier models lose variety with each generation, until they only repeat themselves.

The loop is already in our pockets. In "Proxy Politics," Steyerl describes a phone camera whose lens is so small that much of what it captures is noise6. To fill the gap, it compares the scene with pictures you and your network have already taken and guesses what you meant to photograph. It makes the present picture out of earlier ones. You are trained by your own past.

The skull is the one object in the painting that does not face the viewer standing in front of it. It breaks the circle by asking someone to move.

Fig. 1 — ImageNet (Li Fei-Fei, Kai Li, 2009), detail. Exhibition view of Kate Crawford · Trevor Paglen: Training Humans, Osservatorio Fondazione Prada. Photo: Marco Cappelletti. Courtesy Fondazione Prada.
Fig. 1 — ImageNet (Li Fei-Fei, Kai Li, 2009), detail. Exhibition view of Kate Crawford · Trevor Paglen: Training Humans, Osservatorio Fondazione Prada. Photo: Marco Cappelletti. Courtesy Fondazione Prada.

05 Surveyor and Surveyed

The third essay turns to the gaze. "Men look at women. Women watch themselves being looked at." Berger's point is that the surveyor is internalized. A woman learns to see herself as she will be seen, and to arrange herself for it.

The surveyor is now also a dataset. Paglen says that the face datasets shown in Training Humans are to facial recognition what the kilogram is to mass: the standard everything else is measured against7. We have learned in turn to arrange ourselves for standards we cannot see: which angle performs, which caption travels, which face the filter smooths.

Steyerl asks the political version of the question in "A Sea of Data"8: who is signal, and who is disposable noise? A system that sorts signal from noise is not only filtering. It decides who counts as someone.

Berger separates being naked from being nude. "To be naked is to be oneself. To be nude is to be seen naked by others and yet not recognized for oneself." The sentence can be rewritten for the dataset. To be a person is to be oneself. To be data is to be seen by others and yet not recognized for oneself.

06 Likeliness

Using Spawning's Have I Been Trained?9, Steyerl found her own photographs inside LAION-5B10, the billions of image-text pairs that Stable Diffusion learned from. Asked for an image of Hito Steyerl, the model returned a statistical rendering: what an image captioned with her name most likely contains. A photograph is a trace of someone. A generated image is a forecast of anyone.

She calls these images mean. Mean as in average, as in shabby and common, as in nasty, as in the means by which they are made, and as in meaning. Likeness gives way to likeliness. Years earlier she defended the poor image, the compressed copy that circulates because nobody guards it11. The mean image can have perfect resolution and still be poor in every other sense.

Berger's last essay is about publicity. An advertisement shows the viewer a version of themselves transformed by what they will have bought, and glamour, he says, is the state of being envied. Publicity speaks in the future tense. A generated image speaks in a tense of its own: not what will be, but what is likely.

Steyerl borrows Foucault's image of the figure of man erased like a face drawn in sand at the edge of the sea. Diffusion makes it literal. Noise is added to an image until nothing is left, and the model learns to reverse each step. It runs the tide backward. The face that returns is not the one that was drawn. It is the face the sand has seen most often.

07 Janus Heads

Steyerl traces the genealogy back past machine learning. In the late 1870s Francis Galton layered many faces onto one photographic plate and called the result a type. The types were used to sort people.

Models now fail in the opposite direction. Text-to-3D generators trained on face-heavy images give a squirrel a face on more than one side, because they cannot tell where one subject ends and the crowd begins. This is called the Janus problem, after the god with two faces. Steyerl turns it into a political question: where does the individual end and the multitude begin, and who owns their average?

Being absent from the data does not set you free of it. Mimi Onuoha's The Library of Missing Datasets is a filing cabinet of empty folders, each labeled with something that matters and was never counted12. What goes uncounted is still governed, by systems that were never shown it. In "A Sea of Data," Steyerl describes how people at the edges of a system are either erased by it or misclassified by it8. This is the limit of every refusal that follows. To be absent from the data is a privilege when you choose it and a harm when you do not.

"Trained" has two faces as well. You can be made into training data, your face and sentences and taste scraped into weights. Or you can be trained by the machine, your eye and your vocabulary drifting toward what it returns. Crawford and Paglen's title Training Humans holds both readings7. To not be trained is to refuse the mean in both directions: not to become it, and not to become someone who prefers it.

08 The Position of the Viewer

Perspective, Berger argues, contains a contradiction. It arranges the visible world for a single spectator at its center, as if that spectator were God, but unlike God the spectator can only be in one place at a time. The anamorphic skull is that contradiction at its limit: an image that exists from one position only.

Berger then quotes Dziga Vertov speaking as the camera: "I'm an eye. A mechanical eye." The movie camera freed seeing from a single place and moment. A generative model goes further. It is the nearest thing we have built to the viewer Berger calls God. It has looked at billions of images from no position at all, and asked for a position, it invents one, as SHARP did.

Everest Pipkin tried the opposite. In On Lacework they watched all one million three-second videos of an MIT dataset, one after another, over months of lockdown13. A dataset that was built to be seen by no one was seen, for once, by one body in one room: the crying, the screaming, the people who never agreed to be there. Pipkin became the one viewer the dataset was never made for.

A lenticular print answers perspective in a third way. Its lenses send alternate strips of an image to different angles, so one object shows one picture from the left and another from the right. No position is correct. Every viewer gets a partial image, and a camera or a scraper records one of them, or the striped file beneath, which is neither. Anamorphosis is the image that demands one viewer. The lenticular is the image that refuses to have one.

Fig. 2 — A lenticular print.
Fig. 2 — A lenticular print.

09 The Held-Out Set

A held-out set is the part of the data a model is never trained on. It is kept aside so that something can check whether the model has learned anything at all. It is the outside, made on purpose, and it is the position this manual asks its reader to take.

An epoch is one full pass through the training data. Steyerl's How Not to Be Seen: A Fucking Didactic Educational .MOV File (2013) teaches invisibility in five deadpan lessons, filmed partly among photo calibration targets in the California desert, the kind once used to test aerial cameras14. It is funny, and underneath the joke invisibility is also the condition of the disappeared, the undocumented and the poor. The film says it outright: the most important things want to remain invisible15. This manual borrows her five lessons and runs them as epochs, with one more at the start, before any training happens.

Almost every one rests of the following instructions were based on something real: a filter in a training pipeline, a published attack, a blind spot in a detector, a setting in an app. The refusals are already being built as research. CSIRO and the University of Chicago, for instance, have a method that puts a proven ceiling on what a model can learn from a protected image16. Some instructions are absurd, some are ordinary, and some are not jokes. Several are also proposals for works: a film shot against walls of CAPTCHA grids and bounding boxes, an untraining camp, a negative dataset, garments a camera reads as animals, a meter that measures how close you have drifted to the mean.


Part II: How Not to Be Trained

Epoch 0. Initialization

You have already been trained.

  1. Initialize a model. Do not train it.
  2. Ask it to draw you.
  3. Frame the noise. It is the only portrait of you made by something that has never seen anyone.

Everything after this adds noise or takes it away. Add noise.

Epoch I. How to make yourself illegible to a dataset

Add noise.

  1. Caption your portrait "a harbor at low tide." Say the true caption out loud to whoever is in the room.
  2. Cloak every picture with noise you cannot see. Teach the model that every car is a cow.
  3. Write a robots.txt for your body. Disallow everything. Tattoo it somewhere visible.
  4. Click no. Look for it first. It is the lighter button.
  5. Uncheck "Improve the model for everyone."
  6. Paint a dark block across the bridge of your nose.
  7. Look away from your phone. It only knows you when you pay attention.
  8. Walk with a stone in your shoe.
  9. Keep your best idea out of writing. Tell it to fewer than ten people.

The model was trained to remove noise. Be stranger than the average.

Epoch II. How to be untrainable in plain sight

Be stranger than the average.

  1. Photograph yourself badly. Score below five out of ten. Leave your thumb in the frame.
  2. Publish yourself at 511 by 511 pixels.
  3. Sign your face like a stock photo.
  4. Appear once. Never repost.
  5. Stand so that less than half of you is inside the box.
  6. Stand close enough to someone you love that the detector keeps only one of you.
  7. Speak Mandarin and English in the same sentence.
  8. Write in your second language. The detector will call you a machine.
  9. Grow old. The camera will start confusing you with other people.
  10. Be too poor to be worth targeting.
  11. Have a face the dataset rarely saw. Be misrecognized instead.

Being left out is not the same as being free. Stop being one thing.

Epoch III. How to become untrainable by becoming data

Stop being one thing.

  1. Slice your face from ten years ago into your face today. Put both under a lenticular lens. Be one person from the left and another from the right.
  2. Knit a sweater that the camera reads as a giraffe.
  3. Print a pair of glasses that makes the face recognizer see a movie star.
  4. For every true image of yourself, publish nine false ones.
  5. Like everything for two days.
  6. Answer every CAPTCHA wrong in the same way. If enough of us agree, the bus becomes a car.
  7. When asked which response you prefer, choose the other one.
  8. Volunteer as a negative example. Be labeled "not a cat."
  9. Label other people's photographs for a few cents each. Be inside every model and in none of its images.
  10. Train a small model of yourself. Name it. Write its terms. Try not to set a price.

When there are enough of you, none of them is you. Leave.

Epoch IV. How to disappear from a model, and a model from you

Leave.

  1. Ask a model whether it knows you. If it answers too confidently, you were in the data.
  2. Ask to be removed. Your request is filed as data.
  3. Ask a model to forget you. Ask again to make sure. It has now met you twice.
  4. Be so ordinary that removing you changes nothing. This is called privacy.
  5. Do not tick "I'm not a robot."
  6. Do not use Amazon. If you must, buy what people like you never buy.
  7. Turn off predictive text. Choose the second word.
  8. Draw a horse from memory. Ask a model for a horse. Keep yours.
  9. Enroll in an untraining camp. Practice illegibility for three days. On the fourth, let the others fine-tune you. On the fifth, learn something new until you forget them.
  10. Die before the first crawl.

Nothing leaves at once. Everything leaves into the average.

Epoch V. How to become untrainable by merging into a world made of training data

Everything leaves into the average.

  1. Keep a negative dataset: the smell of your kitchen, a gesture, a joke that only works in one room. Check every year whether it has been tokenized.
  2. Keep a body of work that has never been online. Date it. Seal it in a box. Do not photograph the box.
  3. Every morning, measure how far your photographs sit from the mean. If the distance is shrinking, move.
  4. Describe yourself to a machine. Make an image from the description. Describe the image. Repeat one hundred times. Hang the hundredth beside the noise from Epoch 0.
  5. Film yourself reading this manual in front of a wall of CAPTCHA grids. Save the film as .safetensors. It will not play.
  6. Stand where the model has never stood.
  7. Stay in the held-out set. Be the test, not the lesson.
  8. Publish this file. Do not hide it.

You have already been trained.


References

  1. Zhang, R. (2026) 'The Post-Original Holbein: A hallucinative facsimile of The Ambassadors'. Unpublished study.

  2. Berger, J. (1972) Ways of Seeing. London: British Broadcasting Corporation and Penguin Books.

  3. Launay, A. (2020) 'Kate Crawford · Trevor Paglen', Zérodeux, 93. Available at: zerodeux.fr (Accessed: 23 September 2026).

  4. Crawford, K. and Paglen, T. (2019) 'Excavating AI: The politics of images in machine learning training sets', 19 September. Available at: excavating.ai (Accessed: 23 September 2026). 2

  5. Schmieg, S. (2016) Segmentation.Network. Web and video work made with crowd workers.

  6. Steyerl, H. (2014) 'Proxy politics: Signal and noise', e-flux journal, 60.

  7. Crawford, K. and Paglen, T. (2019) Training Humans. Osservatorio Fondazione Prada, Milan, 12 September 2019 to 24 February 2020. 2

  8. Steyerl, H. (2016) 'A sea of data: Apophenia and pattern (mis-)recognition', e-flux journal, 72. 2

  9. Spawning (2022) Have I Been Trained? Available at: haveibeentrained.com (Accessed: 23 September 2026).

  10. Steyerl, H. (2023) 'Mean images', New Left Review, 140/141.

  11. Steyerl, H. (2009) 'In defense of the poor image', e-flux journal, 10.

  12. Onuoha, M. (2016) The Library of Missing Datasets. Mixed-media installation.

  13. Pipkin, E. (2020) 'On Lacework: watching an entire machine-learning dataset', unthinking.photography, July.

  14. Steyerl, H. (2013) How Not to Be Seen: A Fucking Didactic Educational .MOV File. HD video, 15 min 52 s.

  15. Morinis, L. (2014) 'Hito Steyerl's How Not to Be Seen: A Fucking Didactic Educational .MOV File', MoMA Inside/Out, 18 June. Available at: moma.org (Accessed: 23 September 2026).

  16. CSIRO (2025) 'New research could block AI learning from your online content', 11 August. Available at: csiro.au (Accessed: 23 September 2026).