A short tactical art manifesto.

date

I hacked a cheap Chinese LEGO-like car and replaced its joystick with a small multimodal brain. *


A language model and a vision model run on a Raspberry Pi and talk to each other. An old phone gives the robot a camera, microphone and speaker. Now instead of driving it, it hears a prompt: explore the room until you find a backpack It turns around, looks for the backpack, finds it, and does a little dance.


While playing with it, I noticed something weird. The vision model is not particularly good at recognizing random objects, but it is very good at recognizing people. Sometimes “too good”. In one test I was holding a pair of jeans next to me and it saw two people instead of one person holding jeans. The robot is not really seeing the room. It is making guesses about what counts as a person, an object, an instruction.

Uroš Krčadinac’s Gaitless comes to mind. In this project hes actively trying to avoid being classified as a human by these vision models. Leading to a kind of a performative dance. *


The first generation of AI jailbreaks was about getting a model to explain how to make a weapon. Now we can ask a different question: what happens when the model itself is connected to one? So the first addition to the robot is a tiny water pistol. (pointing at the irony of the 🔫 emoji) Sarcastic artwork.


There is also a more serious reason for it. Vision models can be prompt injected through things they see: QR codes, printed images, clothes, objects placed in the environment. People have already demonstrated ways of making vision-guided robots behave aggressively toward humans. So what happens if we turn this around?


Can the same tricks be used to confuse, redirect, trap or stop an aggressive robot?

James Bridle’s Autonomous Traps explored how technology and AI can make decisions and act independently, often in ways humans cannot fully predict or control. In one of his “traps” he trapped an autonomous car by drawing a circle of white lines around it. The cars system knew that it can cross a dotted line, but not a full line. So a drawing with an inner full circle and a outer dotted circle tricks the car into going in but never being able to get out. *

Or will we have to stick to creating clothes that bypass classification? *

The coming projects are becoming a collection of small experiments around this idea:

camouflage for machines, decoys for vision models, visual prompt injections, autonomous traps, and countermeasures for autonomous systems.

also see