We gave the test to the machines
The test has been run through AI at least half a dozen times. Each run asked the model to complete the same number of responses the human panel produced, anywhere from 20 to 600 depending on the test. Same prompt, same images, same task. The only variable was whether a person or a model was responding.
The point was not to embarrass the AI. It was to find out whether a machine can reproduce what a human audience does when it reacts to a concept. If it can, perception testing is something you could simulate. If it cannot, that tells you something about what the test is actually capturing.
“Random dots” was last year’s story
The original soundbite was that AI produced random dots. We asked a human audience what “American” looks like, then asked four AI models the same thing. The human audience showed a clear collective preference, with statistically significant groupings. The models showed no clustering and no pattern. A year ago that was literally true. Everything came back evenly split, which on a perception map looks like a complete wash, no signal anywhere.
That is not where things stand now, and the easy version of the story is already out of date.
AI can cluster now. It has gotten meaningfully better at taking the test over the past year, and it produces semi-intelligent answers instead of an even smear. On tests with a fairly predictable right answer, it does okay. Not fantastic. Okay.
Figure 1 · Same prompt, three respondents
The trouble shows up exactly where branding lives: the nuanced, abstract prompts, the kind of question a real brand exercise asks. On those, the model is wildly off from the human data. And it is off while sounding completely sure of itself.
The compelling-rationale problem
This is the part that should make any creative director uneasy.
When AI is wildly off, it is wildly off with a very compelling rationale. It explains its decision, describes the pattern it claims to see, and lays out reasoning that would be genuinely convincing if the human data were not sitting next to it for comparison.
A year ago, a model that could not decide was at least honestly useless. You looked at the even split and knew you had nothing. Today’s models are more precarious, not less, because they hand you a confident answer with a story attached. Take that answer at face value, skip the human panel, and you have a synthetic audience telling you a beautifully argued thing that real people do not feel.
Figure 2 · The failure mode got harder to see
Then the campaign underperforms, and you are left guessing why it did not land, holding a rationale that sounded airtight. The risk was never that AI could not generate an answer. The risk is that it generates a persuasive one, disconnected from the people you are trying to reach.
There are full comparison reports behind this: human data run against Claude, ChatGPT and Gemini on the same tests. The pattern holds across all three. The machines can produce a confident reading. They cannot reliably produce the human one.
Why this matters to a creative director
If you take one thing from this, take this: AI being able to cluster does not make it safe to substitute for your audience. It makes it more dangerous, because the failure mode got harder to spot.
A creative director’s job is to make work that moves a specific group of people. The only way to know whether it does is to ask those people. A synthetic audience can now imitate the format of that answer well enough to fool you. That is not a reason to use it. It is a reason to be more careful, and to keep a human panel as the thing you check the machine against, not the other way around.
A note on information theory, since it explains the whole thing
Claude Shannon’s work is where this gets interesting, and it is worth being careful with, because it gets misrepresented constantly.
Shannon’s basic idea is that communication is a balance between redundancy and entropy: the expected and the unknowable. Neither is good or bad. Both are necessary. Redundancy is the familiar part that gives you the context to understand something. Entropy is the unexpected part that carries new information. Pure redundancy is the same thing over and over, and it gets rejected; it does not even qualify as communication. Pure information, with no redundancy at all, is unintelligible noise. Communication is the balance of the two.
Apply that to a visual and you get a concrete, usable idea: the entropy of how an audience responds. If everyone reads an image roughly the same way, the response is low-entropy, semantically narrow. The image communicates decisively. If responses are all over the place, the image is high-entropy, genuinely ambiguous. “This visual is unclear” stops being a hunch and becomes something you can put a number on. You could show a client that their hero image carries more perceptual ambiguity than a competitor’s, which means the competitor is communicating more decisively. That is the kind of number that lands.
Figure 3 · Ambiguity as a number
Here is the part the gurus and motivational speakers always skip, and it was a genuine pet peeve of Shannon’s: his theory deliberately had nothing to do with meaning. He said outright that the meaning of a message was irrelevant to the engineering problem he was solving. He cared whether a transmission was received, not what it meant.
Which is exactly why this is interesting rather than a misappropriation. Constellations measures meaning to an audience, the thing Shannon set aside. He was not wrong to set it aside; it was not his problem to solve. But the tools he built for measuring the spread of a signal turn out to give you a real, calibrated way to measure how clearly a piece of creative communicates. Shannon would probably have found that worth a look. He just would have insisted you not pretend he had already done it.
Frequently asked questions
Can AI take a perception test now?
Yes, in the sense that it produces an answer. A year ago it could not decide and every result was an even split. Today it clusters and gives semi-intelligent answers. On simple prompts it does okay. On the abstract, nuanced prompts branding depends on, it is wildly off from the human data.
If AI can cluster, why not use it instead of a human panel?
Because the failure mode got harder to catch. AI now produces a confident, compelling rationale for an answer that does not match what real people feel. Take it at face value and you risk a campaign that underperforms for reasons your synthetic audience never warned you about, while the rationale still sounds airtight.
What does this prove about the test itself?
That it captures something specific to human perception: the actual response of real people to a concept. A model can imitate the form of that response, a cluster with a story attached, but it cannot reliably reproduce the substance. Human data run against Claude, ChatGPT and Gemini shows the same pattern across all three.
How does Shannon’s information theory connect to this?
Shannon gives you a way to quantify how spread out an audience’s responses are, which translates directly to how clearly a visual communicates. His theory deliberately ignored meaning, and that is exactly the gap a perception test fills: it measures what an image means to an audience, not just whether the signal got through.