What 500 People See · Proof

Show AI 100 Brand Directions

A year ago the machines could not decide. Now they decide with confidence, and on the prompts branding depends on, they decide wrong.

Quick answer

Run a Constellations perception test through an AI model and ask it to answer the way a human panel would, and you get a specific result. A year ago the models could not decide at all: every result came back an even split. Today they cluster, and on simple, predictable prompts they do okay. On the nuanced, abstract prompts branding depends on, they are wildly off from the human data. A confident answer with a story attached does not qualify as audience data, and neither does a synthetic panel that happens to agree with the room. The danger is not that AI gets it wrong. It is that it now argues its wrong answer convincingly, and a campaign built on that answer underperforms for reasons nobody warned you about.

We gave the test to the machines

The test has been run through AI at least half a dozen times. Each run asked the model to complete the same number of responses the human panel produced, anywhere from 20 to 600 depending on the test. Same prompt, same images, same task. The only variable was whether a person or a model was responding.

The point was not to embarrass the AI. It was to find out whether a machine can reproduce what a human audience does when it reacts to a concept. If it can, perception testing is something you could simulate. If it cannot, that tells you something about what the test is actually capturing.

“Random dots” was last year’s story

The original soundbite was that AI produced random dots. We asked a human audience what “American” looks like, then asked four AI models the same thing. The human audience showed a clear collective preference, with statistically significant groupings. The models showed no clustering and no pattern. A year ago that was literally true. Everything came back evenly split, which on a perception map looks like a complete wash, no signal anywhere.

That is not where things stand now, and the easy version of the story is already out of date.

AI can cluster now. It has gotten meaningfully better at taking the test over the past year, and it produces semi-intelligent answers instead of an even smear. On tests with a fairly predictable right answer, it does okay. Not fantastic. Okay.

Figure 1 · Same prompt, three respondents

Three perception maps for the same prompt: human panel, AI a year ago, AI now Three panels. The human panel shows two tight clusters of responses. The AI of a year ago shows dots spread evenly across the whole map with no grouping. The AI of today shows one tight, confident cluster sitting nowhere near the ghosted outlines of where the humans clustered. HUMAN PANEL Clear collective preference AI, A YEAR AGO An even smear. No signal. AI, NOW Clusters — in the wrong place humans humans “confident”
The middle panel was easy to dismiss and easy to see through. The right panel is the problem: it looks exactly like the left one until you overlay the human data, and if you skipped the human panel there is nothing to overlay.

The trouble shows up exactly where branding lives: the nuanced, abstract prompts, the kind of question a real brand exercise asks. On those, the model is wildly off from the human data. And it is off while sounding completely sure of itself.

The compelling-rationale problem

This is the part that should make any creative director uneasy.

When AI is wildly off, it is wildly off with a very compelling rationale. It explains its decision, describes the pattern it claims to see, and lays out reasoning that would be genuinely convincing if the human data were not sitting next to it for comparison.

A year ago, a model that could not decide was at least honestly useless. You looked at the even split and knew you had nothing. Today’s models are more precarious, not less, because they hand you a confident answer with a story attached. Take that answer at face value, skip the human panel, and you have a synthetic audience telling you a beautifully argued thing that real people do not feel.

Figure 2 · The failure mode got harder to see

How an AI perception test fails a year ago versus now Two rows of three boxes. A year ago: an even split leads to an honestly useless result, which leads to running the human panel anyway. Now: a confident cluster with a rationale reads as an answer, and leads to a campaign that underperforms while the story still sounds airtight. A YEAR AGO An even split no clustering, no pattern Honestly useless you can see you have nothing So you run the human panel the test still gets done NOW A confident cluster plus a rationale for it Reads as an answer and isn’t one The campaign underperforms and the story still sounds airtight The machine got better at the format. The failure mode got harder to spot.
The top row cost you nothing, because a wash sends you straight back to real people. The bottom row is the expensive one: it looks finished, so nobody runs the check that would have caught it.

Then the campaign underperforms, and you are left guessing why it did not land, holding a rationale that sounded airtight. The risk was never that AI could not generate an answer. The risk is that it generates a persuasive one, disconnected from the people you are trying to reach.

There are full comparison reports behind this: human data run against Claude, ChatGPT and Gemini on the same tests. The pattern holds across all three. The machines can produce a confident reading. They cannot reliably produce the human one.

Why this matters to a creative director

If you take one thing from this, take this: AI being able to cluster does not make it safe to substitute for your audience. It makes it more dangerous, because the failure mode got harder to spot.

A creative director’s job is to make work that moves a specific group of people. The only way to know whether it does is to ask those people. A synthetic audience can now imitate the format of that answer well enough to fool you. That is not a reason to use it. It is a reason to be more careful, and to keep a human panel as the thing you check the machine against, not the other way around.

A note on information theory, since it explains the whole thing

Claude Shannon’s work is where this gets interesting, and it is worth being careful with, because it gets misrepresented constantly.

Shannon’s basic idea is that communication is a balance between redundancy and entropy: the expected and the unknowable. Neither is good or bad. Both are necessary. Redundancy is the familiar part that gives you the context to understand something. Entropy is the unexpected part that carries new information. Pure redundancy is the same thing over and over, and it gets rejected; it does not even qualify as communication. Pure information, with no redundancy at all, is unintelligible noise. Communication is the balance of the two.

Apply that to a visual and you get a concrete, usable idea: the entropy of how an audience responds. If everyone reads an image roughly the same way, the response is low-entropy, semantically narrow. The image communicates decisively. If responses are all over the place, the image is high-entropy, genuinely ambiguous. “This visual is unclear” stops being a hunch and becomes something you can put a number on. You could show a client that their hero image carries more perceptual ambiguity than a competitor’s, which means the competitor is communicating more decisively. That is the kind of number that lands.

Figure 3 · Ambiguity as a number

Two response histograms: a low-entropy image and a high-entropy image Two bar charts of how an audience read an image. On the left one bar towers over the others, labelled low entropy, one reading dominates. On the right the bars are all roughly the same height, labelled high entropy, readings scatter. LOW ENTROPY what people read into the image → narrow spread One reading dominates. The image communicates decisively. HIGH ENTROPY what people read into the image → wide spread Readings scatter. The image is genuinely ambiguous. Same audience, same question, two images. The spread is a number, not a hunch.
This is what Shannon’s tools give a creative team: not whether the image was received, but how far apart the readings landed. The left image is doing its job. The right one is asking the audience to guess, and they are.

Here is the part the gurus and motivational speakers always skip, and it was a genuine pet peeve of Shannon’s: his theory deliberately had nothing to do with meaning. He said outright that the meaning of a message was irrelevant to the engineering problem he was solving. He cared whether a transmission was received, not what it meant.

Which is exactly why this is interesting rather than a misappropriation. Constellations measures meaning to an audience, the thing Shannon set aside. He was not wrong to set it aside; it was not his problem to solve. But the tools he built for measuring the spread of a signal turn out to give you a real, calibrated way to measure how clearly a piece of creative communicates. Shannon would probably have found that worth a look. He just would have insisted you not pretend he had already done it.

Frequently asked questions

Can AI take a perception test now?

Yes, in the sense that it produces an answer. A year ago it could not decide and every result was an even split. Today it clusters and gives semi-intelligent answers. On simple prompts it does okay. On the abstract, nuanced prompts branding depends on, it is wildly off from the human data.

If AI can cluster, why not use it instead of a human panel?

Because the failure mode got harder to catch. AI now produces a confident, compelling rationale for an answer that does not match what real people feel. Take it at face value and you risk a campaign that underperforms for reasons your synthetic audience never warned you about, while the rationale still sounds airtight.

What does this prove about the test itself?

That it captures something specific to human perception: the actual response of real people to a concept. A model can imitate the form of that response, a cluster with a story attached, but it cannot reliably reproduce the substance. Human data run against Claude, ChatGPT and Gemini shows the same pattern across all three.

How does Shannon’s information theory connect to this?

Shannon gives you a way to quantify how spread out an audience’s responses are, which translates directly to how clearly a visual communicates. His theory deliberately ignored meaning, and that is exactly the gap a perception test fills: it measures what an image means to an audience, not just whether the signal got through.