How can creative agencies use AI without the work becoming generic? Start by taking the threat seriously, because it is measured, not imagined: in controlled studies, writers using AI produced individually better work that was collectively more alike. Better prompting will not fix that. Moving your agency's own context into the system the AI works inside will.
The sameness is measured
In a Science Advances experiment published in 2024, Anil Doshi and Oliver Hauser had writers produce short stories, some with ideas from a large language model. Access to AI made individual stories score higher on creativity, writing quality and enjoyment, with the biggest lift going to the weakest writers. The same stories were also more similar to one another, on the study's own metrics, than stories written without help.
A second study locates the mechanism. At ICLR 2024, Vishakh Padmakumar and He He compared essays written with a base model, with an instruction-tuned model, and with no model. The base model left diversity intact. The instruction-tuned model produced a statistically significant reduction, pulling different authors toward the same phrasings and the same points. Instruction tuning is what makes a model feel helpful, and every assistant your team touches is tuned this way. The pull toward the middle is a property of the category, not a bug in one tool.
Averages in, averages out
The economics follow from who gets the lift. Doshi and Hauser found the largest gains at the bottom, which means AI raises the floor for everyone, including the eleven other agencies pitching the same account. Once the floor rises everywhere at once, floor-level work stops winning pitches. A blank prompt box knows nothing about the brand, the brief or the last four rounds of feedback, so it answers from the middle of its distribution, and the middle is now the one thing every agency owns equally.
This is also why the consultant-built prompt library ages so badly. A prompt is portable: the person who wrote yours can sell its twin to the agency across town, and does. Context cannot be resold, because it is made of things only you have: your client history, your strategy documents, your feedback, the work you refused to ship.
Context is the input the studies were missing
Both experiments gave every writer the same starting point, which is exactly what a blank chat gives your team. The variable an agency controls is what the model reads before it writes. In Multiply that is a hierarchy set once rather than pasted per prompt: an organisation profile and company instructions that apply to everything, a workspace per client holding brand guidelines, strategy and past work, and projects that inherit both. Two teammates briefing the same client start from the same memory, and the output starts from your position instead of the average.
A test you can run this week
Take a live brief. Run it through a blank chat window, then through whatever context system you have, and put both outputs in front of someone who knows the client. If they cannot tell which one came from your agency, the model is not your problem, and no model upgrade will fix it. Judgement stays where it always was: AI proposes inside the constraints you set, and the team decides what ships.

