Blog · Design

Why every AI-generated app looks the same, and the token fix

Why language models converge on one look, the familiar font and the purple gradient, and how named design languages and shared tokens break the pattern.

By DraftScreens · · 5 min read

Ask five AI tools for a fitness app and you get five versions of the same screen: a white or near-black background, a familiar sans-serif, a purple-to-blue gradient on the main button, rounded cards floating on soft blobs. People call it AI slop. It is less a taste problem than a statistics problem, and it has a fix.

Why models converge

A language model writing an interface predicts the most likely next piece. The most likely interface in its training data is the kind that tutorials, templates and starter kits are made of: a component library with default styling. A model left alone reproduces the centre of that distribution, and the centre is the same for everyone who asks.

It gets worse across screens. Each screen is a fresh generation, so even inside one app the palette shifts, the corner radius changes and the tab bar gains or loses an item. In effect the app is sampled rather than designed.

Why it matters beyond looks

An app that looks like many other apps gives a user no reason to remember it. It also invites a harder look in store review: Apple's guidelines treat apps that are near-copies of each other, or built from a template with little changed, as spam.

The fix is constraints, not adjectives

Telling a model "make it unique" moves it a little. Telling it exactly which language to commit to moves it all the way. Three constraints do most of the work.

Named design languages. A style is a paragraph of concrete commitments. For example, Brutalist: mono type, hard shadows, zero radius. Or Luxury: serif headlines, gold on near-black, sharp corners. Each one rules out the default look in specific terms. DraftScreens has ten of these built in; design languages explained lists them.

Shared tokens across the flow. Plan the flow first and write the palette, fonts, radius and spacing down as data. Then generate every screen with the identical token block in its instructions, told to use exactly those colours and nothing else. Drift stops because there is nothing to drift from.

Your brand as the source. If you have a brand, the tokens should not be invented at all. A design system, with palette roles, body and display fonts, a voice, do and don't rules and a logo, replaces the planner's guess. Every flow, screen and edit reads it.

What it looks like in practice

Take two prompts for the same habit tracker. The first is the bare prompt, and the model returns the default. The second adds a design language, say Luxury, and the token block from the planned flow. The model is the same in both cases. What changed is how much it was told, and how much it was told not to do.

Try it

Write your prompt, pick a design language instead of leaving it open, and generate the same screen in two different languages. How DraftScreens works shows where the tokens are planned and applied, and the home page has the prompt box.

Keep reading