Design Arena creators secure $7.9 million to enhance AI models with flavor

Design Arena creators secure $7.9 million to enhance AI models with flavor

As described by co-founder Grace Li, her venture began weeks prior to their 2025 graduation, as a group of college friends aimed to make their AI game engine functional. The models were capable of producing operational games, yet none were enjoyable — posing the intriguing question of how to determine if a game is entertaining.

They concluded that there was no alternative to human assessment, leading them to generate ideas for gathering genuine human feedback on a larger scale. This effort culminated in the creation of Design Arena, an AI platform now utilized by 5.3 million users globally. It turned out that numerous AI enterprises were seeking scalable user insights — and many were prepared to invest in it.

“It represented the crucial bottleneck for many of these models to enhance their design capabilities,” Li states. “Approximately a week later, we secured our initial significant agreement with a frontier lab, and the subsequent events unfolded as history.”

On Monday, the entity behind Design Arena — titled Intelligence — revealed a $7.9 million seed funding round spearheaded by Index Ventures, alongside contributions from Conviction (Sarah Guo and Mike Vernal), A*, Valkyrie, and others.

For non-enterprise users, engaging with Design Arena resembles using an advanced model router. There’s a ChatGPT-esque interface for entering prompts, with distinct dropdown menus for websites, images, and various other visual formats. Upon submitting your request, format, and style, you’ll receive a series of “A vs. B” comparisons until you’ve ranked the small selection of outputs from best to worst.

While it’s a beneficial feature, the real advantage of the platform lies in the enterprise aspect, where participating models can utilize it as a continuous source of instantaneous feedback for their media-generating models. The users typically don’t care which models they’re evaluating — as Li notes, they are simply after the best results — hence their rankings can provide critical insights into user desires.

For frontier labs, that’s a service worth financing, Li remarks, mentioning that the site is currently achieving $60 million in ARR, strengthening its role as a vital provider of human-centric evaluation data for the AI sector.

Importantly, users must log in to access their output, allowing Intelligence to monitor how preferences evolve across various regions and over time. (Li observes that web dashboards in Asia often display a more maximalist design aesthetics.) These metrics are a crucial complement to automated benchmarks, which can function on a larger scale yet are frequently vulnerable to manipulation, as evidenced by the Hugging Face breach that dramatically highlighted these issues last week.

However, it is essential to note that crowdsourced human feedback might not guarantee market success. Less than a year after its debut, Yupp closed down earlier this year after securing $33 million from a16z crypto’s Chris Dixon. It also acquired some frontier models as clients and claimed to have over 1.3 million users, yet failed to establish a viable long-term business.

Nevertheless, other startups focused on human evaluation appear to be flourishing. LM Arena, which adopts a similar method for text-based feedback, raised $150 million in a Series A in January, merely four months after officially launching its paid service.

When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.

Leave a Reply