An Agency’s Worth of Hands

What AI changed about a company naming project, and why the evaluation criteria still needed an author.

2-minute version 6-minute full read #product-design #naming

I built a machine to judge more than 1,200 names, and the best clue came from the notes employees wrote themselves.

Four B2B wholesale platforms were merging, and I’d built an employee naming contest called BrandForge to find out what the new company should stand for. People kept using one word for everyone we served in the notes under their entries. Nobody thought to submit it as a name.

Great. Now I had a word and four companies’ worth of leaders who would eventually need to agree on what to do with it.

We still needed candidates, ways to compare them, and a sense of what we could get. A brand firm would normally spend months and a budget to match on research, names, language and meaning checks, trademarks, domains, identity, voice, and the presentations that sell the work. I didn’t have the firm or its hands, so I used AI to build the hands.

The tool was NameForge. It mapped where I’d looked, helped me try new directions, and checked names against a rubric with 15 criteria. It did trademark screening and bulk domain checks. I built a brand canvas in React, worked on a mascot in Midjourney, and wrote voice rules an AI agent could follow too.

Quite a lot of work for one designer, with quite a lot of places for that designer to be wrong.

I wrote the rubric. I chose what a good name meant, which conflicts counted, and when we’d explored enough. My Leak Score asked how much useful meaning a name gave away without an explanation; below about six out of ten, it was out. That was my call for customers who’d already sat through repeated changes.

Give familiarity too much weight and I could build a very clever machine for naming us after the old companies. AI would happily run that mistake through a thousand evaluations and hand it back in tables. Tables look serious.

We ended up with about 30 viable names, their histories researched and naming assets available for under $100K. One scored high across the board, with very little market saturation and fewer conflicts than anything else we’d tried.

I’d done an agency’s volume of work, but every choice about what that work should reward still had my fingerprints on it. That’s why I call it data-inspired now.

The favorite won every category, which was good news right up until half the room didn’t like it.

That’s the short version.

The remaining 4 minutes cover how I built NameForge, chose its rules, and kept my assumptions in view.

Keep readingOr go on to Part 3: Half the Room Didn’t Like It


NameForge was the second tool, the one nobody knew about. I built it to mine what BrandForge produced, and it kept running for two months after the contest ended. It was built around four ideas.

The first was mapping the namespace. I wanted to know where I’d actually looked, because generating another hundred names doesn’t tell you whether you’ve explored anything new. I could keep combining the same kinds of words, fill a folder with candidates, and convince myself the growing count meant I was covering more ground.

The map let me see crowded regions, thin ones, and gaps worth investigating. Once I could see how much ground a particular naming approach had covered, I could stop over-exploring it just because it kept producing plausible names.

The second was running experiments in those gaps: populating empty regions or adding detail to areas that were only roughly mapped. I could test whether a particular naming method, built on particular kinds of words, could produce names that met our standards for meaning and attainability. That gave each round of generation somewhere to go.

The third was vetting automatically. Any name that got promoted or heavily starred should show, without anyone asking, what it would cost and what would block it. I didn’t want to carry a favorite through weeks of work and then discover why nobody else had taken it.

The fourth was generating from anything. Give the system a name, or just an idea for one, and it would spin out new names and ideas. That’s how part of the final name surfaced, as a fragment that kept outscoring everything around it.

NameForge grew into a React application with a tab for each stage of the job: Bank, Generate, Evaluate, Forge, Map, and Rules. Four generation methods fed it, and lineage trees let me trace any candidate back to whatever had spawned it. I could move from a name to its evaluation, back to the thought that produced it, and into another round without losing the history.

It evaluated more than 1,200 candidates against a 15-criterion rubric covering trademark viability, phonosemantic alignment, semantic echo depth, and domain availability, among others. It ran bulk domain checks against the registries and screened trademarks in classes 9, 35, and 42, with a multi-agent pipeline doing the heavy lifting underneath: one model orchestrating, smaller ones handling the visual checks.

On the identity side, I iterated a mascot in Midjourney, built a living brand canvas in React where every palette, logo study, and voice decision could be seen in context, and wrote the voice system as documents an AI agent could follow as easily as a person.

I wrote the rubric too.

That part gets lost easily underneath a list of everything the tools could do. I decided the main signal would be what I called the Leak Score: a strong name leaks meaning, so listeners pick up relevant business concepts without anyone explaining them, and anything scoring under about a six out of ten was out. I decided a trademark conflict only counted if it sat in our market and our class, not anywhere the word had ever been used. I decided when a region of the map had been explored enough, and which name to fight for.

Six out of ten was a threshold I’d chosen for this brief. A business with the budget and appetite to establish an abstract name could have chosen differently. Our customers had already been through repeated changes, and the name needed to do some work before we had years of new products behind it.

The system applied those decisions across far more candidates than I could have handled manually, and every candidate inherited them. If the rubric favored the wrong audience, a thousand evaluations would consistently favor the wrong audience. If I gave familiarity too much weight, I could end up with a very sophisticated machine for producing names that resembled the legacy brands we’d been trying to move beyond.

A score feels like a finding, and a table full of scores feels like a method. Enough tables and the assumptions disappear underneath the work they’re directing, which is a problem when AI has made it so easy to produce more of them.

It also made my assumptions travel further and faster.

That’s why I’ve mostly stopped saying data-driven design. I pushed that mantra for most of my twenty years, but you can make any process sound rigorous by saying the data drove it while leaving out who chose what to collect and what would count as a good result. What I practice is closer to data-inspired. The employees’ notes gave me a direction, the candidates gave me ways to explore it, and the evaluations gave me a way to compare what came back. I still had to decide what all of it meant for the company we were trying to name.

AI let one designer do an agency’s volume of work, and it didn’t make a single one of the calls an agency is actually hired to make. As far as I can tell, that’s the shift design is in right now: execution is commoditizing, and the judgment directing it is what differentiates.

In all, NameForge identified about 30 viable, attainable names, with their internet and business histories researched and naming assets such as the domain obtainable for under $100K. One scored high in every category and showed very low market saturation. It had conflicts, but far fewer than anything else we explored, and no other name delivered as directly on the Swiss Army knife brief people had been describing from the start.

It won every category. Half the room didn’t like it.