Most people ask one AI one question, get back something clean and confident, and walk into a Monday meeting calling it research. It isn’t. One model hands you an opinion written well enough to pass for a finding. And the thing a single model does best is agree with you, which is the last quality you want in the room when you’re deciding where to point a build budget.

A tasting where everyone nods is a wasted tasting

I’ve sat in menu tastings where the chef presents, the room nods, the dish goes on the menu and dies by week three. The tastings that actually saved money were the ugly ones, where the pastry chef and the exec sous went at each other over whether the dish held eight minutes under a lamp on a 900-cover night. Nobody enjoyed those rooms. They were the only ones that told me anything before the money was spent.

So I built the argument into the software

Verdikt is a tool I built in my studio, and inside it is something I call the Harness. Several frontier models get handed the same question. Then they red-team each other’s answers, and a synthesis arbiter names a winner and the single best answer on the table. Anthropic, OpenAI and Perplexity all run in it live, right now.

I read the fight, not just the verdict. The arbiter is a foreman, not a judge. I’m the judge, because I’m the one who has to look a client in the eye and say this concept is worth eleven weeks and a fabrication budget.

The disagreement is the signal

Three models are no smarter than one model. What they hand me is a map of where they disagree, and that map is the whole reason I run the thing. Where they converge, I move, and I stop researching. That’s a settled question, and spending another afternoon on it is procrastination wearing a lab coat.

Where they split is the real find. A split means I’ve located the actual open question inside my own concept, and that one earns a human afternoon, a phone call to somebody who has run the format on a live floor, and a rehearsal with real product on a real pass. Consensus tells me the idea is legible. Disagreement tells me where the surprise is still unbuilt, and the surprise is the entire product.

Certain goes to code, uncertain goes to a model

This is the rule everything in my studio runs on. A model may retrieve, extract, and distill. It may never decide, and it may never kill. If I can write a condition down as a number, it belongs in code, because a deterministic rule can be audited six weeks later by somebody who wasn’t in the room, and a model’s mood on a Tuesday cannot.

Killing belongs to rules and to taste. Both of those are mine to defend. Neither one should be handed off to a well-worded paragraph.

Kill it early, kill it cheap

Alongside the Harness I run a Kill Test that is deliberately biased toward death. It fires twice. The first pass is desk research: is the job real and underserved, is the window rising rather than closing, and is there an identifiable buyer with a budget line already open. The second pass is behavior: did real people take a real costly action, a deposit or an opt-in, above a threshold I set before the test ran.

That last part is the whole trick. The threshold gets written down in advance, because a number chosen after you see the result is just a feeling with a decimal point. Any single knock-out No kills the concept, or recycles it into a different one.

Biasing toward death sounds harsh for a company built on wonder. It’s what pays for the wonder. Making a kill cheap and early is what protects the budget for the three or four concepts that deserve the full build, the prototype, the rehearsal, and the unveiling. You get more first-of-its-kind moments by killing faster, not by being precious with the ones you started with.

What search volume will never tell you

Here’s the limit, and I’d rather say it than have you find it. A first-to-market moment has near-zero search volume by definition. If I’d asked a keyword tool in advance whether anyone was searching for a dish nobody had served yet, the answer would be zero, and zero would be both accurate and useless.

So demand data does a narrower job in my process. It maps adjacent demand, it shows me the gaps my competitors left open and what’s rising next door, and it never gets a veto over something genuinely new. Any process that lets a search volume report kill an invention will only ever produce things that already exist.

Try this on Monday

Take the last AI answer you were about to carry into a meeting and run the identical question past two more models before you present it. Read all three side by side. Bring the disagreement into the room instead of the consensus, put the split on the whiteboard, and ask your team which one of the three is wrong and why. That argument is your research. The clean single answer was just a well-written guess.