What Does Good Look Like?
I keep asking founders a version of the same question, and I keep getting the same silence.
The question is not complicated. How are you measuring this today? How would you know if the AI you are using got worse next week? What does good look like for the thing you are actually trying to do?
The silence is not because they have not thought about their AI. It is because they have thought about everything except this. They have thought about the model. The vendor. The rent-versus-own question, endlessly. The stack decisions. The board slide. What they have not done is define, in their own words, what good looks like for the specific thing they are trying to do.
I wrote in April about depth being the filter that lets you tell when a recommendation is wrong. This is the same argument, one step earlier. Before you can tell if the AI is wrong, you have to know what right looks like.
The first question is not the one you are being asked
The industry conversation is loud right now, and it is almost entirely about the wrong layer. Rent or own. Open weights or closed. This model or that one. Fine-tune or prompt-engineer. Every panel, every newsletter, every LinkedIn post is arguing about the choice above the one that actually matters.
None of these questions have an answer until you have defined what good looks like for your task. Every one of them is downstream. You cannot say a model is cheaper if you have not defined what performance you are pricing. You cannot say your data is a moat if you cannot describe what quality that data is supposed to produce. You cannot compare two vendors if the only benchmark you have is whichever one their marketing put on the front page.
The rent-versus-own decision is not the first decision. It is the second.
What you build when you skip it
A founder who skips this step does not end up with nothing. They end up with worse than nothing.
They accumulate a pile of AI outputs, spread across three tools and two spreadsheets and the memories of whoever was on the last call. It looks, from outside, like a data moat. It is not. It is a receipt for decisions made without a definition of good. You cannot fine-tune against it. You cannot A/B two vendors on it. You cannot notice when the model quietly gets worse, because you have no baseline of what better was.
The people who talk about data as a competitive advantage are correct. But the advantage is not in having data. The advantage is in having data paired with a definition of quality specific enough to let the data teach a model to be better. Without the definition, the pile is just a pile.
This is why the loudest voices in the room right now are the ones talking about evaluation. The evaluation is not a technical artefact. It is the founder-shaped decision made explicit. Someone had to sit down and say: here is what good looks like for what we do, in enough detail that a machine could grade it. Most founders have not done that sitting-down.
Right-size before you optimise
Most tasks do not need frontier intelligence. Once you have set a bar, most tasks turn out not to need it.
A classification task on your CRM does not need the same intelligence as a legal review. A first-draft-of-an-email tool does not need the same intelligence as a system that decides which returns to accept. Once you have defined what good looks like for a given task, you often discover that a much smaller, much cheaper model hits your bar. The right question is not what is the smartest AI I can use. It is what is the smallest AI that still hits the bar I set.
You cannot ask that question at all until you have set the bar. This is why founders who skip the definition-of-good conversation end up over-buying intelligence for tasks that did not need it, under-buying it for tasks that did, and unable to explain the difference to anyone.
The version that holds up
Speed is real. The tools are real. The AI decisions are being made whether you sit down for this conversation or not.
But every decision made without a definition of good is a decision made in the dark. Sometimes it will be right. You will not know when it was and when it was not. That is the whole problem.
Sit down. Write down, in language a colleague could argue with, what good looks like for the thing you are trying to do. If you cannot yet, that is the work. Everything else is downstream.
The AI decision is not rent or own. It is whether you know what you are measuring against.