Bryan Korba on AI: “Being right more often than a human is not the qualifying condition for authority”

The founder of Popcorn Labs talks to Nuvastra about AI’s limits, the pressures that distort business decisions and why giving machines more authority means asking what they can be held to account for.

Bryan P. Korba has spent much of his career investing in businesses and working inside them. Based in Dallas, the founder and chief executive of Popcorn Labs has been principal of the single-family office JDFIU Capital since 2001. His background spans commercial real estate, healthcare, software, furniture and distribution, with operating roles alongside his investments. It is that experience of how organisations behave under pressure that informs his approach to artificial intelligence.

His current work centres on Equilibrium, a platform designed to measure the distance between an organisation’s stated standards and what is happening in practice. On his website, Korba describes an architecture intended to make its outputs traceable to their inputs and reproducible. In this interview, he makes a distinction that matters well beyond his own business: a system that gives the same answer every time can still be wrong, but he argues that reproducibility makes its mistakes easier to investigate and correct.

That concern with accountability runs through a conversation that moves from insurance and boardrooms to employment, education and the intimacy of talking to a machine. Korba is forthright about the AI industry’s claims and equally concerned with the incentives inside the institutions adopting its products. We began by asking what had surprised him most about AI over the past year. His answer was about the people expected to insure its failures.

1. What has surprised you most about AI over the past year?

The insurers. Not a technical development, a market one. In 2026, AIG, W.R. Berkley and Great American filed with regulators to exclude artificial intelligence risk from corporate policies, and Berkshire Hathaway and Chubb had already obtained approval to do the same. W.R. Berkley's proposed language reaches any claim arising from any actual or alleged use of AI.

That is not a price increase. An insurer that thinks a risk is expensive raises the premium and writes the business. An insurer that files to exclude a risk is saying it cannot find a number at which writing it is rational. The industry whose entire profession is pricing uncertainty looked at this one and declined.

It surprised me that almost nobody covered it. Equity analysts are still debating multiples. The underwriters already left the table, quietly, in state regulatory filings, in the language of endorsements.

2. You argue that making AI models bigger won't solve some of their fundamental problems. What is the clearest example?

Hallucination, and the argument against scale fixing it is older than the field.

Take any enumerable collection of computable models. You can construct a computable ground truth that every single one of them gets wrong, by the same diagonal argument Cantor used to show some infinities are larger than others. Ziwei Xu and colleagues did exactly this in 2024. Dataset size does not appear anywhere in that proof, because it is a statement about cardinality rather than about sample complexity. You cannot train your way past it.

Then researchers at OpenAI bounded it from the other side. Kalai, Nachum, Vempala and Zhang proved that generative error rates are bounded below by the model's own misclassification rate on whether a statement is valid. Generation is provably harder than classification. A system cannot reliably produce what it cannot reliably recognise.

Scale improves the number. It never changes the kind. And the risks increase. It’s math.

3. A system can give the same answer every time and still be wrong. What does making it deterministic actually solve?

This is the right question and the answer is not correctness.

Determinism buys accountability. A wrong answer that reproduces can be found, traced to the inputs that produced it, attributed to a specific step, and corrected once for every future case. A wrong answer that does not reproduce cannot be investigated at all. You cannot debug what you cannot reproduce, and you cannot litigate it either. Reproducibility is what turns an error from a dispute into a fact.

The second thing it solves is drift detection. Measuring whether something has moved requires a reference that has not. If your instrument changes between readings, you cannot distinguish a change in the thing from a change in the instrument. A meter that varies with the weather is not a meter.

So determinism is not a claim to be right. It is the precondition for being correctable, measurable, governable, and for being able to say anything about a specific case rather than about a population of cases.

4. What have you seen in business that changed your mind about how good decisions get made?

That the best decisions did not come from the smartest people in the room.

Twenty-five years of watching this, across several institutional contexts, and the pattern held: capability is roughly constant in a given person, and access to that capability varies enormously with conditions. The same executive who is clear-eyed in March is captured in October, and nothing about their intelligence, experience or integrity changed. What changed was the pressure and their distance from the standard they said they were operating to.

So I stopped asking whether we had the right people. The more useful question is whether our people are in a state where their capability is available. That question has the advantage of being measurable, and the first one never was.

5. How would you judge whether a board was doing a good job without relying on what it says about itself?

Watch what gets rewarded, tolerated and punished under pressure. The charter is the declared standard. The reinforced behavior is the operating standard, and the second one is what actually governs.

Three specific things I would want. What happened to the last person who brought the board bad news, and where are they now. How long it took the board to correct after a deviation was identified, which is a rate and can be compared. And whether anyone can point to a decision the board reversed on evidence, because a body that has never reversed itself is either infallible or not listening.

Note what is missing from that list. Composition, credentials, meeting frequency, committee structure. All of it is the declared standard, and all of it is what governance assessment currently measures.

6. What do you lose when you reduce a whole company to one risk score?

Direction, and attribution. Both of which are the parts you could act on.

A score is a position at a moment. Two companies sitting at the same number, one improving and one deteriorating, are not the same risk and should not be priced or governed the same way. Trajectory tells you which is which and the score cannot.

And aggregation destroys attribution. A single number tells you something is wrong and not which force is producing it. Internal pressure, external conditions and distance from the declared standard are different problems with different remedies, and if you have collapsed them into one figure you know you have a problem and have no idea what to do about it. Worse, sequencing matters. If environmental pressure is not reduced first, internal work does not take hold. A score cannot tell you that.

7. Suppose an airline has passed its safety audit. What would you want to know about it six months later?

Whether the standard it passed against is still the standard it is operating to.

An audit is a snapshot against a declared baseline. Six months later I want the direction and the rate. Has the distance from that baseline grown or closed, how fast, and is the rate itself accelerating. I want to know what the tolerance band was and whether anything has been quietly widened.

And I want the reported deviation count, with a specific instruction: be suspicious of zero. An organisation reporting no deviations is either exceptional or suppressing them, and those two look identical in a compliance file. The way to tell them apart is to ask what happened to the last several people who filed a report. If the answer is uncomfortable, the zero is not good news.

8. How can a company spot trouble several layers down its supply chain, beyond the suppliers it deals with directly?

Not by asking, which is what most programmes do. Direct suppliers report on themselves, and a tier-three problem reaches you through a tier-one that has every reason to absorb it quietly.

Two things work. First, treat your own observable signals as instruments rather than noise. Lead-time variance, quality drift, substitution requests, requests to change payment terms, changes in who answers your emails. Upstream stress has a fingerprint downstream and it usually shows up in those before it shows up in a disclosure.

Second, and more important, check whether your diversification is real. If your three independent tier-one suppliers all depend on the same tier-three, you have one supplier and three invoices. Correlated exposure through a shared dependency is the thing that actually takes companies down, and it is invisible on a supplier map that only goes one layer deep.

9. How would insurance change if insurers could see how a business was behaving throughout the year?

It would stop being a bet on a snapshot and become what the rest of insurance already is, which is continuous measurement. Telematics did this to motor insurance and nobody now regards it as exotic.

The specific change is that pricing would move from position to trajectory. Today a company that finds and fixes its problems reports more incidents than one that suppresses them, and the reporting company is penalised for it. That is a measurement system selecting for concealment. Price the rate of correction instead and the incentive inverts.

The harder consequence is that it would reopen a market that has closed. AI risk became uninsurable because correlated exposure removes the benefit of pooling, and the condition for it becoming insurable again is not more premium. It is uncertainty turned into a contractual object that is bounded, documented, monitored and limited. Every one of those four words is a measurement requirement.

10. When does waiting for more information become a bad decision?

When the cost of the delay exceeds the value of what you would learn, and you can usually tell that moment has arrived because the new information stops changing the decision. If the last three data points all pointed the same way, the fourth is not being gathered for the decision. It is being gathered for the record.

The thing I would add is that waiting is itself a decision and it should be recorded as one, with a name against it. Most institutional failure I have seen was not a bad call. It was an unmade call that no individual owned, which is why the post-mortem cannot find anyone responsible. The delay was the decision and nobody wrote it down.

11. What should an AI system be allowed to change on its own once it has been put to work?

Operational parameters inside a declared band, yes. Its own calibration, never, without a governed approval step that leaves a record.

The distinction matters more than it sounds. A system that adjusts how hard it works within stated limits is doing its job. A system that adjusts how it scores is rewriting the standard it is measured against, and once it has done that you can no longer trust any output it produced before the change or determine when the change occurred.

The property worth protecting is that any output, at any point in its history, can be replayed and trusted. A system that silently recalibrates destroys that property permanently and retroactively. An engine that calibrates itself invisibly is a contradiction in terms, and the approval step is the only thing keeping the word deterministic meaning anything.

12. When might better information threaten the people running a company?

Exactly in proportion to how much of their authority rests on the gap between what is declared and what is happening.

That is not a cynical claim about executives. It is structural. If your position depends on a judgment nobody can check, then an instrument that checks it removes something you had. Most people in that situation do not experience it as a threat to their interests. They experience it as the instrument being wrong, or crude, or missing important context, and they may be sincere about that.

Which is why the objection is diagnostic. When measurement is proposed, watch who objects and on what grounds. Objections about methodology are usually real and worth answering. Objections about the very idea of measuring tend to come from the part of the organisation that has the most to lose from being measured, and they arrive dressed as sophistication.

13. What would justify slowing down AI development?

I would not frame it as slowing down, because that is an argument about speed and the problem is not speed.

Global spending on AI safety runs between four hundred and six hundred million dollars a year against roughly seven hundred and twenty-five billion in capital expenditure from four companies. That is about one dollar of proof for every four to six thousand dollars of capability. Correcting the ratio does not require slowing anything.

The reason it matters is not moral. The cost of building capability grows polynomially. The cost of proving anything about the result is NP-hard or worse, and for some properties it is undecidable. So the two costs sit in different complexity classes and the gap between them widens as models grow, by construction, regardless of anyone's intent. Building more produces a larger ungoverned surface. That is arithmetic, not advocacy.

What would justify genuinely stopping: deployment into a domain where the consequence is irreversible and no instance-level account of a decision can be produced. That describes several current deployments.

14. Who do you expect to hold the most power in the AI industry five years from now?

Probably not the model builders, and I would hold that loosely.

A model is a lossy compression of the data it was trained on. It retains whatever statistical pattern the training process extracted and nothing else. A governed operational record retains every observation, every score, every deviation and every attribution, with the evidence intact. Past a certain depth the model's approximation stops adding information the record did not already have, and governed data cannot be scraped, purchased or synthesised. It accumulates through operation or it does not exist. Whoever holds that has something no amount of capital reproduces quickly.

The answer nobody gives is the carriers and the reinsurers. They have more control over where this technology gets deployed than any regulator, because they decide what risk can be transferred, and a system whose failures cannot be transferred does not get deployed in a regulated industry however well it performs.

15. How should an AI adviser make money without compromising the advice it gives you?

By never being paid by the thing it assesses. That is the audit problem and it is structural rather than ethical.

An audit firm is hired and paid by the company it is supposed to independently assess, and every incentive in that arrangement points toward findings the client finds acceptable. The independence is real on the engagement letter and nowhere else. Nobody has to be dishonest for it to fail. The arrangement is what fails.

So the rule is to charge for the measurement and not for the verdict, and to hold no financial interest in the entities being measured. If the fee varies with the answer, you have built a system that will eventually produce the answer that pays. And it should apply its own measurement to itself and publish the result, because an instrument that only ever reports on other people's drift is not an instrument. It is a one-way mirror.

16. If every business can use the same AI, where does competitive advantage come from?

From what you can prove, and from how fast you find your own problems.

If every competitor runs the same models on the same substrates, capability is table stakes and cannot be a differentiator by definition. What remains is the accumulated record of your own operation, which nobody else can acquire because it only exists as a byproduct of you operating. That is a moat that compounds and cannot be bought.

The second source is less obvious and probably larger. An organisation where surfacing a problem is cheaper than concealing it receives its signals early, corrects early, and pays early-correction costs. An organisation where concealment is rational receives nothing, corrects late, and pays late-failure costs. Those differ by orders of magnitude, not by percentages, and the gap compounds every cycle. That is available to any company today without any technology at all, and almost nobody does it, because it requires not punishing the messenger and that turns out to be very hard.

17. What happens to people who do their jobs well but can no longer compete with AI on cost?

I do not want to give a comfortable answer to this, because the comfortable answers have not survived contact with any previous transition.

The part I am confident about: cost competition against a machine is not winnable in work that is purely generative. If the output is text, image or code that nobody has to answer for personally, that work is being repriced now and no amount of doing it well changes the price.

The part that is genuinely different this time, and I think underdiscussed: the requirement for someone to be accountable for a specific decision does not go away, and it cannot be automated, because the systems cannot produce a statement about a particular case. Every regulated domain needs a human who can be asked why this happened and answer. That category is growing rather than shrinking.

But I am not going to pretend that maps cleanly onto the people being displaced, or that the timing works out for them. It often will not. Saying the market sorts it out is the sort of claim people make when the cost is not landing on them.

18. What would you tell an eighteen-year-old choosing what to study now?

Choose something with a standard. A field where you can be measured against something external, be wrong in a way that is unambiguous, and learn what correction feels like.

Engineering, medicine, accounting, law, trades, the physical sciences. The discipline of being demonstrably wrong and then fixing it is the thing that transfers, and it is very hard to acquire later.

Three specific capabilities I would build. Learn to measure things, which means statistics done properly rather than as a requirement to pass. Learn to write clearly, because the ability to state precisely what you mean is now scarcer than it was, not more common. And learn one domain deeply enough to know when a number is lying to you, because you are going to be handed a great many confident numbers.

What I would avoid is any field whose entire value is producing plausible text. That is the one thing the machines do cheaply and well.

19. What do you make of the idea of an AI becoming someone's closest confidant?

I would be careful, and not for the reason people usually give.

The objection is not that it is artificial. It is that a confidant who has no stake cannot hold you to anything. The value of a real one is that they can tell you something you do not want to hear and then bear the cost of the friendship changing, and they do it anyway. That cost is what makes the honesty mean something. A system that risks nothing by telling you the truth also risks nothing by telling you what you want, and you have no way to tell which you are getting.

These things can be genuinely useful for thinking. I use one for exactly that. But it should not be the only voice, and if someone's closest relationship is with a system, the right response is to ask about that gently rather than to sell them a better subscription. The people building these products know which way the engagement metrics point, and I do not think that question is being asked carefully enough by anyone with a commercial interest in the answer.

20. How much authority should we give a machine that consistently makes better decisions than we do?

Consistently better is a statement about a population. Authority gets exercised in individual cases. Those are different logical types and the whole question turns on the difference.

So the answer is not a percentage of decisions. It is a structure. Where the system can produce a reproducible, attributable, replayable account of how it reached this conclusion from these inputs, give it a great deal of authority, and more than most organisations currently give their own people. Where it can only tell you it is usually right, give it none, however good usually is.

Being right more often than a human is not the qualifying condition for authority. Being answerable is. We do not grant authority to people on their average accuracy either. We grant it on the expectation that they can be called to account for a particular act, and we remove it when they cannot.

21. What problem would you most like to see AI solve in your lifetime?

I would like to retire a sentence.

Every major institutional failure contains the same line in its post-mortem: people inside the organisation knew. The engineers knew about the O-rings. The risk managers knew the exposure models were wrong. The physicians saw the pattern. They knew, they did not speak, and the failure landed.

That is not a knowledge problem and it never was. The information was present, available and unacted upon, because in an environment that punishes disclosure, silence is the rational choice and the consequences land on the institution rather than the silent individual. I would like to see a measurement architecture that makes the drift visible while it is still cheap, so the question stops being whether anyone will be brave enough to say it.

Not because people should not be brave. Because a system that requires courage to function correctly is badly designed.

22. Which part of your own thinking about AI are you least certain of?

Not whether this is being done deliberately. I have stopped being uncertain about that. What I cannot tell you is how far it has already gone, and who has used it.

Start with what is documented. Stanford scores the major developers on disclosure, and the average fell from 58 out of 100 in 2024 to 40 in 2025. Read that carefully. Transparency did not fail to improve. It was withdrawn. The worst categories are training data, training compute, and post-deployment usage and impact: what went in, what it cost, and what it is doing now. Of 95 notable models released in 2025, 80 shipped with no training code. Reported parameter counts from OpenAI, Anthropic and Google have simply stopped.

Nobody forgot how to count parameters. That is a decision, made in a room, by people who understood exactly what they were choosing to stop telling the public.

And it happened in the same year the law moved the other way. The European Union made training-content disclosure an obligation under Article 53 of the AI Act, published a mandatory template in July 2025, and brought it into force that August. Voluntary disclosure fell by a third against a legal requirement to increase it. That is not inertia and it is not a resourcing problem. It is a posture.

So ask the only question that matters about any opaque system: who benefits from the opacity. Not the public, which cannot know what trained the thing answering its questions. Not the regulator, which cannot inspect it. Not the courts, which cannot reproduce it. There is exactly one beneficiary of a control surface nobody can see, and it is whoever is holding it.

I do not accept that this is an accident. An industry that withdrew its own disclosures while multiplying its reach over what hundreds of millions of people read, believe and repeat has told you what it is doing. You do not need a leaked memo. The filings are the memo.

And then there is the word. They call it intelligent. That word is doing specific work. It makes the thing sound like a mind, and a mind cannot be audited, only consulted and trusted. What it actually is, is a control surface over information at population scale, and a control surface can be audited, which is precisely why the vocabulary matters to them. Calling it intelligence converts a question about control into a question about capability and buries it. That is not marketing sloppiness. It is the most effective piece of positioning in the history of the technology industry.

So what am I least certain of. How deep it already runs. Whether narrative shaping is currently deliberate policy at any of these companies or merely available on demand. Whether it has been used, by whom, on what, and whether anybody will ever be able to prove it. I do not know, and the reason I do not know is that the systems were built so that nobody could find out. That is not an unfortunate side effect of competitive pressure. It is the design working as intended.

Without governance of the instrument itself, every harm people worry about is available to whoever holds the controls and unprovable by everyone else. That is not a risk to our productivity or our jobs. It is a risk to whether a population can establish what is true, and a population that cannot do that has lost the thing every other freedom depends on.

23. Where do you draw the line in your own use of AI?

Three rules, and I hold to them.

It does not make a claim under my name that I have not verified. If it produces a number, a citation or a fact, I check the source before it goes out, and I have caught it inventing plausible ones. That is not a criticism of the tool. It is what the tool does.

It does not get the last word on a fact. It is an instrument for finding and drafting, and the judgment stays with me, because I am the one who can be asked why and has to answer.

And I do not use it for judgment about people. Hiring, firing, evaluation, anything where a person bears a consequence. Not because it would be bad at it, but because a person subject to a decision is owed someone who can be held to account for it, and a system cannot be.

24. Finally, give us your honest prediction: is AI going to save us, kill us, or make the future much stranger than either of those?

I will not take the first two, and not out of diplomacy. Both are prophecies, and prophecy is what you produce when you cannot measure. An industry that cannot tell you what its system will do with this input on this occasion has no standing to tell you what it will do to civilization, and neither do its critics. Everyone shouting about salvation and extinction is working from the same absence of instrumentation.

Stranger, then, but let me say in which direction, because stranger on its own is a dodge.

I think we are heading for a world in which the question can you prove that becomes answerable in places it has never been answerable. Not because machines got smarter, but because measurement became continuous and cheap and hard to argue with. That changes technology less than people expect and changes institutions far more, because every institution in human history has run partly on the gap between what it declared and what it did. Some of that gap is corruption. Most of it is ordinary drift that nobody measured and everybody tolerated.

Close that gap and a great deal becomes visible at once. Not all of it flattering, including in places that have been comfortable for a long time. That is the strange part, and it has nothing to do with the machines.

Next
Next

Claude Enters Financial Advice. The Real Prize Is the Workflow.