Ask an AI model how to configure a load balancer health check, and it will tell you. Confidently. Immediately. In a tone that suggests it read the vendor documentation, tested the configuration, and is now reporting back with the authority of someone who has done this a hundred times. It hasn’t done any of that. It generated a statistically plausible answer based on patterns in its training data, and there’s no way to know, just from reading the answer, whether it’s right.
This is the trade engineers made without really noticing they made it. We used to have to go find the answer. Now the answer finds us. The problem is that “finding us” and “being correct” are two completely different things, and the tools that deliver the answer don’t distinguish between them.
The Black Box Between Question And Answer#
When you look up something on Stack Overflow, you get a chain of evidence. There’s a question, an answer, a vote count, a date, sometimes a comment thread arguing about whether the answer still applies to the current major version. You can see who wrote it, when, and whether anyone pushed back. You’re not just getting an answer. You’re getting the provenance of that answer, and provenance is what lets you decide how much to trust it.
An AI chat response gives you none of that. You get a paragraph of text that reads like it came from an expert, but there’s no visible trail connecting the claim to a source. The reasoning happened somewhere inside the model, and it isn’t coming back out. You can ask the model to “explain its reasoning,” and it will produce another paragraph of text that sounds like reasoning, but there’s no guarantee that paragraph reflects what actually happened when the model generated the original answer. It’s a plausible-sounding explanation generated after the fact, not a log file.
This is the core issue. AI models are inscrutable by design. Nobody, including the people who built them, can point to the exact chain of internal computations that produced a specific answer and say “here’s why it said that.” You’re not debugging a function. You’re interrogating a black box that’s very good at sounding like it isn’t one.
Wikipedia Had The Receipts#
It’s worth remembering what we had before this. Wikipedia, for all the grief it took early in its life, actually solved the trust problem reasonably well. Every article has an edit history. Every contested claim has a citation, or a “citation needed” tag flagging that it doesn’t. You can see who edited a page, when, and what the previous version said. If two editors disagreed about a technical detail, you could read the talk page and watch them argue it out with sources.
That’s an audit trail. It’s imperfect, and plenty of bad information slipped through over the years, but the mechanism for checking it was right there in the open. You could verify a Wikipedia claim without needing anyone’s permission. You just clicked “View history.”
That said, the audit trail wasn’t as untouchable as it looked. Wikipedia has quietly deleted edit histories from its own databases over the years, whether to comply with legal takedown requests, scrub privacy violations, or bury editorial disputes that someone decided shouldn’t be searchable anymore. When that happens, the receipts you thought you could always go check simply aren’t there. That’s a real crack in the trust model, not because the mechanism was fake, but because an institution that will selectively erase its own record isn’t fully committed to the transparency it advertises.
AI chat interfaces didn’t just inherit that weakness. They removed the mechanism entirely and replaced it with a delivery format that feels more authoritative, not less. The output is cleaner. The prose is more confident. There’s no visible disagreement, no edit war, no “citation needed” tag. It reads like a finished product instead of an ongoing negotiation between sources, which is exactly what makes it more dangerous to treat as a source of truth. The confidence of the delivery has nothing to do with the accuracy of the content.
Treat AI As A First Draft, Not A Source#
None of this means you should stop using AI to work faster. It means you should be precise about what job you’re asking it to do. AI is excellent at accelerating search, drafting configuration you’re going to review anyway, and summarizing a pile of documentation you’d otherwise have to read cover to cover. It is not a substitute for the primary source. It’s a shortcut to finding the primary source, if you use it that way.
In practice, this looks like a specific habit, one you should apply every time an AI model hands you a command, a configuration block, or a piece of architecture advice.
- Trace the claim back to vendor documentation. If the model tells you a flag exists on a particular command, open the actual manual page or vendor docs and confirm it. Flags get deprecated, renamed, or invented outright by a model that’s pattern-matching against a similar tool.
- Check the RFC or spec directly for protocol behavior. If the AI describes how a protocol handles a particular edge case, that’s a testable claim against a written standard. Don’t take the paraphrase. Read the source paragraph.
- Read the actual source code for anything load-bearing. If you’re relying on a library’s behavior for something that touches production, the source is the ground truth, not a model’s summary of what the source “probably” does.
This isn’t paranoia. It’s the same discipline good engineers already apply to an old Stack Overflow answer from six years ago that might reference a deprecated API. You already know not to copy-paste that answer into production without checking it against the current docs. AI-generated answers deserve exactly the same skepticism, arguably more, because they don’t even carry a timestamp telling you when the underlying knowledge went stale.
The Discipline Doesn’t Change, Only The Tool Does#
The uncomfortable truth is that verifying AI output takes real time, and that cuts against the entire appeal of using AI in the first place. That’s fine. The value of AI was never that it eliminates the work of verification. The value is that it makes the first pass faster. It can point you toward the right RFC section, the right man page, the right function in a codebase you’re not familiar with, faster than you could find it yourself.
Where engineers get into trouble is when they let the speed of the first pass substitute for the second pass. A fast wrong answer is worse than a slow right one, especially when the wrong answer is a firewall rule, an IAM policy, or a database migration.
So treat every AI answer the way you’d treat an intern’s first draft. It might be right. It’s probably close. But you sign your name to the final version, not the model, and that means you’re the one who checks it against the source before it goes anywhere near production.
Featured image by Aniket Narula on Unsplash

