How to Know Which AI Answers in Your Spreadsheet Are Wrong
You can't tell which AI answers in a spreadsheet are wrong by looking at them, because wrong answers read just as smoothly as right ones. What works is a number on every row. Sort by a per-row confidence number, have a human check the lowest rows first, spot check a few of the high ones, and never assume any row is safe just because the AI sounded sure.
Updated October 3, 2026.
What a confidence number is, and what it isn't
A confidence score is a visible estimate of how sure the system is about an answer. It can show up as a percent, a bar or a level, and the point is to let you decide per row whether to accept, verify or escalate.
That's all it is. It's an estimate, not a proof. A row with a high number can still be wrong, and a row with a low number can still be right. The number doesn't tell you the answer is correct. It tells you where to look first.
Plenty of tools give you the answer with nothing next to it. Then your only choices are to trust all 5,000 rows or check all 5,000 rows. The first hides the wrong ones. The second throws away the time you were trying to save.
Why checking everything doesn't scale
The usual advice for checking an AI answer comes from chat. One checklist suggests sending the same question to several models and comparing: if they agree, you can feel better, and if they disagree, at least one is wrong. That's reasonable for a single question. It's one site's advice, and it isn't written for spreadsheets.
On a list it falls apart. Every row now costs you several answers instead of one, and you still have to read the disagreements yourself. Run that on a few thousand rows and you've built a bigger pile to check than the one you started with.
A per-row number is the middle path. You get one answer per row, plus a signal about which rows need a person.
Can you trust the number itself?
Partly, and you should say so out loud before you rely on it.
Research on language models keeps finding the same problem: they tend to be more confident than they should be. That paper looks at AI used as a judge of other AI output, not at lead lists, and I'd read it before leaning on any specific claim. A model that sounds sure is not the same as a model that is right.
One vendor blog puts the cause down to fluency: these systems are built to produce natural-sounding text, so a guess can read as confident. That's one company's explanation, so treat it as an opinion, not a measurement.
The same research line makes a second point worth keeping. Sending low-confidence items to a human works as quality control when the scores reflect real uncertainty. That is why the rule below has two parts. You review the low rows, and you also spot check some of the high ones, because the number might be off.
Columns puts a confidence number on every answer. It doesn't claim those answers are always right, and nothing here says how well any tool's numbers match reality on your list. You find that out by testing, which is the next section.
A made-up example
Here's what it looks like. The question goes in the column header, and the confidence sits beside the answer. This table is an invented example with fake companies and fake numbers. It is not a real result.
| Company (made up) | Is this a B2B SaaS company? | Confidence |
|---|---|---|
| Brightlane Software | Yes | 94 |
| Corner Bakery Supply Co. | No | 91 |
| Northwind Logistics Tools | Yes | 62 |
| Pinecrest Consulting | Yes | 38 |
| Harbor Analytics | Yes | 93 |
Sort by the last column. Pinecrest, at 38, gets a human look first. Northwind, at 62, is next. The three high rows look fine.
Now say, in this invented example, you open Harbor Analytics and find it's a staffing firm that happens to have "analytics" in the name. A high-confidence answer, and wrong. That's the reason for the spot check. If you only ever read the low rows, you'd never catch it.
The review rule
Keep it simple. Three buckets:
- Low confidence rows: a human looks at every one. Open the website, fix the answer or drop the lead.
- High confidence rows: check a small random slice. Pick them without looking at the answers, so you don't only check the ones that seem plausible.
- Expensive mistakes: anything where a wrong answer costs you, like a lead you're about to email in a big campaign, gets checked no matter what the number says.
Then the harder question: how big should that slice be? There's no universal answer. Data-labeling teams don't review everything either. One labeling-services site says checking every annotation in a large dataset isn't practical, so quality checks run on a representative sample. It doesn't hand you a number for your list, and neither will I. It depends on how messy your data is, how sure you need to be, and how much time you have.
Test it in an afternoon before you trust it
Don't pick a cutoff because someone online said 70 is high. Scores run on different scales in different products. Microsoft's docs for its old QnA Maker tool, for instance, describe a 0 to 100 score with its own bands, and those bands belong to that product, not to yours. Find your own.
Here's a test that costs very little:
- Take a real sample. Use a few hundred rows from your own list. The free plan gives you 500 row-questions, no card, so you can do this without paying.
- Plant a few known rows. Add 10 to 20 companies you've already qualified by hand, so you know the right answer. Labeling teams use the same trick, mixing in items with a known correct answer to watch accuracy. The 10 to 20 is my suggestion, not a sourced figure.
- Sort by confidence. Check a few rows from the top, a few from the middle and a few from the bottom.
- Find where the wrong answers start. Note the confidence level where you first see mistakes show up regularly. That's your cutoff, for this question, on this kind of list.
- Check the known rows. If the column gets companies you already know wrong, the problem is probably the question, not the cutoff.
If you find wrong answers among the high rows, widen your spot check. Don't narrow it.
What to do with the low rows
Three moves, in order of effort:
- Look and fix. Open the site, read the homepage, correct the answer.
- Drop it. If a lead is low confidence and low value, skip it. A short list you trust beats a long list you don't.
- Fix the question and re-run. Sometimes a whole cluster of low rows comes from a vague question. "Is this a B2B SaaS company?" is clearer than "Is this a tech company?" Rewrite the header and run it again. Cached answers aren't billed again in Columns, so a re-run on rows that didn't change is cheap.
Don't borrow someone else's numbers
You'll find review rates and quality scores quoted all over the data-labeling world. One vendor says its own tool can lower review rates to around 10 percent while keeping quality scores above 95. That's a claim about that vendor's product on that vendor's data. It tells you nothing about a lead list or about any column you build. Don't copy a percentage from another field and call it your process.
Your own sample, your own cutoff, and your own spot checks are the only numbers that count.
From the test to the whole list
Once you have a cutoff you trust, the rest is routine: run the full column, sort by confidence, work the low rows, spot check the high ones. If you want the whole walkthrough for a big list, How to Qualify 5,000 Leads Without Reading Every Website covers the qualification step end to end.
Try it on your own list
Write the question as a column header, run it on a sample, and sort by the confidence column. You'll see in an afternoon which rows deserve your time and which answers hold up. Start free, no card. The free plan covers 500 row-questions, enough to run the test above. Starter is $29/month for 25,000 row-questions and Pro is $99/month for 150,000, priced by usage, never per seat.
Frequently asked questions
- How do I know which AI answers in my spreadsheet are wrong?
- You can't see it from the answer alone. Use a confidence number on every row, sort by it, and have a person check the lowest rows first. Spot check some high rows too, since a high number can still sit on a wrong answer.
- Is a confidence score the same as accuracy?
- No. A confidence score is an estimate of how sure the system is, shown as a percent, bar or level. It helps you decide where to look, but it doesn't guarantee a row is right or wrong.
- How many rows should I spot check?
- There is no universal number. It depends on how messy your data is and how costly a wrong answer is. Test on a small sample, check rows across the confidence range, and widen the check if you find mistakes among the high rows.
- Can I skip the high-confidence rows?
- Not entirely. AI tends to sound sure even when it's guessing, so check a small random slice of the high rows. Always check any row where a wrong answer would be expensive, like a lead in a big email campaign.
- What should I do if my sample shows too many wrong answers?
- Look at the question first. A vague column header often causes a cluster of bad rows. Reword it and re-run on the same sample. Cached answers in Columns aren't billed again, so a re-run on unchanged rows is cheap.