How to Qualify 5,000 Leads Without Reading Every Website
You don't read 5,000 websites. You write the qualifying question once, as a column header, get an answer for every row, and then read only the rows the answer is unsure about. Add a confidence number next to each answer, sort by it, and your list of 5,000 guesses turns into a short pile worth a human's time.
That is the whole method. The rest of this post is how to set it up so the shortlist is one you can trust, and where it can go wrong.
Why most advice for this question misses
Search for how to qualify leads and you mostly get guides about inbound leads: form fills, chatbots, website visitors and what they do on your pricing page. That is a real problem, but it is not yours. You have a CSV. Maybe it came from Apollo, Sales Navigator or a scraper. Every row is a company that never contacted you, and the only way to know if it fits is to look at what it does.
That is the job a VA does when you pay someone to open each site and fill in a column. It is slow, and it is hard to keep the same standard from the first row to the last.
The second group of results is tools. Several of them score a list of domains against your ideal customer: you describe the customer in plain language and get a 0 to 100 fit score per domain, with reasons. Superlead's scorer scrapes company websites, classifies industry, estimates size and detects tech stack. Valley takes a Sales Navigator or CSV upload and has a button that deletes low fits. Cleanlist scores on company, tech, job title and location signals.
Those are fine if you want a score and a ready-made workflow. They answer "how well does this company fit?" They don't always answer the question that decides how you spend your afternoon: how sure is the tool about this particular answer?
A fit score and a confidence number are different things
Take a company whose homepage is three sentences and a contact form. A tool can say "B2B SaaS: yes" and be guessing. Another company has a pricing page, a docs site and a careers page, and the same answer is on solid ground. The answer is identical. The reliability is not.
Here is a made-up example, to show the shape. This is not real data.
| Company (example) | Is this a B2B SaaS company? | Confidence |
|---|---|---|
| acme-example.com | Yes | High |
| northwind-example.io | Yes | Low |
| plumbing-example.net | No | High |
| studio-example.co | Unclear | Low |
If you only had the Yes/No column, you would either trust all four or check all four. With the confidence column, you check two. At 5,000 rows, that difference is the whole job.
How to do it, step by step
1. Write the question like you'd ask a person
Put the question in the column header, in plain English. Be specific about what counts.
- Bad:
Good fit? - Better:
Does this company sell software to other businesses, and does it have a sales team?
If your export has no industry column, this is where you fix that. Ask about the website directly: Does this company sell software to other businesses? You don't need the data to already be in the sheet. You need the question to be answerable from the company's own site.
One question per column. Don't stack three conditions in one header, because when the answer is wrong you won't know which condition failed. Split it into columns: Sells software to businesses?, Has a sales team?, Based in the US?
2. Test on a sample, not the whole list
Before you run 5,000 rows, run 200 or so, picked at random. Read every answer. You are checking two things: is the question wording doing what you meant, and where do the low-confidence rows cluster? If they are all one kind of company (say, tiny sites with no content), that tells you something about your list, not just the tool.
On Columns, the free plan has 500 row-questions and needs no card, so a 500-row sample with one question fits inside it. That is enough to find out if it works on your list before you pay anything.
3. Run the full list
Now run all of it. Each row gets an answer and a confidence number.
4. Sort by confidence and read the bottom
Sort the confidence column. Read the low rows by hand, since that is where the answers are most likely to be wrong. Fix or throw out what you find.
Then do something people skip: read a small random sample of the high-confidence rows too. Not all of them, just enough to see if the answers hold up on your list. If you find errors there, tighten the question and run it again.
5. Don't trust the number blindly
A confidence score is a guide for where to look. It is not a promise. One engineering team that measured its own LLM's confidence scores found the answers were wrong 16 percent of the time even at maximum confidence, according to the post's title. That is one team's result on its own task, so don't read it as a rate for your list. But it is a good reason to spot check the top of the pile as well as the bottom. A paper on confidence-based evaluation points the same way, reporting that model confidence scores are often poorly calibrated. Treat the number as a way to decide where to spend your attention.
That is also why nobody should tell you the answers are always right. They are not. The point of a confidence column is that you stop pretending they are.
What this costs, with the math
Say you have 5,000 rows and one question. That is 5,000 row-questions, because a row-question is one question against one row. On Columns that fits inside the Starter plan, which is $29 a month for 25,000 row-questions. Run five columns like that and you have used the month. This is an example calculation, not a quote for your list.
Two things keep the cost down. Cached answers are not billed again, so re-running a list after you tweak one column doesn't re-charge the rest. And pricing is by usage, never per seat, so a team of three doesn't pay three times.
Compare that with a VA. You already know how long a few hundred rows takes by hand. I'm not going to put a number on it here, because the honest answer depends on your VA, your sites and how strict your criteria are.
When this is not the right tool
- Your list is small. If you have 80 leads, read them. It's faster than setting anything up.
- The answer isn't on the website. Revenue, headcount and funding often aren't. A question about something the site doesn't say will get a low confidence answer at best.
- You need a ready-made scoring model. If you want points for job title, company size and region, one of the scoring tools above fits better.
- A wrong answer is expensive. If one bad row costs you a lot, put a human on every row.
A simple checklist
- Write one clear question per column, in plain English.
- Run a random sample of about 200 rows and read every answer.
- Fix the wording until the sample looks right.
- Run the full list.
- Sort by confidence. Read the low rows. Spot check some high rows.
- Export the shortlist.
Where to start
Take your CSV, pull a random sample, write one question as a column header and see which rows come back unsure. That's usually a quicker way to judge the method than any amount of reading about it. If you want to try it, start free, no card and use the 500 row-questions on a sample of your own list.
Frequently asked questions
- How do I qualify thousands of leads without reading every website?
- Write your qualifying question as a column header, have it answered for every row, and review only the rows with low confidence. Spot check a small sample of the high-confidence rows too.
- Can AI tell if a company fits my ICP from its website?
- Often, if the answer is visible on the site, such as what the company sells and to whom. It can be wrong on thin sites, so read the low-confidence rows and check a sample of the rest.
- What is the difference between a fit score and a confidence number?
- A fit score says how well a company matches your ideal customer. A confidence number says how sure the tool is about its own answer. A company can fit well and still have a low-confidence answer if its site says little.
- How many rows should I spot check?
- Read all the low-confidence rows, then a small random sample of the high-confidence ones, enough to see if the answers hold up. There is no one number that fits every list.
- What does it cost to qualify 5,000 leads with one question?
- One question on 5,000 rows is 5,000 row-questions. On Columns, Starter is $29 a month for 25,000 row-questions, so that is about a fifth of a month's allowance. Cached answers are not billed again.