How to Filter a Lead List by ICP Fit When There Is No Industry Column
Updated October 5, 2026
If your lead list has no industry column, turn your ICP into one plain-English yes or no question per row, answer it from the company's website or domain, and review only the rows with low confidence. Split the list into clear fits, clear misses and a review pile. You never need an industry field to do it.
That is the whole method. The rest of this post is how to set it up so it holds up on a real list.
Why the usual advice doesn't help here
Most ICP filtering guides assume the list already has firmographic fields: industry, headcount, revenue band, region, buying role. Instantly's roundup of ICP filtering tools is a typical example, and Hunter's help page on filtering leads works the same way: you filter on fields the tool already holds. That is fine when the data came out of a database that has them.
It is no help when you are holding a CSV from a scraper, a Sales Navigator export or a client's old CRM dump, with a company name, a domain, a job title and nothing else. The industry cell is empty or missing, and a formula can't fill it, because "does this company sell software to other businesses?" is a judgment question, not a lookup.
You have three options:
- Buy industry codes for every domain. Dataprovider, for example, sells GICS, NAICS and SIC data for websites. That works, but it is another vendor, another cost and another import.
- Have a person open every site. This is the VA route, and it is why qualifying a list takes days.
- Ask a question of each row and let the answer carry its own confidence. This is what the rest of the post covers.
Step 1: write your ICP as questions, not as a filter
An ICP is usually written as a sentence: "B2B software companies with a small sales team." A sentence can't be sorted. A question per row can.
Break the ICP into the parts you can actually judge from what the list has:
- What they sell: "Does this company sell software to other businesses?"
- Who they sell to: "Is this company's main customer another business, not consumers?"
- Rough size: "Does this company look like a small team of under 50 people?" (Use your headcount column for this one if you have it. A real number beats a guess.)
- Disqualifiers: "Is this an agency, consultancy or reseller?"
Start with one question, not four. Every extra question multiplies your cost and your review work, and one well-worded question often does most of the filtering. You can add a second later.
Two rules for wording:
- Make it answerable from a website. "Is this company a good lead?" is too vague. "Does this company sell software to other businesses?" can be read off a homepage.
- Ask for a yes or no, then add a second column that asks for the reason in a few words. The reason column lets a reviewer judge a row in seconds instead of opening the site.
Step 2: run it and read the confidence column
In Columns, the question is the column header. You upload the CSV, type the question as a new header, and every row gets an answer plus a confidence number. There is no prompt to engineer and no formula to write.
Here is what it looks like. This is a made-up example with invented companies and invented confidence numbers, not real output.
| Company | Domain | Is this a fit: B2B software company? | Confidence | Reason |
|---|---|---|---|---|
| Northline Metrics | northlinemetrics.example | Yes | 94 | Sells analytics software to finance teams |
| Brightcup Coffee | brightcup.example | No | 91 | Sells coffee to consumers |
| Harbor & Pine | harborpine.example | Yes | 41 | Site is mostly a logo and a contact form |
| (parked domain) | oldsite.example | No | 22 | Placeholder page, nothing to read |
The first two rows are easy and the answers are confident. The third is a "yes" the tool isn't sure about, and the fourth is a "no" it has almost nothing to base an answer on. Those last two are exactly the rows you want in front of a human. The confidence number is what tells you that. Without it, all four rows would look equally trustworthy.
One caution that applies to any AI answer on a list: the answers can be wrong. A company with a vague website can be misjudged. That is why the confidence column exists, and it is why you should not treat a high number as proof. If you want the longer version of how to decide which answers to doubt, the article on how to know which AI answers in your spreadsheet are wrong covers it.
Step 3: split the list into three piles
Sort by the answer, then by confidence. You end up with three piles:
| Pile | What's in it | What you do |
|---|---|---|
| Clear fits | "Yes" answers with high confidence | Move straight to outreach |
| Clear misses | "No" answers with high confidence | Drop them |
| Review pile | Anything with low confidence, either answer | A person looks at each row |
The review pile is where your time goes, and it is smaller than the whole list. That is the point: instead of reading 5,000 sites, you read the few hundred the tool was unsure about.
How to pick the cutoff
There is no universal number for "high confidence," so don't borrow one from someone else's list. Pick yours on a small sample:
- Take a sample of rows, say 250.
- Run your question on them.
- Sort by confidence and open the sites from the top, the middle and the bottom of the list.
- Find the point where you start disagreeing with the answers. Set your cutoff a bit above that.
If you disagree with answers even at high confidence, your question is probably too vague. Reword it and run the sample again before you spend anything on the full list.
What to do with thin and broken sites
Some rows will never have much to read. A parked domain, a login-only page or a homepage that is one big image gives no real evidence about what the company does.
Those rows should come back with low confidence, and low confidence means they go to the review pile. You don't want a confident answer built on nothing, so check a few of these rows on your sample before you trust the pattern. If you find rows where the answer is confident but the site obviously has no content, that is a sign your question needs tightening, or that the row needs a human no matter what the number says.
Some classifier vendors say their tools fall back to an about page when the homepage is empty. One Apify classifier describes doing this, and that is the vendor's own claim. Whatever tool you use, test what it does on a few bad sites before you trust it with a whole list.
How to estimate the cost before you run it
Columns prices by usage. One row-question is one question against one row. Cached answers aren't billed again, so re-running a list you already ran doesn't cost you twice.
The arithmetic is rows times questions. Here are made-up examples:
- 250 rows with 2 questions (fit answer and a reason) is 500 row-questions. That is exactly the free plan: 500 row-questions, no card.
- 5,000 rows with 2 questions is 10,000 row-questions. That fits inside Starter, which is $29 a month for 25,000 row-questions.
- 20,000 rows with 3 questions is 60,000 row-questions. That is past Starter, so you would look at Pro, which is $99 a month for 150,000.
The row counts here are invented to show the math. Columns is never priced per seat, so adding a teammate to review the pile doesn't change the bill.
Clay prices in credits, and a credit and a row-question are different units, so a per-lead comparison between the two would mislead you. If you are weighing tools, check the current pricing page of each one instead of trusting a blog post, including this one.
A workflow you can run today
- Open your CSV and confirm you have a domain or company name for every row. Rows with neither can't be judged, so set them aside.
- Write one yes or no question that captures your ICP.
- Add a second header that asks for the reason in a few words.
- Run a 250-row sample. That is 500 row-questions, so it fits the free plan.
- Sort by confidence, open a handful of sites at each level and set your cutoff.
- Fix the question wording if you disagree with confident answers.
- Run the full list.
- Split into fits, misses and the review pile. Have a person work the review pile only.
For the broader version of this flow, including how to handle very large lists, the post on qualifying 5,000 leads without reading every website goes through it end to end.
Common mistakes
- Asking too many questions at once. Each one costs row-questions and adds another column to review. Start with one.
- Trusting every answer equally. The whole reason to have a confidence number is that some answers deserve more doubt than others.
- Skipping the sample. Running 20,000 rows before you know whether your question works is the expensive way to find out it doesn't.
- Writing the question like a pitch. "Is this company a perfect fit for our amazing product?" gets a mushy answer. State the criteria plainly.
- Setting the cutoff by feel. Check it against the sites on your sample.
Try it on your own list
Pull 250 rows from the list you have been avoiding, write one question as a header and look at the confidence column. If the three piles come out sensible, run the rest. If they don't, you've lost nothing: start free, no card and the 500 free row-questions cover the whole test.
Frequently asked questions
- Can AI tell what a company does from its domain alone?
- Sometimes, but a domain name is thin evidence. A question answered from the website content is more reliable, and rows where the site says little should come back with low confidence and go to a human.
- How many rows should I review by hand?
- Only the review pile: rows where the confidence is below the cutoff you set on a small sample. How big that pile is depends on your list and how clearly the question is worded.
- What if my list has no website or domain column?
- Then the company name is all there is to go on, and answers will be less certain. Set those rows aside for a person, or find the domains first.
- How do I test this without paying?
- The free plan includes 500 row-questions with no card. A 250-row sample with two questions per row uses exactly 500.
- Does a high confidence number mean the answer is right?
- No. It tells you which answers deserve less doubt, not that any answer is always right. Spot-check a few high-confidence rows on your sample before you trust the cutoff.