Can ChatGPT Qualify a List of Leads From a CSV?
Updated October 9, 2026
Yes, ChatGPT can qualify a CSV of leads. It cannot tell you which of its answers to trust, and that is the part that costs you. By the end of this post you will know which route fits your list size, have a paste-ready prompt that adds an "Unclear" label and a confidence column, and be able to work out how many rows you actually have to read.
ChatGPT can judge a row. It can't flag the rows it guessed on.
Give it a company name, a website description and a one-sentence ICP, and it will say yes or no for each row. That part works.
The trouble starts when 5,000 answers come back and every one of them reads the same. A confident yes and a guess look identical. So you either trust the whole column or you re-check the whole column. Both cost you days.
If that is your exact problem, How to Know Which AI Answers in Your Spreadsheet Are Wrong goes deeper on it. This post is about the ChatGPT route specifically.
The upload cap is a file size, not a row count
You will see people ask "how many rows can ChatGPT take?" There is no row number. OpenAI's help page says a CSV or spreadsheet cannot exceed about 50MB, depending on the size of each row, and that any file uploaded to ChatGPT has a hard limit of 512MB.
So the question is how big your rows are. Use this:
Rows that fit = 50MB (about 50,000 KB) divided by your average row size in KB.
Example: say your export averages 2 KB a row (name, URL, a short description). That is roughly 25,000 rows. Say instead you scraped the homepage copy into every row and it averages 20 KB. Now it is roughly 2,500 rows. Same cap, ten times fewer leads.
To find your own average, take the file size in KB and divide by the row count.
Other tools have their own caps, and they are not the same. Zoho CRM users, for instance, have reported an import rejected for a file over 24,689 rows. Check the cap of whatever tool the list passes through before you plan the job around it.
And fitting under the cap tells you nothing about quality. ChatGPT can misread dates and handle formats inconsistently, has limits on how much it processes at once, and does better when you break a request into smaller parts. A file that uploads is not a file where every row got the same attention.
Pick the route by list size
| Your list | Do this | Skip this |
|---|---|---|
| Under a few dozen rows, one-off | Paste the rows into a chat, state the ICP once, read every answer | Building a workflow |
| Hundreds of rows | Split into batches, identical criteria text in every batch, add the Unclear label below | One giant paste |
| Thousands of rows | Run it row by row through a bulk tool or workflow, with a confidence number on each row | Reading every answer, or trusting every answer |
Bulk tools exist because a chat window doesn't scale. Datablist lets you define labels, pick the text column and run the whole file, and it asks the model for one label from your list per row. GPT for Work publishes a how-to for grading and scoring leads in bulk in Sheets or Excel. n8n has a template that takes a CSV through validation, enrichment and sorting with a GPT model. All of that gets you answers. It's one label per row, and the label doesn't say how sure the model was.
If you're running thousands of rows, the scale side of this is in How to Qualify 5,000 Leads Without Reading Every Website. I won't repeat it here.
The paste-ready prompt: add "Unclear" and a confidence number
This is the copy-and-use part. It fixes the biggest problem with a plain yes or no: a forced choice makes the model guess on thin rows, and the guess looks like a decision.
You are qualifying leads against one rule.
Qualified = [B2B software company, 10 to 200 employees, sells to agencies].
For each row, read only the Company, Website and Description columns.
Return a table with these columns:
Company | Verdict | Confidence | Reason
Verdict must be exactly one of: Qualified, Not qualified, Unclear.
Use Unclear when the description is too thin to judge. Do not guess.
Confidence is a whole number from 0 to 100.
Reason is one sentence, 15 words or fewer.
Return one output row for every input row, in the same order.
Swap the bracketed sentence for your own ICP and leave the rest alone. Three details matter:
- "Unclear" gives thin rows somewhere to go. Every Unclear is a row for a human.
- "One output row for every input row" lets you count. Compare rows out to rows in. A mismatch is the cheapest way to catch a dropped row.
- The criteria sentence is identical in every batch. Reword it between batches and the same company gets a yes in one and a no in another.
Treat the confidence number from a chat prompt as a triage signal, not a measurement. It tells you where to look first. It does not tell you how likely the answer is to be right.
Here is what the output looks like. Example: these two rows are invented to show the shape.
| Company | Verdict | Confidence | Reason |
|---|---|---|---|
| Northfield Digital | Unclear | 35 | Description is two words, no product named. |
| Relay Metrics | Qualified | 92 | Product page describes reporting software for agencies. |
Sort by Confidence, lowest first. Read those. Accept the high end as your shortlist.
Work out your review budget before you start
Review budget = rows × share flagged low-confidence or Unclear.
Example: say you have 5,000 rows and 6 percent come back low-confidence or Unclear. That is 300 rows to read by hand instead of 5,000. If it's 20 percent, it's 1,000, and that tells you something too: your descriptions are too thin, or your criteria are too vague. Fix the input before you run more rows.
Run a 20-row sample twice before you run the list
Pick 20 rows. Run them. Run the same 20 again, same prompt.
If a row's verdict flips between the two runs, your criteria sentence is vague and the model is filling the gap differently each time. Tighten the sentence (add a number, name the buyer) and run the 20 again. Don't rerun the whole file to "check" it. You pay for the full run a second time and still can't tell which of two different answers is right.
Five mistakes, and what each one costs
| Mistake | What it costs you |
|---|---|
| Pasting thousands of rows into one chat | Later rows get thinner treatment, or the output stops partway. The list looks done and isn't. |
| Judging from the company name alone | Confident-sounding wrong answers. Give it the website or a description column. |
| Forcing yes or no with no Unclear option | Wrong rows that look exactly like right ones. |
| Uploading every export column | A bigger file for the same question, and a cap you didn't need to hit. |
| Rewording the criteria between batches | An inconsistent shortlist nobody can audit. |
Where Columns fits
If you'd rather not keep the prompt, the batches and the row counts in your head, that is what Columns does. You upload the CSV and write the question as a column header in plain English, for example "Is this a B2B SaaS company that sells to agencies?" Columns answers it for every row, and every answer carries a confidence number, so you sort and read the low end.
The numbers on cost: one row-question is one question against one row. Five thousand rows with two questions each is 10,000 row-questions, which fits inside Starter at $29 a month (25,000 row-questions). Pro is $99 a month for 150,000. Cached answers are not billed again, and pricing is by usage, never per seat. Columns still gets rows wrong, which is why the number is there.
Do this in the next ten minutes
- Open your CSV and check its file size. Divide by the row count to get KB per row.
- Delete every column the question doesn't need.
- Write your ICP as one sentence with a number in it.
- Pull 20 rows and run the prompt above twice. Note which rows flip.
Then you know whether your list is a chat job, a batch job or a bulk job, and which rows are going to need your eyes.
Try the confidence column on your own list
The free plan gives you 500 row-questions with no card, which is enough for 250 rows with two questions each. Take a sample from the list you were about to paste into ChatGPT and see which rows come back low. Start free, no card.
Frequently asked questions
- How many rows can ChatGPT handle in a CSV?
- There is no published row count. OpenAI says a CSV or spreadsheet cannot exceed about 50MB, depending on the size of each row, so divide 50,000 KB by your average row size to estimate it.
- Why does ChatGPT give a different answer when I rerun the same lead list?
- Vague criteria leave the model to fill gaps differently each time. Rerun a 20-row sample twice, and if any verdict flips, tighten the one-sentence rule and test again.
- Can ChatGPT give a confidence score for each lead?
- You can ask for a 0 to 100 number and a one-line reason in the prompt, then sort on it. Treat it as a triage signal for where to look first, not a measured probability.
- Should I let ChatGPT answer yes or no only?
- No. Add an Unclear option so thin rows have somewhere to go, then send every Unclear to a human instead of letting the model guess.
- How do I check that ChatGPT did not drop any rows?
- Ask for one output row per input row in the same order, then compare the row count going in with the count coming out. A mismatch means rows were dropped or merged.