ChatGPT can do part of keyword research well and one part of it badly. It is genuinely strong at expanding a topic into the phrasings people actually search, reading the intent behind a query, and grouping a messy list into the pages those keywords belong to. It cannot tell you how many people search a term, how hard a term is to rank for, or who currently ranks. Ask it anyway and it will answer with a confident number it made up.
That split is the whole answer, and it is worth being precise about, because "AI can do keyword research now" and "AI cannot do keyword research" are both wrong in ways that cost money.
The six steps, and how a chatbot scores on each
Keyword research is not one task. Splitting it up makes the useful and useless parts obvious.
- Expansion (strong). Give it a seed and it produces phrasings, question forms and modifiers, including the beginner wording you have forgotten how to use because you know the subject too well. This is the step it is best at.
- Intent classification (strong). Informational, commercial, transactional, plus the finer read of whether someone is comparing vendors or still naming their problem. Language judgment, which is what these models are for.
- Clustering (strong). Grouping keywords by meaning rather than shared words. It will correctly put "affordable" and "cheap" variants together, and correctly separate a definition query from a buying query.
- Search volume (unreliable). No access to the data. The numbers it produces look like keyword tool output and are not connected to anything. More on this below, because it is the step that causes real damage.
- Difficulty (unreliable). Difficulty depends on which pages currently rank and how strong they are, which is live information the model does not have. It will pattern-match from training data and be systematically out of date.
- Briefing (strong). Turning a cluster into an outline, listing the subtopics implied by the long tail, spotting what the query needs answered. Fine work, especially when you paste in the current top results for it to react to.
Four out of six is a good tool. It is also a tool with a specific hole in it, positioned exactly where people are least likely to notice.
The invented volume problem
Ask ChatGPT for the monthly search volume of "commercial espresso machine repair" and you will get a number. It will be a round-ish, plausible figure, delivered in the same tone as everything else it says. There is no dataset behind it. The model has learned what keyword volume answers look like and is producing something shaped like one.
This matters more than a normal hallucination because of what happens next. Nobody sanity-checks a search volume. You would notice if it invented a statistic about your industry, but 1,300 monthly searches for a term you have never measured is unfalsifiable at a glance, and it goes straight into a spreadsheet and then into a quarterly plan. Teams have committed months of writing to terms nobody searches on the strength of a number a chatbot produced in half a second.
The same applies to difficulty, with an extra wrinkle: even when the model is reasoning from something real, its picture of the results page is months or years old. Results pages move. A term that was open when the training data was gathered may now be defended by three strong pages and an AI overview.
The rule that keeps you safe is simple. Use the model for anything that is a language judgment. Use a data source for anything that is a measurement. If a tool will not tell you where a number came from, treat the number as decoration.
A workflow that actually works
Here is the version that gets the benefit without the failure mode. It takes about an hour for a topic area.
Step one: expand with the model. Give it your seed, your business context and who you sell to, and ask for the queries that audience would type at each stage of their problem. Ask for the wording a beginner would use and the wording a specialist would use separately; you will get two different and both useful lists. Push for a hundred or more, then deduplicate.
Step two: get real numbers on the list. Take the expanded set to a source with actual data. Google Keyword Planner gives banded volumes free if you have an Ads account, and dedicated tools give you difficulty alongside it. This is the step you cannot skip and cannot delegate to a chatbot.
Step three: bring the scored list back for intent and clustering. Now the model is working on real data instead of imagined data, and this is where it earns its keep. Ask it to classify intent per keyword and group the ones a single page could satisfy. Review the groups; over-clustering is the common failure, where a definition query and a vendor comparison get merged into one page that serves neither.
Step four: check the live results yourself. For the five or six clusters you plan to act on, open the actual results page. Twenty minutes of looking tells you more about winnability than any score, and it is the only way to catch a term whose competitive picture changed recently.
Step five: brief and write. Back to the model for outlines, with the current top results pasted in so it can react to what exists rather than guess. The full method, including the parts a chatbot has no view of, is laid out on our SEO keyword research pillar.
Why not just use a tool that does both
The four-window workflow above works, and it is tedious. Copying a list out of a chat window into a keyword tool, exporting a CSV, pasting it back into the chat, then rebuilding the result in a spreadsheet is forty minutes of clerical work per topic. Doing that weekly across several sites is where good intentions go to die.
That plumbing is what an AI keyword research tool is supposed to remove: model on the language steps, real data on the measurement steps, joined so you get a prioritized set of pages instead of four browser tabs. The test for whether any of them is honest is the one above. Ask where the volume number came from. A real answer names a source; a marketing answer says "our AI".
What ChatGPT is better at than any keyword tool
Two things, and both are underused.
The first is describing an audience in its own words. Keyword tools show you what was searched; a model will tell you why, and reason about the person behind the query. "Someone searching this has already tried the free option and hit a limit" is a claim you can verify by reading the results page, and it changes the angle of the page you write more than a volume number ever will.
The second is the strategy conversation around the list. Which of these clusters serves someone who will pay us. What are we missing that a competitor would cover. Which of these should be one page instead of three. Those are the questions the analysis is for, and they are the ones no keyword tool answers.
There is a growing practical reason to care about the second one. A rising share of buyers now ask an assistant for a recommendation instead of scrolling results, and the pages those assistants quote tend to be the ones that answer a specific question cleanly and are recently updated. If you want to know whether your pages are actually being cited in those answers, that is a brand monitoring problem rather than a ranking one, and it is worth tracking separately from your Search Console numbers.
The short version
Can ChatGPT do keyword research? It can do the thinking parts and not the counting parts. Use it to expand, to read intent, to cluster and to brief. Get volume and difficulty from something with data behind it. Check the live results yourself before you commit real writing time. Anyone telling you a chatbot replaces the data layer is describing a workflow that produces confident plans built on numbers nobody measured.
If you want the joined-up version without the copy-paste, drop a seed into the free Keywordpilot demo and see the expanded set come back already scored and clustered.
Put it into practice
The fastest way to apply this article: run your own niche through the free Keywordpilot demo. One seed keyword, twenty seconds, a clustered mini content plan. No account.