“LSI keywords” is an SEO name for words related to your main keyword. It borrows from latent semantic indexing, a real retrieval method patented in 1989. That is where the link ends. Google’s John Mueller has said there is no such thing as LSI keywords, and nobody has shown that Google ever used LSI to rank pages. The advice sold under the name is half right. Covering related subtopics, using your readers’ words and naming things clearly all help. Pasting a generated keyword list into a page does not.
This post goes back to the original 1990 paper, sets out what Google says it uses instead, and ends with a topic coverage method that works for search and for AI answers. Every claim links to its source at the bottom.
The first three figures come from the 1990 paper that introduced latent semantic indexing[2]. The last comes from Google’s post announcing BERT in Search[10].
What are LSI keywords?
When tools and guides say LSI keywords, they usually mean one of three things. Synonyms, such as “cheap flights” and “low-cost flights”. Words that tend to appear together, so a page about coffee mentions beans, grind and roast. Or related searches. Some tools sell these lists without saying how they were made. Bill Slawski, who wrote about search engine patents at SEO by the Sea, noted that one such site “doesn’t provide any information about how they generate those keywords or use Latent Semantic Indexing (LSI) technology”. His summary: “Google does like synonyms and Semantics, but they don’t call it Latent Semantic Indexing”[6].
So if you searched for “LSI keywords SEO” because a tool or a brief told you to add some, the short answer is this. Don’t add a list. Cover the topic. The rest of this page explains why, starting with what LSI actually was.
What latent semantic indexing actually is
Latent semantic indexing came out of Bell Communications Research. Its patent was filed in September 1988, granted in June 1989, and names seven inventors, including Scott Deerwester, Susan Dumais and Thomas Landauer. Its abstract presumes “an underlying, latent semantic structure in the usage of words”[1]. The method was published in 1990 as “Indexing by Latent Semantic Analysis” in the Journal of the American Society for Information Science[2].
The problem it set out to solve is one SEO writers know well. In the paper’s words, “the words searchers use often are not the same as those by which the information they seek has been indexed”. The authors split this into synonymy, many words for one thing, and polysemy, one word with many meanings. They cite research showing that “two people choose the same main key word for a single well-known object less than 20% of the time”[2].
The paper states the payoff plainly: “terms that did not actually appear in a document may still end up close to the document”. LSI can also measure how similar two terms are inside the collection[2]. But what it produced was a ranking of documents for a query. It was not a list of words for a writer to add, and the word relationships it found belonged to the one collection it was built from.
What the 1990 tests found
The authors tested LSI on two standard collections of titles and abstracts. On MED, 1,033 medical abstracts with 30 test queries, LSI’s average precision was .51 against .45 for plain word matching, which they called “a 13% improvement over raw term matching”. On CISI, 1,460 information science abstracts with 35 queries, both methods averaged .11. There, the paper says, the latent structure was “no more useful than raw term overlap”[2]. A promising result on one collection, at a scale of about a thousand documents.
The limits its authors listed
- Words with several meanings. LSI “offers only a partial solution to the polysemy problem”. A word like “bank” becomes one averaged point, which “may create a serious distortion”[2].
- Size. “Computational constraints have limited us to around 7000 terms”[2].
- Change. The paper describes a way to “fold-in” new documents, but adds that “how much of this updating can be done without having to perform a new decomposition is unknown”[2].
None of this makes LSI a bad idea. It was careful early research, and its authors said what they did not know. It was built for fixed collections. As Slawski put it, the patent “doesn’t discuss how a process such as this could handle something the size of the Web because nothing that size had quite existed yet then”[6].
Sources[1][2][9][3][10][11][8][4].
Does Google use LSI keywords?
No, according to Google. John Mueller of Google posted on Twitter in July 2019: “There’s no such thing as LSI keywords”, adding that “anyone who’s telling you otherwise is mistaken, sorry”[3]. In January 2023 someone asked him whether LSI keywords work better in headings or in body text. He replied that both have no effect, and that “Anyone who tells you to use LSI keywords is ... still wrong after all these years”[4]. Search Engine Journal also quotes him saying “we have no concept of LSI keywords. So that’s something you can completely ignore”[5].
What Google says it uses instead
Google has described, in its own words, how it gets past exact word matching. Three points matter for writers.
Matching words still counts. “The most basic signal that information is relevant is when content contains the same keywords as your search query”[7]. Use the words of the question you are answering.
Synonyms are handled for you. Google describes a “sophisticated synonym system” that finds relevant pages “even if they don’t contain the exact words you used”. Its example: a search for “change laptop brightness” can match a manufacturer’s page that says “adjust laptop brightness”[7].
Repetition is not relevance. “When you search for ‘dogs,’ you likely don’t want a page with the word ‘dogs’ on it hundreds of times.” Instead, Google says, its algorithms look for “other relevant content beyond the keyword”, such as pictures of dogs, videos or a list of breeds[7].
| What Google says it does | In Search since | |
|---|---|---|
| RankBrain | Understands how words are related to concepts, so pages can match without every exact word | 2015 |
| Neural matching | Understands representations of concepts in queries and pages, and matches them | 2018 |
| BERT | Understands how combinations of words express different meanings and intent | 2019 |
| Passage ranking | Identifies individual sections of a page to judge how relevant the page is | Not dated in the guide |
| MUM | Understands and generates language. Not currently used for general ranking | Announced 2021; specific uses only |
BERT is the clearest case. When Google brought it to Search in 2019, it said BERT would help it understand “one in 10 searches in the U.S. in English”, and that people often type “keyword-ese” instead of asking naturally[10]. By 2022, Google said BERT “plays a critical role in almost every English query”[9]. MUM, announced in 2021 as “1,000 times more powerful than BERT”[11], is “not currently used for general ranking in Search”[8].
Notice what none of these systems asks of a writer: a list of related words. RankBrain is described as returning relevant content “even if it doesn’t contain all the exact words used in a search”[8]. Google’s SEO Starter Guide makes the same point: “don’t worry if you don’t anticipate every variation of how someone might seek your content”, because its “language matching systems are sophisticated”. The same guide warns that “Excessively repeating the same words over and over (even in variations) is tiring for users”, and that keyword stuffing is against Google’s spam policies[12]. “Even in variations” is a polite description of a sprinkled LSI list. Our post on keyword stuffing covers the policy and the density myth.
What the LSI advice gets right, and what it gets wrong
Most LSI keyword advice is a good idea with a wrong explanation. The wrong explanation then leads to the wrong tactic: a list to work in, instead of a topic to cover.
| Holds up? | Why | Do this instead | |
|---|---|---|---|
| Google uses LSI to understand pages | ✕No | Mueller says there is no such thing as LSI keywords. Google's docs name other systems | Drop the label, keep the reader |
| Use synonyms and natural variations | ~ | Readers use different words, like 'charcuterie' and 'cheese board'. Google says you need not catch every variation | Use the words your readers use, where they read naturally |
| Cover related subtopics | ✓Yes | Google looks for relevant content beyond the keyword, and its AI features search across subtopics | Answer the questions a reader asks next |
| Name related entities | ✓Yes | Precise names tell search engines and models which thing you mean | Name products, people, places and versions exactly |
| Add a generated LSI keyword list | ✕No | Repeating words, even in variations, is what Google's starter guide warns against | Turn the list into questions, then write answers |
| Use each term a set number of times | ✕No | Google's own example: a page with 'dogs' on it hundreds of times is not what searchers want | Check which phrases dominate, then cut |
The synonym row deserves one more line. Google’s starter guide notes that “some users might search for ‘charcuterie’, while others might search for ‘cheese board’”, and that writing with those differences in mind “could produce positive effects”[12]. That is advice to know your readers, not to collect every variant.
Entities and topic coverage
If LSI keywords are the wrong idea, two ideas do the real work: entities and coverage.
An entity is a specific thing, such as a company, a product, a person or a place. “Python” could be a programming language or a snake. A page that says which one it means, and names the things around it precisely, leaves less to guess for a search engine or a model. Our entity SEO guide covers how to make your own company one clear entity across the web.
Coverage is the page-level half. Google’s helpful content guidance asks: “Does the content provide a substantial, complete, or comprehensive description of the topic?”[13]. Its “dogs” example makes the same point from the other side: pictures, videos, a list of breeds[7]. A page that covers a topic fully will contain related words, because you cannot explain a topic without them. A page that contains related words has not necessarily covered anything. That is the flaw in LSI keyword lists: they measure the side effect and miss the cause. For coverage across a whole site, see topical authority for AI answers.
How AI answer engines pick passages
Start with what is not known. No AI company publishes, in any detail, how it picks the passage it quotes or the page it cites. Any claim that an engine rewards a set of related keywords is a guess. What is documented is how some engines search, and it points the same way as everything above.
Google says AI Overviews and AI Mode may use a “query fan-out” technique, issuing “multiple related searches across subtopics and data sources” to build a response[14]. Its AI guide gives an example. For “how to fix a lawn that’s full of weeds”, fan-out queries might include “best herbicides for lawns”, “remove weeds without chemicals” and “how to prevent weeds in lawn”. The same guide says “AI systems can understand synonyms and general meanings”, so “you don’t have to worry that you don’t have enough ‘long-tail’ keywords”. And there is no need to cut content into tiny pieces, because “Google systems are able to understand the nuance of multiple topics on a page and show the relevant piece to users”[15].
OpenAI describes something similar for ChatGPT. When it searches through partner search providers, “ChatGPT search typically rewrites your query into one or more targeted queries”, and after reviewing the first results it may send “additional, more specific queries”[16].
Our reading, and it is an inference: each of those follow-up searches is a subtopic. A page that answers the main question and the obvious next ones, each in a clearly headed section, gives an engine something to retrieve for more of them. A page padded with related terms gives it nothing new to quote. For what is known, inferred and marketed about AI citations, see how AI assistants decide which sources to cite.
A topic coverage method that replaces LSI keyword lists
Every step below produces questions, not words to sprinkle. Work through them in order.
Search Console’s Performance report shows “what search queries are most likely to show your site”[17]. Google’s AI guide advises: “Don’t just recycle what others on the internet have already said”[15].
Two steps deserve more detail. For the second, our People Also Ask guide includes a planner that turns those questions into a page plan. For the third, you are reading competitor headings for the list of subtopics, not for wording. Your own data, examples and opinions are what make a section worth quoting, and they are the part a competitor cannot give you.
A worked example
Take Google’s own fan-out example. A page answering “how to fix a lawn that’s full of weeds” might plan these sections:
- Which weeds you have, and why it changes the fix.
- Herbicides that are safe for lawns.
- Removing weeds without chemicals.
- Stopping the weeds from coming back.
- When to reseed the bare patches.
The middle three come from the fan-out queries in Google’s guide[15]. The first and last are our additions, the kind of next question a reader asks. A related-terms list for the same page might suggest words like “dandelion”, “turf” or “broadleaf”. Each will appear on its own once the sections are written, because you cannot explain the fix without them.
What an LSI keyword generator actually gives you
An LSI keyword generator is a related-terms tool with a borrowed name. Depending on the tool, the list may come from related searches, autocomplete, or words that appear on top-ranking pages. Slawski’s complaint about one such site applies widely: it gave no “information about how they generate those keywords”[6].
Used with care, a list can still help you brainstorm. Read each term and ask one question: would a reader expect this page to explain it? If yes, it points to a missing section, so write that section. If not, leave it out. Never paste the list into the page, and never aim to use each term a set number of times.
Check which phrases dominate your page
After writing, check the page from the other side. A page written around a keyword list tends to lean on a few phrases. The free readability checker has a keyword density tab that shows which words and phrases a page repeats most. If one phrase dominates and sounds forced when read aloud, that is the repetition Google’s starter guide warns about. The same tool scores reading ease and flags passive voice; see Flesch-Kincaid explained and active and passive voice for what those numbers mean.
Finding coverage gaps with CoreCited
The method above starts from search questions. CoreCited’s Content Briefs start from AI answers: a question your buyers ask, where an engine names competitors and not you. The brief shows which sources the engines cited and what those pages cover that yours does not. Starter includes 10 briefs a month, and every limit is on the pricing page.
Questions people ask
What are LSI keywords?
An SEO name for words related to a target keyword: synonyms, related topics and terms that often appear alongside it. The name borrows from latent semantic indexing, a retrieval method patented in 1989. The lists sold as LSI keywords are not that method, and Google's John Mueller has said there is no such thing as LSI keywords.
Does Google use LSI keywords?
No, according to Google's John Mueller. He said in 2019 that there is no such thing as LSI keywords, and in 2023 that anyone telling you to use them is still wrong. Google's own documentation describes a synonym system, RankBrain, neural matching and BERT, and does not mention LSI.
What is latent semantic indexing?
A document retrieval method from researchers at Bell Communications Research, patented in 1989 and published in 1990. It builds a table of which words appear in which documents, compresses it into about 100 factors, and matches queries to documents in that compressed space, so a document can match a query even when they share few words.
Is an LSI keyword generator worth using?
Only as a brainstorming list. The terms are related words under a borrowed name, and they are a side effect of good coverage, not a cause of it. Read each term, ask whether a reader would expect the page to explain it, and write that section if so. Never paste the list into the page.
What should I do instead of using LSI keywords?
Cover the topic. Answer the main question first, then the questions a reader asks next, which you can find in People Also Ask, Search Console and conversations with buyers. Use your readers' words, name entities precisely, and check that no single phrase dominates the page.
Do LSI keywords help in ChatGPT or Google AI Overviews?
No AI company says they do. Google says its AI systems understand synonyms, so you do not need every keyword variation, and that AI Overviews and AI Mode may run several related searches across subtopics. ChatGPT search rewrites a prompt into its own search queries. Covering those subtopics clearly is what gives an engine more to work with.
