CoreCited
Content & authority

LSI keywords: why Google does not use them, and what topic coverage really means

LSI keywords explained: where the term came from, why Google says it does not use them, and how to cover a topic so search and AI engines understand it.

AlexFounder & Developer14 min read

“LSI keywords” is an SEO name for words related to your main keyword. It borrows from latent semantic indexing, a real retrieval method patented in 1989. That is where the link ends. Google’s John Mueller has said there is no such thing as LSI keywords, and nobody has shown that Google ever used LSI to rank pages. The advice sold under the name is half right. Covering related subtopics, using your readers’ words and naming things clearly all help. Pasting a generated keyword list into a page does not.

This post goes back to the original 1990 paper, sets out what Google says it uses instead, and ends with a topic coverage method that works for search and for AI answers. Every claim links to its source at the bottom.

Key takeaways
1LSI keywords are not a Google concept. John Mueller of Google said in 2019 there is no such thing, and in 2023 that anyone telling you to use them is still wrong.
2Latent semantic indexing is real: a 1990 method tested on collections of 1,033 and 1,460 abstracts. It ranked documents. It never produced keyword lists for writers.
3Google names its own language systems instead: a synonym system, RankBrain, neural matching, BERT and passage ranking. None of them asks you to add related words.
4What the advice gets right is coverage: answer the next questions, use your readers' words, and name the entities you mean.
5AI answer engines search for subtopics too. How they pick a passage is not published, so cover the follow-up questions clearly rather than chase terms.
1460 docs
in the largest test collection of the 1990 LSI paper
Deerwester et al., 1990
~100
factors LSI compressed every term and document into
Deerwester et al., 1990
<20%
how often two people pick the same main keyword for a well-known object
Furnas et al., cited in the 1990 paper
15%
of the searches Google sees each day are ones it has not seen before
Google, Oct 2019

The first three figures come from the 1990 paper that introduced latent semantic indexing[2]. The last comes from Google’s post announcing BERT in Search[10].

What are LSI keywords?

Definition · LSI keywords
An SEO term for words and phrases related to a target keyword: synonyms, related topics, and terms that often appear alongside it. The name suggests they come from latent semantic indexing. They do not.

When tools and guides say LSI keywords, they usually mean one of three things. Synonyms, such as “cheap flights” and “low-cost flights”. Words that tend to appear together, so a page about coffee mentions beans, grind and roast. Or related searches. Some tools sell these lists without saying how they were made. Bill Slawski, who wrote about search engine patents at SEO by the Sea, noted that one such site “doesn’t provide any information about how they generate those keywords or use Latent Semantic Indexing (LSI) technology”. His summary: “Google does like synonyms and Semantics, but they don’t call it Latent Semantic Indexing”[6].

So if you searched for “LSI keywords SEO” because a tool or a brief told you to add some, the short answer is this. Don’t add a list. Cover the topic. The rest of this page explains why, starting with what LSI actually was.

What latent semantic indexing actually is

Latent semantic indexing came out of Bell Communications Research. Its patent was filed in September 1988, granted in June 1989, and names seven inventors, including Scott Deerwester, Susan Dumais and Thomas Landauer. Its abstract presumes “an underlying, latent semantic structure in the usage of words”[1]. The method was published in 1990 as “Indexing by Latent Semantic Analysis” in the Journal of the American Society for Information Science[2].

The problem it set out to solve is one SEO writers know well. In the paper’s words, “the words searchers use often are not the same as those by which the information they seek has been indexed”. The authors split this into synonymy, many words for one thing, and polysemy, one word with many meanings. They cite research showing that “two people choose the same main key word for a single well-known object less than 20% of the time”[2].

Definition · Latent semantic indexing (LSI)
A retrieval method that builds a table of which words appear in which documents, then compresses it with a technique called singular value decomposition into about 100 “factors”. Documents and queries become points in that smaller space, so a query can find a relevant document even when the two share few words.

The paper states the payoff plainly: “terms that did not actually appear in a document may still end up close to the document”. LSI can also measure how similar two terms are inside the collection[2]. But what it produced was a ranking of documents for a query. It was not a list of words for a writer to add, and the word relationships it found belonged to the one collection it was built from.

Even 'semantic' meant less than it sounds
A footnote in the paper says that by “semantic structure” the authors meant “only the correlation structure in the way in which individual words appear in documents”[2]. LSI counted which words turn up together in which documents. That was the whole of its “semantics”.

What the 1990 tests found

The authors tested LSI on two standard collections of titles and abstracts. On MED, 1,033 medical abstracts with 30 test queries, LSI’s average precision was .51 against .45 for plain word matching, which they called “a 13% improvement over raw term matching”. On CISI, 1,460 information science abstracts with 35 queries, both methods averaged .11. There, the paper says, the latent structure was “no more useful than raw term overlap”[2]. A promising result on one collection, at a scale of about a thousand documents.

The limits its authors listed

  • Words with several meanings. LSI “offers only a partial solution to the polysemy problem”. A word like “bank” becomes one averaged point, which “may create a serious distortion”[2].
  • Size. “Computational constraints have limited us to around 7000 terms”[2].
  • Change. The paper describes a way to “fold-in” new documents, but adds that “how much of this updating can be done without having to perform a new decomposition is unknown”[2].

None of this makes LSI a bad idea. It was careful early research, and its authors said what they did not know. It was built for fixed collections. As Slawski put it, the patent “doesn’t discuss how a process such as this could handle something the size of the Web because nothing that size had quite existed yet then”[6].

From the LSI patent to Google's language systems
Sep 1988
Patent filed
Bell Communications Research files for 'computer information retrieval using latent semantic structure'.
Jun 1989
Patent granted
US patent 4,839,853 is granted, naming seven inventors.
Sep 1990
The paper
'Indexing by Latent Semantic Analysis' is published, tested on 1,033 and 1,460 abstracts.
Sep 2008
Patent expires
The anticipated expiration date listed by Google Patents.
2015
RankBrain
Google's first deep learning system in Search, to understand how words relate to concepts.
2018
Neural matching
Google starts matching fuzzier representations of concepts in queries and pages.
Jul 2019Mueller
No such thing
John Mueller of Google: there's no such thing as LSI keywords.
Oct 2019
BERT in Search
Google says BERT will help it understand one in 10 English searches in the US.
May 2021
MUM announced
Google's ranking guide now says MUM is not used for general ranking.
Jan 2023Mueller
Still wrong
Mueller: anyone telling you to use LSI keywords is still wrong after all these years.
Google Patents, JASIS, Google, Search Engine Roundtable

Sources[1][2][9][3][10][11][8][4].

Does Google use LSI keywords?

No, according to Google. John Mueller of Google posted on Twitter in July 2019: “There’s no such thing as LSI keywords”, adding that “anyone who’s telling you otherwise is mistaken, sorry”[3]. In January 2023 someone asked him whether LSI keywords work better in headings or in body text. He replied that both have no effect, and that “Anyone who tells you to use LSI keywords is ... still wrong after all these years”[4]. Search Engine Journal also quotes him saying “we have no concept of LSI keywords. So that’s something you can completely ignore”[5].

What this does and does not prove
These are statements by a Google employee, reported by trade press. They are not a page in Google’s documentation. Google does not publish a list of methods it does not use, so nobody outside Google can prove a negative. Search Engine Journal put it carefully: “there is no evidence that Google has ever used LSI to rank results”[5]. What Google does publish is a description of its own language systems, below. LSI is not among them.

What Google says it uses instead

Google has described, in its own words, how it gets past exact word matching. Three points matter for writers.

Matching words still counts. “The most basic signal that information is relevant is when content contains the same keywords as your search query”[7]. Use the words of the question you are answering.

Synonyms are handled for you. Google describes a “sophisticated synonym system” that finds relevant pages “even if they don’t contain the exact words you used”. Its example: a search for “change laptop brightness” can match a manufacturer’s page that says “adjust laptop brightness”[7].

Repetition is not relevance. “When you search for ‘dogs,’ you likely don’t want a page with the word ‘dogs’ on it hundreds of times.” Instead, Google says, its algorithms look for “other relevant content beyond the keyword”, such as pictures of dogs, videos or a list of breeds[7].

Google's language systems, in Google's words
Shortened from Google's ranking systems guide and its 2022 post on AI in Search. Toggle columns.
What Google says it doesIn Search since
RankBrainUnderstands how words are related to concepts, so pages can match without every exact word2015
Neural matchingUnderstands representations of concepts in queries and pages, and matches them2018
BERTUnderstands how combinations of words express different meanings and intent2019
Passage rankingIdentifies individual sections of a page to judge how relevant the page isNot dated in the guide
MUMUnderstands and generates language. Not currently used for general rankingAnnounced 2021; specific uses only

BERT is the clearest case. When Google brought it to Search in 2019, it said BERT would help it understand “one in 10 searches in the U.S. in English”, and that people often type “keyword-ese” instead of asking naturally[10]. By 2022, Google said BERT “plays a critical role in almost every English query”[9]. MUM, announced in 2021 as “1,000 times more powerful than BERT”[11], is “not currently used for general ranking in Search”[8].

Notice what none of these systems asks of a writer: a list of related words. RankBrain is described as returning relevant content “even if it doesn’t contain all the exact words used in a search”[8]. Google’s SEO Starter Guide makes the same point: “don’t worry if you don’t anticipate every variation of how someone might seek your content”, because its “language matching systems are sophisticated”. The same guide warns that “Excessively repeating the same words over and over (even in variations) is tiring for users”, and that keyword stuffing is against Google’s spam policies[12]. “Even in variations” is a polite description of a sprinkled LSI list. Our post on keyword stuffing covers the policy and the density myth.

What the LSI advice gets right, and what it gets wrong

Most LSI keyword advice is a good idea with a wrong explanation. The wrong explanation then leads to the wrong tactic: a list to work in, instead of a topic to cover.

LSI keyword advice, claim by claim
Each row is sourced in the sections above and below. Toggle columns.
Holds up?WhyDo this instead
Google uses LSI to understand pages✕NoMueller says there is no such thing as LSI keywords. Google's docs name other systemsDrop the label, keep the reader
Use synonyms and natural variations~Readers use different words, like 'charcuterie' and 'cheese board'. Google says you need not catch every variationUse the words your readers use, where they read naturally
Cover related subtopics✓YesGoogle looks for relevant content beyond the keyword, and its AI features search across subtopicsAnswer the questions a reader asks next
Name related entities✓YesPrecise names tell search engines and models which thing you meanName products, people, places and versions exactly
Add a generated LSI keyword list✕NoRepeating words, even in variations, is what Google's starter guide warns againstTurn the list into questions, then write answers
Use each term a set number of times✕NoGoogle's own example: a page with 'dogs' on it hundreds of times is not what searchers wantCheck which phrases dominate, then cut

The synonym row deserves one more line. Google’s starter guide notes that “some users might search for ‘charcuterie’, while others might search for ‘cheese board’”, and that writing with those differences in mind “could produce positive effects”[12]. That is advice to know your readers, not to collect every variant.

Myth
LSI keywords are how Google understands context.
What is true
Google's documentation describes a synonym system, RankBrain, neural matching and BERT. It does not mention LSI, and John Mueller says there is no such thing as LSI keywords.
Myth
The more related terms on a page, the better it ranks.
What is true
Related terms are a side effect of covering a topic, not a cause. Google warns against repeating words 'even in variations'.
Myth
AI search needs every keyword variation on the page.
What is true
Google says its AI systems understand synonyms, so you do not have to capture every variation of how someone might search.

Entities and topic coverage

If LSI keywords are the wrong idea, two ideas do the real work: entities and coverage.

An entity is a specific thing, such as a company, a product, a person or a place. “Python” could be a programming language or a snake. A page that says which one it means, and names the things around it precisely, leaves less to guess for a search engine or a model. Our entity SEO guide covers how to make your own company one clear entity across the web.

Coverage is the page-level half. Google’s helpful content guidance asks: “Does the content provide a substantial, complete, or comprehensive description of the topic?”[13]. Its “dogs” example makes the same point from the other side: pictures, videos, a list of breeds[7]. A page that covers a topic fully will contain related words, because you cannot explain a topic without them. A page that contains related words has not necessarily covered anything. That is the flaw in LSI keyword lists: they measure the side effect and miss the cause. For coverage across a whole site, see topical authority for AI answers.

Related words are a side effect of covering a topic. Adding them does not create the coverage. Answering the next question does.

How AI answer engines pick passages

Start with what is not known. No AI company publishes, in any detail, how it picks the passage it quotes or the page it cites. Any claim that an engine rewards a set of related keywords is a guess. What is documented is how some engines search, and it points the same way as everything above.

Google says AI Overviews and AI Mode may use a “query fan-out” technique, issuing “multiple related searches across subtopics and data sources” to build a response[14]. Its AI guide gives an example. For “how to fix a lawn that’s full of weeds”, fan-out queries might include “best herbicides for lawns”, “remove weeds without chemicals” and “how to prevent weeds in lawn”. The same guide says “AI systems can understand synonyms and general meanings”, so “you don’t have to worry that you don’t have enough ‘long-tail’ keywords”. And there is no need to cut content into tiny pieces, because “Google systems are able to understand the nuance of multiple topics on a page and show the relevant piece to users”[15].

OpenAI describes something similar for ChatGPT. When it searches through partner search providers, “ChatGPT search typically rewrites your query into one or more targeted queries”, and after reviewing the first results it may send “additional, more specific queries”[16].

Our reading, and it is an inference: each of those follow-up searches is a subtopic. A page that answers the main question and the obvious next ones, each in a clearly headed section, gives an engine something to retrieve for more of them. A page padded with related terms gives it nothing new to quote. For what is known, inferred and marketed about AI citations, see how AI assistants decide which sources to cite.

A topic coverage method that replaces LSI keyword lists

Every step below produces questions, not words to sprinkle. Work through them in order.

Planning topic coverage
Click any step to stop the animation and read it.
Name the reader and the question
Who is searching, and what do they need to do after reading? Write the main question in their words, not yours.
Step 1: Name the reader and the question. Who is searching, and what do they need to do after reading? Write the main question in their words, not yours.
Step 2: Collect the next questions. Look at the People Also Ask box for the main query and its variants, at autocomplete, and at the queries your page already appears for in Search Console.
Step 3: Read the top pages' headings. List the subtopics the top results cover. Look for gaps and for questions they answer badly. Do not copy their outline.
Step 4: Ask the people who buy. Sales calls, support tickets, reviews and forum threads hold the questions searchers type, in the words they use.
Step 5: Map every question. Each question gets a section on this page, a link to the page that owns it, or a decision to leave it out.
Step 6: Write, then check what dominates. Name entities precisely and use your readers' words. Then check which phrases the page repeats most, and cut any you repeated for a machine.

Search Console’s Performance report shows “what search queries are most likely to show your site”[17]. Google’s AI guide advises: “Don’t just recycle what others on the internet have already said”[15].

Two steps deserve more detail. For the second, our People Also Ask guide includes a planner that turns those questions into a page plan. For the third, you are reading competitor headings for the list of subtopics, not for wording. Your own data, examples and opinions are what make a section worth quoting, and they are the part a competitor cannot give you.

A worked example

Take Google’s own fan-out example. A page answering “how to fix a lawn that’s full of weeds” might plan these sections:

  • Which weeds you have, and why it changes the fix.
  • Herbicides that are safe for lawns.
  • Removing weeds without chemicals.
  • Stopping the weeds from coming back.
  • When to reseed the bare patches.

The middle three come from the fan-out queries in Google’s guide[15]. The first and last are our additions, the kind of next question a reader asks. A related-terms list for the same page might suggest words like “dandelion”, “turf” or “broadleaf”. Each will appear on its own once the sections are written, because you cannot explain the fix without them.

What an LSI keyword generator actually gives you

An LSI keyword generator is a related-terms tool with a borrowed name. Depending on the tool, the list may come from related searches, autocomplete, or words that appear on top-ranking pages. Slawski’s complaint about one such site applies widely: it gave no “information about how they generate those keywords”[6].

Used with care, a list can still help you brainstorm. Read each term and ask one question: would a reader expect this page to explain it? If yes, it points to a missing section, so write that section. If not, leave it out. Never paste the list into the page, and never aim to use each term a set number of times.

Check which phrases dominate your page

After writing, check the page from the other side. A page written around a keyword list tends to lean on a few phrases. The free readability checker has a keyword density tab that shows which words and phrases a page repeats most. If one phrase dominates and sounds forced when read aloud, that is the repetition Google’s starter guide warns about. The same tool scores reading ease and flags passive voice; see Flesch-Kincaid explained and active and passive voice for what those numbers mean.

Before you publish: coverage, not keyword lists
0/7

Finding coverage gaps with CoreCited

The method above starts from search questions. CoreCited’s Content Briefs start from AI answers: a question your buyers ask, where an engine names competitors and not you. The brief shows which sources the engines cited and what those pages cover that yours does not. Starter includes 10 briefs a month, and every limit is on the pricing page.

Questions people ask

What are LSI keywords?

An SEO name for words related to a target keyword: synonyms, related topics and terms that often appear alongside it. The name borrows from latent semantic indexing, a retrieval method patented in 1989. The lists sold as LSI keywords are not that method, and Google's John Mueller has said there is no such thing as LSI keywords.

Does Google use LSI keywords?

No, according to Google's John Mueller. He said in 2019 that there is no such thing as LSI keywords, and in 2023 that anyone telling you to use them is still wrong. Google's own documentation describes a synonym system, RankBrain, neural matching and BERT, and does not mention LSI.

What is latent semantic indexing?

A document retrieval method from researchers at Bell Communications Research, patented in 1989 and published in 1990. It builds a table of which words appear in which documents, compresses it into about 100 factors, and matches queries to documents in that compressed space, so a document can match a query even when they share few words.

Is an LSI keyword generator worth using?

Only as a brainstorming list. The terms are related words under a borrowed name, and they are a side effect of good coverage, not a cause of it. Read each term, ask whether a reader would expect the page to explain it, and write that section if so. Never paste the list into the page.

What should I do instead of using LSI keywords?

Cover the topic. Answer the main question first, then the questions a reader asks next, which you can find in People Also Ask, Search Console and conversations with buyers. Use your readers' words, name entities precisely, and check that no single phrase dominates the page.

Do LSI keywords help in ChatGPT or Google AI Overviews?

No AI company says they do. Google says its AI systems understand synonyms, so you do not need every keyword variation, and that AI Overviews and AI Mode may run several related searches across subtopics. ChatGPT search rewrites a prompt into its own search queries. Covering those subtopics clearly is what gives an engine more to work with.

Keep reading

Sources

[1]US4839853A: Computer information retrieval using latent semantic structure — Google Patents (original assignee Bell Communications Research), filed 15 September 1988, granted 13 June 1989
[2]Indexing by Latent Semantic Analysis — Deerwester, Dumais, Furnas, Landauer and Harshman, Journal of the American Society for Information Science 41(6), pages 391 to 407 (copy hosted by CSU Stanislaus), September 1990
[3]Google: There Is No Such Thing As LSI Keywords — Search Engine Roundtable (Barry Schwartz, reporting John Mueller), 31 July 2019
[4]Google: LSI Keywords Have No Effect Again & Again — Search Engine Roundtable (Barry Schwartz, reporting John Mueller), 4 September 2023
[5]Latent Semantic Indexing (LSI): Is It A Google Ranking Factor? — Search Engine Journal (Miranda Miller), 23 January 2022
[6]Does Google Use Latent Semantic Indexing (LSI)? — SEO by the Sea (Bill Slawski), 22 January 2018, updated 7 March 2022
[7]How Does Google Determine Ranking Results — Google Search, How Search Works, read 2 October 2026
[8]A guide to Google Search ranking systems — Google Search Central, updated 10 December 2025
[9]How AI powers great search results — Google (Pandu Nayak), 3 February 2022
[10]Understanding searches better than ever before — Google (Pandu Nayak), 25 October 2019
[11]MUM: A new AI milestone for understanding information — Google (Pandu Nayak), 18 May 2021
[12]SEO Starter Guide — Google Search Central, updated 10 December 2025
[13]Creating helpful, reliable, people-first content — Google Search Central, updated 1 October 2026
[14]AI features and your website — Google Search Central, updated 10 December 2025
[15]Optimizing your website for generative AI features on Google Search — Google Search Central, updated 10 July 2026
[16]Searching the web with ChatGPT — OpenAI Help Center, read 2 October 2026
[17]Performance report (Search results) — Search Console Help, read 2 October 2026

Find out where you actually stand

One real question, real AI engines, and the answer they gave — including who was named in it. No account, no card.