Multilingual Keyword Research: Do Each Language From Scratch

Multilingual keyword research has exactly one iron law: run a complete, independent study for each language market, and never translate the source language’s keyword list. Different markets use different words, phrase questions differently, care about different sub-topics, and face different competitors. A translated list targets the keywords you assume exist; native research targets the keywords locals are actually searching for.

Why translating a keyword list is doomed to fail

Keywords are fossils of search behavior, and search behavior grows out of language and culture. A translated list steps on at least three landmines:

  1. Wrong synonym. A single concept usually has several phrasings in the target language, and their search volumes can differ 10x. The Traditional Chinese term for “search engine optimization” is 搜尋引擎優化; in Simplified Chinese it is 搜索引擎优化. Translate “SEO audit” literally into Chinese and you miss that most people actually search for “SEO 健檢” (SEO health check) or “網站健檢” (site check-up).
  2. Different sub-topic distribution. Under the same theme, each market cares about different questions. On multilingual site architecture, the English-speaking world discusses ccTLDs heavily (lots of multinationals), while the Chinese-speaking world asks far more about Traditional/Simplified cannibalization — a translated list makes you write a pile of articles nobody local is asking about.
  3. Different competitive terrain. A keyword that is a bloodbath in English may be almost unwritten in Traditional Chinese. Only separate research surfaces each market’s low-competition entry points.

The four-step process: run one pass per language

Step 1: Define the “market,” not the “language”

Language ≠ market. Simplified Chinese alone splits into at least two market decisions: Mainland China runs on Baidu (Google is unavailable there) — a completely different game — while overseas Simplified Chinese (Singapore/Malaysia, North American Chinese communities) runs on Google/Bing. Traditional Chinese mostly maps to Google in Taiwan and Hong Kong. Write down “who does this language version serve, and what engine do they use” first — that’s what makes you pick the right tools and data sources later.

Step 2: Native brainstorming plus field observation

Step 3: Validate and expand with tools

Feed your brainstormed seed terms into the tools (Google Keyword Planner, Ahrefs, Semrush, etc., and remember to set the region and language to the target market) to validate volume and expand the long tail. Note: tool data for Chinese keywords is generally coarser than for English — treat volume as a relative reference and don’t agonize over absolute numbers.

Step 4: Build the intent matrix for that language

Cross-expand a list of micro-intents for the language with [topic] × [action] × [context], tagging each as informational (I) or commercial (C), and mapping it to that language’s Boss page. The point is breadth first: the job of a long-tail article is to cover the topic map and accumulate entity authority, not to win volume on a single post. Maintain each language’s matrix independently — never translate one into another.

Traditional vs Simplified worked example: same theme, different list

Theme Terms a Traditional Chinese matrix might use Terms a Simplified Chinese matrix might use
Software reviews 軟體推薦、工具比較 软件推荐、工具对比
Video SEO 影片 SEO、YouTube 排名 视频 SEO、视频排名
Site health check SEO 健檢、網站體檢 SEO 诊断、网站优化分析

Look at the third row: it isn’t just different characters — even the conceptual packaging differs (“健檢/health check” vs “诊断/diagnosis”). Only native research catches a gap like that.

Three things to line up after the research

  1. Every language’s title, slug, and meta use terms from that language’s matrix — never shared.
  2. hreflang correctly pairs each language version, so you don’t do the research right only to have traffic served the wrong version.
  3. Track query strings by language directory in GSC, verify whether the queries you actually win are local phrasings, and revise the matrix each quarter.

Frequently asked questions (FAQ)

Q1: Can I do keyword research for a language I don’t speak? You can operate the tools, but seed terms and intent judgments need native input — at minimum, get a native speaker for 2–3 hours of terminology review and phrasing interviews, or the whole matrix gets built on a translated foundation.

Q2: Do the article topics across the three languages need to map one-to-one? No. Each language’s matrix will share core themes, but every market should also have articles unique to it (answering questions unique to that market). Forcing a one-to-one mapping means abandoning each market’s unique opportunities.

Q3: Chinese keyword tools return almost no data — what do I do? Accept the coarseness and read relative signals instead: Google autocomplete, related searches, competitor titles, forum question frequency. Long-tail Chinese terms often show 0 volume in tools yet get real searches, and the cost to cover them is low — write more, not less.

Q4: How often should I rerun keyword research for each language? Make small quarterly additions to each matrix (mine new terms back from actual GSC queries), and do one full rerun per year, adjusting the topic layout to match market and algorithm shifts.


Once the keyword research is done for each language, make sure the technical layer keeps up: GeoSeoToday’s hreflang generator and GEO checker ensure every language version is served to the right market. Related reading: localization vs translation (why you can’t just ship machine translation) and Traditional vs Simplified Chinese SEO (it’s not just character conversion); for the full cluster overview, see the pillar guide on multilingual and international SEO.