Get a Free Quote

The Japanese keyword hack in WordPress: how to find it and clear it

Your site loads normally. Nothing in the admin looks wrong. But search for your domain and Google lists hundreds of pages in Japanese selling branded goods, under your name. This is the Japanese keyword hack, and the reason it is so often found late is that it is deliberately invisible to you: the injected pages are served to search engine crawlers and hidden from ordinary visitors. This guide explains the mechanism, how to confirm it in a couple of minutes, where the code actually lives, and how to get the fake pages out of the index afterwards. Everything here is diagnostic and non-destructive until the removal section, which says plainly what to back up first.

What the hack is

The Japanese keyword hack injects thousands of auto-generated pages into your site, almost always in Japanese and almost always advertising counterfeit branded goods. The pages use your domain's reputation to rank, and the traffic is monetised by whoever placed them.

Three features make it different from ordinary defacement, and all three explain why it tends to run undetected for weeks.

  • It is cloaked. The spam is served to search engine crawlers and hidden from ordinary visitors, including you.
  • It scales. It does not create one page, it creates a URL pattern that generates pages on demand, so the index fills up quickly.
  • It registers itself. It usually adds its own sitemap, so Google is actively invited to crawl the spam rather than having to discover it.

The usual way people find out is a Search Console message, a sudden spike in indexed pages, or a customer asking why the site shows Japanese text in Google.

Confirming it in two minutes

Do this before changing anything. Several other problems look similar from the outside and the fixes are completely different.

1. Ask Google what it has

Search site:example.com with your own domain. Then narrow it:

site:example.com 財布
site:example.com ブランド
site:example.com -inurl:www

Japanese results under your domain, for pages you never wrote, is the first signal. In Search Console, open the Pages report and sort indexed URLs: an infection shows as a large block of URLs sharing a pattern you do not recognise.

2. Request a page as Googlebot

This is the decisive test. Take one of the spam URLs and fetch it twice.

# as an ordinary visitor
curl -s "https://example.com/the-spam-url/" | head -40

# as Googlebot
curl -sA "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)" \
  "https://example.com/the-spam-url/" | head -40

If the two responses differ, the site is cloaking and the diagnosis is settled. The Googlebot response will usually contain Japanese text and outbound links. You can also use the URL Inspection tool in Search Console and view the crawled HTML, which has the advantage of being Google's own fetch rather than a claim about it.

3. Look for a sitemap you did not make

curl -s https://example.com/robots.txt
curl -sI https://example.com/sitemap.xml

Read robots.txt carefully. An injected Sitemap: line pointing at a file you do not recognise, often with a random name, is how the spam gets crawled so fast. Fetch that file and you will usually find thousands of the fake URLs listed in it.

Thinking about a new WordPress website?Get a free consultation and a fixed-scope quote. A senior engineer replies within 24 hours. No obligation.
Get a Free Quote

How it works, which is why the obvious fixes fail

Understanding the mechanism saves you from the two most common wasted days: deleting pages that were never files, and cleaning the files while leaving the database.

  1. Entry. Almost always a known vulnerability in an outdated plugin or theme, or a reused administrator password. The hack is not exotic; the delivery is.
  2. Persistence. A rogue administrator account is created, so losing the original entry point does not matter.
  3. Generation. Code hooks into WordPress request handling and generates a page for any URL matching a pattern. Nothing is stored as a post, which is why the pages are invisible in your admin.
  4. Cloaking. The code inspects the request and decides what to return. Crawler gets spam, human gets the real site.
  5. Invitation. A rogue sitemap is added and often submitted, so crawling is immediate rather than incidental.

So the fake pages have no files to delete and no posts to trash. Removing them means removing the generator, the persistence and the invitation, in that order.

Finding the mechanism

Work through these in order. They only read and report.

Verify everything that has a published checksum

wp core verify-checksums
wp plugin verify-checksums --all

This catches modified core and repository plugin files instantly. It cannot check premium plugins or themes, so a clean result does not clear those.

Find what changed, and when

find . -name "*.php" -type f -mtime -30 -printf "%TY-%Tm-%Td %TH:%TM  %p\n" | sort -r | head -60

Look for a cluster of files written within the same minute. Check that moment against your own update history. This hack commonly drops files into wp-content/ directories with plausible names, and into the active theme.

Look for the generator and the cloaking logic

The cloaking decision needs the user agent, so the code almost always references it.

grep -rEn --include="*.php" \
  "HTTP_USER_AGENT|HTTP_REFERER|googlebot|bingbot" ./wp-content ./wp-includes ./wp-admin | head -60

Plenty of legitimate code reads the user agent, so read the matches rather than counting them. What you want is a check for crawler names combined with output of unrelated content, or sitting inside a file whose name and location make no sense.

# obfuscation, which the payload usually relies on
grep -rEn --include="*.php" \
  "eval\(|base64_decode\(|gzinflate\(|str_rot13\(|create_function\(" ./wp-content | head -60

# PHP where it should never execute
find ./wp-content/uploads -name "*.php*" -type f

Check the database, which is where it usually survives

# administrators, with registration dates
wp user list --role=administrator --fields=ID,user_login,user_email,user_registered

# scheduled events you cannot account for
wp cron event list --fields=hook,next_run_relative,recurrence

# injected markup in options
wp db query "SELECT option_name, LEFT(option_value,120) FROM wp_options
  WHERE option_value REGEXP '<script|base64_decode|eval\\(' LIMIT 50;"

Use your real table prefix, which is set in wp-config.php and is often not wp_. The scheduled-event check is the one most often skipped, and a rogue cron entry is the usual reason a cleaned site reinfects on a timer.

Check what is autoloaded

Some variants store their payload in an autoloaded option, so it runs on every request without any file looking unusual.

wp db query "SELECT option_name, LENGTH(option_value) AS bytes FROM wp_options
  WHERE autoload='yes' ORDER BY bytes DESC LIMIT 25;"

An enormous autoloaded option with a name you do not recognise deserves a close look. The same query is useful for ordinary performance work, which we cover in wp_options autoload bloat.

Removing it

Take a full backup first, files and database, even though the site is compromised. You will want to refer back to it, and a clean removal sometimes takes two attempts. Keep it offline.

  1. Put the site into maintenance or restrict access while you work, so you are not cleaning a moving target.
  2. Replace core. Do not clean it file by file. wp core download --force overwrites core with the official release and keeps your content.
  3. Reinstall every plugin and theme from source. Replace rather than repair. Delete anything you do not actively use, because an inactive plugin is still a file an attacker can reach.
  4. Remove the injected files you identified, including anything executable in uploads.
  5. Clean robots.txt and delete the rogue sitemap file.
  6. Delete rogue administrators and demote any account that should not have the role.
  7. Remove rogue scheduled events with wp cron event delete.
  8. Rotate everything. New salts and keys, new database password, new admin passwords, new hosting and FTP credentials. Rotating salts logs every session out, which is the point.
# official core, over the top, content untouched
wp core download --force --skip-content

# new salts and keys, which invalidates every existing session
wp config shuffle-salts

# re-check once you are done
wp core verify-checksums && wp plugin verify-checksums --all

Then repeat the Googlebot fetch from the confirmation step. If the crawler response now matches what a visitor sees, the cloaking is gone. If it does not, something is still generating pages and you have missed a file or a database entry.

Ready to bring your WordPress project to life?Get a free consultation and a fixed-scope quote. A senior engineer replies within 24 hours. No obligation.
Get a Free Quote

Getting the fake pages out of search

Removing the infection does not remove the URLs from Google's index. Until Google recrawls them, your domain still shows Japanese spam, and this is the part that damages a business.

The important decision is what those URLs now return. A 404 says not found; a 410 says deliberately gone, which crawlers treat as a stronger and faster signal.

  • Return 410 for the spam pattern. If the fake URLs share a path pattern, return 410 for it rather than letting them 404 quietly.
  • Do not redirect them to your homepage. It looks like an attempt to keep the equity, it confuses the signal, and it slows removal.
  • Submit your real sitemap and confirm the rogue one is gone and returns 404 or 410.
  • Use Removals in Search Console for the worst offenders. It is temporary, roughly six months, but it clears them from results while recrawling catches up.
  • Request a security review if Search Console flagged the site, once you are genuinely clean. Requesting one while still infected resets the clock.

Expect weeks rather than days for the index to settle, and watch the Pages report rather than guessing. Impressions for the spam URLs falling to zero is the signal that it is finished.

Staying clean

Reinfection is common and it is almost always one of three things: a backdoor that was missed, a credential that was not rotated, or the original vulnerability still being present because a plugin was repaired rather than replaced.

MeasureWhy it matters here specifically
Update plugins and themes promptlyThis hack enters through known vulnerabilities, not novel ones
Delete unused plugins and themesAn inactive plugin is still reachable code
Block PHP execution in uploadsRemoves the most common place a dropper lands
Disable file editing in the adminDISALLOW_FILE_EDIT stops a stolen login becoming code execution
Unique admin passwords with two-factorCredential reuse is the other entry route
Watch indexed page count in Search ConsoleA sudden jump is the earliest signal you will get
Keep off-site backups with historyA single overwritten backup is worthless once the infection is a week old

Add this to wp-config.php above the line telling you to stop editing:

define( 'DISALLOW_FILE_EDIT', true );

The monitoring line is the one that pays for itself. Everything about this hack is designed so you do not notice, so the only reliable early warning is the indexed page count, and it is free to watch.

If you have removed it twice and it keeps returning, something is still present and continuing to guess is expensive. We clean compromised WordPress sites, establish how the entry happened, and harden what allowed it. Send us the domain and we will tell you what we find before you commit to anything, and our maintenance and support plans exist to keep the updates from falling behind again.

Hamza Hai

Hamza Hai writes about WordPress development, performance, and growth for businesses.

FAQ

Frequently asked questions

A search spam compromise that injects large numbers of auto-generated Japanese-language pages into your site, usually selling counterfeit goods, and serves them only to search engine crawlers. Visitors see your normal site, so it is typically discovered through Search Console or a customer mentioning odd search results.

The payload checks the request before deciding what to serve, usually the user agent, the referrer, and whether the visitor has a login cookie. Crawlers get the spam, you get the real page. This is called cloaking, and it is why you must request the page as Googlebot to see what Google sees.

Run a site: search for your domain, check the Pages report in Search Console for indexed URLs you never created, and fetch one of those URLs with a Googlebot user agent. If the Googlebot response differs from your browser's, it is confirmed.

No, because the pages do not exist as content. They are generated on request by injected code, and often registered through a rogue sitemap. Until you remove the code and the rogue administrator accounts, they regenerate.

Removal is usually a day's work. Getting the fake URLs out of the index takes longer, typically weeks, because Google has to recrawl them and see that they are gone. Returning a proper 410 speeds that up.

Usually not. A verified core, clean plugin and theme copies, a cleaned database and rotated credentials are normally enough. Rebuild when you cannot establish what was changed, or when the same infection returns after a careful clean.

Have a project?

Let's Build Your Next WordPress Website

Get a free consultation and a fixed-scope quote. No obligations.

  • Free Consultation
  • No Hidden Costs
  • 100% Confidential

Request your free quote

Tell us what you are building. A senior engineer replies within 24 hours.

Please enter your name.

Please enter a valid email address.

Please tell us a little more about your project (10+ characters).

No obligation. Your details are only used to prepare your quote.