What the hack is
The Japanese keyword hack injects thousands of auto-generated pages into your site, almost always in Japanese and almost always advertising counterfeit branded goods. The pages use your domain's reputation to rank, and the traffic is monetised by whoever placed them.
Three features make it different from ordinary defacement, and all three explain why it tends to run undetected for weeks.
- It is cloaked. The spam is served to search engine crawlers and hidden from ordinary visitors, including you.
- It scales. It does not create one page, it creates a URL pattern that generates pages on demand, so the index fills up quickly.
- It registers itself. It usually adds its own sitemap, so Google is actively invited to crawl the spam rather than having to discover it.
The usual way people find out is a Search Console message, a sudden spike in indexed pages, or a customer asking why the site shows Japanese text in Google.
Confirming it in two minutes
Do this before changing anything. Several other problems look similar from the outside and the fixes are completely different.
1. Ask Google what it has
Search site:example.com with your own domain. Then narrow it:
site:example.com 財布
site:example.com ブランド
site:example.com -inurl:wwwJapanese results under your domain, for pages you never wrote, is the first signal. In Search Console, open the Pages report and sort indexed URLs: an infection shows as a large block of URLs sharing a pattern you do not recognise.
2. Request a page as Googlebot
This is the decisive test. Take one of the spam URLs and fetch it twice.
# as an ordinary visitor
curl -s "https://example.com/the-spam-url/" | head -40
# as Googlebot
curl -sA "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)" \
"https://example.com/the-spam-url/" | head -40If the two responses differ, the site is cloaking and the diagnosis is settled. The Googlebot response will usually contain Japanese text and outbound links. You can also use the URL Inspection tool in Search Console and view the crawled HTML, which has the advantage of being Google's own fetch rather than a claim about it.
3. Look for a sitemap you did not make
curl -s https://example.com/robots.txt
curl -sI https://example.com/sitemap.xmlRead robots.txt carefully. An injected Sitemap: line pointing at a file you do not recognise, often with a random name, is how the spam gets crawled so fast. Fetch that file and you will usually find thousands of the fake URLs listed in it.
How it works, which is why the obvious fixes fail
Understanding the mechanism saves you from the two most common wasted days: deleting pages that were never files, and cleaning the files while leaving the database.
- Entry. Almost always a known vulnerability in an outdated plugin or theme, or a reused administrator password. The hack is not exotic; the delivery is.
- Persistence. A rogue administrator account is created, so losing the original entry point does not matter.
- Generation. Code hooks into WordPress request handling and generates a page for any URL matching a pattern. Nothing is stored as a post, which is why the pages are invisible in your admin.
- Cloaking. The code inspects the request and decides what to return. Crawler gets spam, human gets the real site.
- Invitation. A rogue sitemap is added and often submitted, so crawling is immediate rather than incidental.
So the fake pages have no files to delete and no posts to trash. Removing them means removing the generator, the persistence and the invitation, in that order.
Finding the mechanism
Work through these in order. They only read and report.
Verify everything that has a published checksum
wp core verify-checksums
wp plugin verify-checksums --allThis catches modified core and repository plugin files instantly. It cannot check premium plugins or themes, so a clean result does not clear those.
Find what changed, and when
find . -name "*.php" -type f -mtime -30 -printf "%TY-%Tm-%Td %TH:%TM %p\n" | sort -r | head -60Look for a cluster of files written within the same minute. Check that moment against your own update history. This hack commonly drops files into wp-content/ directories with plausible names, and into the active theme.
Look for the generator and the cloaking logic
The cloaking decision needs the user agent, so the code almost always references it.
grep -rEn --include="*.php" \
"HTTP_USER_AGENT|HTTP_REFERER|googlebot|bingbot" ./wp-content ./wp-includes ./wp-admin | head -60Plenty of legitimate code reads the user agent, so read the matches rather than counting them. What you want is a check for crawler names combined with output of unrelated content, or sitting inside a file whose name and location make no sense.
# obfuscation, which the payload usually relies on
grep -rEn --include="*.php" \
"eval\(|base64_decode\(|gzinflate\(|str_rot13\(|create_function\(" ./wp-content | head -60
# PHP where it should never execute
find ./wp-content/uploads -name "*.php*" -type fCheck the database, which is where it usually survives
# administrators, with registration dates
wp user list --role=administrator --fields=ID,user_login,user_email,user_registered
# scheduled events you cannot account for
wp cron event list --fields=hook,next_run_relative,recurrence
# injected markup in options
wp db query "SELECT option_name, LEFT(option_value,120) FROM wp_options
WHERE option_value REGEXP '<script|base64_decode|eval\\(' LIMIT 50;"Use your real table prefix, which is set in wp-config.php and is often not wp_. The scheduled-event check is the one most often skipped, and a rogue cron entry is the usual reason a cleaned site reinfects on a timer.
Check what is autoloaded
Some variants store their payload in an autoloaded option, so it runs on every request without any file looking unusual.
wp db query "SELECT option_name, LENGTH(option_value) AS bytes FROM wp_options
WHERE autoload='yes' ORDER BY bytes DESC LIMIT 25;"An enormous autoloaded option with a name you do not recognise deserves a close look. The same query is useful for ordinary performance work, which we cover in wp_options autoload bloat.
Removing it
Take a full backup first, files and database, even though the site is compromised. You will want to refer back to it, and a clean removal sometimes takes two attempts. Keep it offline.
- Put the site into maintenance or restrict access while you work, so you are not cleaning a moving target.
- Replace core. Do not clean it file by file.
wp core download --forceoverwrites core with the official release and keeps your content. - Reinstall every plugin and theme from source. Replace rather than repair. Delete anything you do not actively use, because an inactive plugin is still a file an attacker can reach.
- Remove the injected files you identified, including anything executable in uploads.
- Clean robots.txt and delete the rogue sitemap file.
- Delete rogue administrators and demote any account that should not have the role.
- Remove rogue scheduled events with
wp cron event delete. - Rotate everything. New salts and keys, new database password, new admin passwords, new hosting and FTP credentials. Rotating salts logs every session out, which is the point.
# official core, over the top, content untouched
wp core download --force --skip-content
# new salts and keys, which invalidates every existing session
wp config shuffle-salts
# re-check once you are done
wp core verify-checksums && wp plugin verify-checksums --allThen repeat the Googlebot fetch from the confirmation step. If the crawler response now matches what a visitor sees, the cloaking is gone. If it does not, something is still generating pages and you have missed a file or a database entry.
Getting the fake pages out of search
Removing the infection does not remove the URLs from Google's index. Until Google recrawls them, your domain still shows Japanese spam, and this is the part that damages a business.
The important decision is what those URLs now return. A 404 says not found; a 410 says deliberately gone, which crawlers treat as a stronger and faster signal.
- Return 410 for the spam pattern. If the fake URLs share a path pattern, return 410 for it rather than letting them 404 quietly.
- Do not redirect them to your homepage. It looks like an attempt to keep the equity, it confuses the signal, and it slows removal.
- Submit your real sitemap and confirm the rogue one is gone and returns 404 or 410.
- Use Removals in Search Console for the worst offenders. It is temporary, roughly six months, but it clears them from results while recrawling catches up.
- Request a security review if Search Console flagged the site, once you are genuinely clean. Requesting one while still infected resets the clock.
Expect weeks rather than days for the index to settle, and watch the Pages report rather than guessing. Impressions for the spam URLs falling to zero is the signal that it is finished.
Staying clean
Reinfection is common and it is almost always one of three things: a backdoor that was missed, a credential that was not rotated, or the original vulnerability still being present because a plugin was repaired rather than replaced.
| Measure | Why it matters here specifically |
|---|---|
| Update plugins and themes promptly | This hack enters through known vulnerabilities, not novel ones |
| Delete unused plugins and themes | An inactive plugin is still reachable code |
| Block PHP execution in uploads | Removes the most common place a dropper lands |
| Disable file editing in the admin | DISALLOW_FILE_EDIT stops a stolen login becoming code execution |
| Unique admin passwords with two-factor | Credential reuse is the other entry route |
| Watch indexed page count in Search Console | A sudden jump is the earliest signal you will get |
| Keep off-site backups with history | A single overwritten backup is worthless once the infection is a week old |
Add this to wp-config.php above the line telling you to stop editing:
define( 'DISALLOW_FILE_EDIT', true );The monitoring line is the one that pays for itself. Everything about this hack is designed so you do not notice, so the only reliable early warning is the indexed page count, and it is free to watch.
If you have removed it twice and it keeps returning, something is still present and continuing to guess is expensive. We clean compromised WordPress sites, establish how the entry happened, and harden what allowed it. Send us the domain and we will tell you what we find before you commit to anything, and our maintenance and support plans exist to keep the updates from falling behind again.