How to Block AI Crawlers from Scraping Your WordPress Site: The Complete Technical Guide

A practical guide to blocking AI crawlers on WordPress - robots.txt rules, .htaccess IP bans, and server-level controls to stop GPTBot, ClaudeBot, and others.

The Wake-Up Call: When Meta’s Crawler Scraped a Site 7.9 Million Times

A few months ago, a web developer posted on Reddit’s r/webdev with a story that sent shockwaves through the WordPress community. Meta’s AI crawler - meta-externalagent - had scraped their website 7.9 million times in 30 days. The result: over 900 GB of bandwidth consumed and server logs bloated beyond recognition, all before the site owner even noticed anything was wrong.

The post collected 527 upvotes and 80 comments. The comment section was a mix of horror, shared experiences, and frantic questions about how to stop it. One commenter described finding similar patterns from multiple AI bots simultaneously. Another calculated the hosting cost impact and the number was not trivial.

This is not an isolated incident. It is the new normal. AI companies are crawling the web at unprecedented scale and speed, and most WordPress site owners have no idea it is happening to them right now.

7.9 Million

requests from Meta’s AI crawler in just 30 days on a single website

What Are AI Crawlers and Why Are They Different?

To understand why AI crawlers are a different beast from traditional search engine bots, you need to understand what they are doing and how they do it.

Traditional Search Crawlers vs. AI Crawlers

Traditional search engine crawlers like Googlebot and Bingbot visit your site to index pages for search results. They follow your robots.txt directives, respect crawl rate limits, and generally behave like polite visitors. Their goal is to understand your site’s structure and content well enough to rank it in search results.

AI crawlers have a fundamentally different objective: they are collecting training data for large language models or feeding real-time information into AI assistants. This difference in purpose leads to dramatically different behavior:

  • Volume: AI crawlers often make 10-100x more requests than traditional search crawlers. They do not just want to index your pages - they want to ingest every piece of text on every page, often repeatedly.
  • Depth: Traditional crawlers prioritize important pages and often skip deep archives. AI crawlers tend to be more thorough, crawling everything they can reach.
  • Frequency: Some AI crawlers revisit pages much more frequently than search engine bots, looking for updated content to feed into their models.
  • Compliance: While major AI companies claim to respect robots.txt, enforcement is inconsistent and many smaller AI crawlers ignore it entirely.

900+ GB

of bandwidth consumed by a single AI crawler in one month

The AI Crawler Landscape: Who Is Crawling Your Site?

Here is a comprehensive list of the major AI crawlers you need to know about, what they do, and who operates them.

Meta

  • meta-externalagent - Meta’s primary AI training crawler. This is the bot that hit 7.9 million requests in the Reddit story. Used to collect training data for Meta’s Llama models and other AI products.
  • FacebookBot - While primarily used for link previews in Facebook/Instagram, it also feeds into Meta’s broader data pipeline.

OpenAI

  • GPTBot - OpenAI’s crawler for collecting training data for GPT models. OpenAI has published guidance on blocking it via robots.txt and claims to respect these directives.
  • ChatGPT-User - A separate user agent used when ChatGPT browses the web in real-time to answer user queries. Blocking this prevents your content from being cited in ChatGPT’s live responses.

Google

  • Google-Extended - Google’s dedicated AI training crawler, used to collect data for Gemini and other AI products. Crucially, this is separate from Googlebot - blocking Google-Extended does not affect your Google Search rankings.

Anthropic

  • ClaudeBot - Anthropic’s crawler for collecting training data for Claude models. Anthropic publishes information about its crawler and claims to respect robots.txt.
  • anthropic-ai - An older user agent string that Anthropic has used.

Other Notable AI Crawlers

  • CCBot - Common Crawl’s bot. Common Crawl is a nonprofit that creates open web datasets, widely used by AI companies for model training.
  • Bytespider - ByteDance’s crawler, used for TikTok and their AI products. Known for aggressive crawling patterns.
  • Amazonbot - Amazon’s crawler for Alexa and AI training data.
  • Applebot-Extended - Apple’s AI training crawler, separate from the regular Applebot used for Siri and Spotlight.
  • PerplexityBot - Perplexity AI’s crawler for their AI-powered search engine.
  • cohere-ai - Cohere’s crawler for training their enterprise AI models.
⚠️ Important Distinction: Some crawlers have “extended” variants (Google-Extended, Applebot-Extended) that are specifically for AI training. Blocking these does NOT affect your regular search engine indexing. However, blocking the base crawlers (Googlebot, Applebot) WILL affect search results.

The Real Cost: What AI Crawlers Do to Your WordPress Site

AI crawler impact goes beyond simple annoyance. For WordPress sites - especially those on shared hosting or metered bandwidth plans - the consequences can be severe and measurable.

Bandwidth and Hosting Costs

The Reddit case study is the most dramatic example, but even less extreme crawling can add up. Consider a typical WordPress site on shared hosting:

  • A standard blog post page might be 500 KB-2 MB including images, CSS, and JavaScript
  • If an AI crawler hits 100,000 pages in a month (modest by AI crawler standards), that is 50-200 GB of bandwidth
  • Many shared hosting plans include only 10-50 GB of bandwidth per month
  • Overage charges or throttling can result, directly impacting your site’s availability

For sites on metered cloud hosting (AWS, DigitalOcean, Google Cloud), every byte of bandwidth has a direct cost. An aggressive AI crawler can add hundreds of dollars to your monthly hosting bill.

10-100x

more bandwidth consumed by AI crawlers compared to traditional search bots

Server Performance Impact

WordPress generates pages dynamically using PHP and MySQL queries. Every crawler request triggers this full processing pipeline unless you have aggressive caching in place. When an AI crawler makes thousands of requests per minute, you are effectively running a denial-of-service attack against your own server.

The symptoms look like:

  • Slow page load times for real human visitors
  • Database connection errors and timeouts
  • PHP worker exhaustion on shared hosting
  • Increased CPU usage leading to hosting provider throttling
  • Cache invalidation as crawlers hit rarely-visited pages

Research shows that 74% of shoppers abandon a site if it takes more than 3 seconds to load. If AI crawler load is causing your WooCommerce store to slow down, you are losing real revenue to bots.

Server Log Bloat

Every request generates a log entry. At 7.9 million requests per month, that is a massive amount of log data consuming disk space and making legitimate log analysis nearly impossible. Finding a real security incident in a sea of AI crawler requests is like finding a needle in a very large, very noisy haystack.

Solution 1: robots.txt - The Polite Request

The simplest and most widely-recommended approach to blocking AI crawlers is adding directives to your robots.txt file. This is a text file in your WordPress root directory that tells crawlers what they can and cannot access.

The Complete robots.txt Configuration

Here is a comprehensive robots.txt configuration that blocks all major AI crawlers while allowing traditional search engines:

# Allow traditional search engines
User-agent: Googlebot
Allow: /

User-agent: Bingbot
Allow: /

# Block AI training crawlers
User-agent: GPTBot
Disallow: /

User-agent: ChatGPT-User
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: anthropic-ai
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: meta-externalagent
Disallow: /

User-agent: FacebookBot
Disallow: /

User-agent: Bytespider
Disallow: /

User-agent: Amazonbot
Disallow: /

User-agent: Applebot-Extended
Disallow: /

User-agent: PerplexityBot
Disallow: /

User-agent: cohere-ai
Disallow: /

User-agent: Diffbot
Disallow: /

User-agent: Omgilibot
Disallow: /

User-agent: YouBot
Disallow: /

# Standard WordPress directives
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php

Sitemap: https://yoursite.com/sitemap_index.xml

WordPress-Specific Implementation

WordPress generates a virtual robots.txt file by default. If you have a physical robots.txt file in your web root, WordPress uses that instead. To use the physical file approach:

  1. Create a file named robots.txt in your WordPress root directory (same level as wp-config.php)
  2. Add the directives above
  3. Verify by visiting yoursite.com/robots.txt in your browser

Alternatively, you can filter WordPress’s virtual robots.txt using the robots_txt filter:

add_filter( 'robots_txt', function( $output, $public ) {
    $ai_bots = array(
        'GPTBot',
        'ChatGPT-User',
        'Google-Extended',
        'ClaudeBot',
        'anthropic-ai',
        'CCBot',
        'meta-externalagent',
        'FacebookBot',
        'Bytespider',
        'Amazonbot',
        'Applebot-Extended',
        'PerplexityBot',
        'cohere-ai',
    );

    foreach ( $ai_bots as $bot ) {
        $output .= "\nUser-agent: {$bot}\nDisallow: /\n";
    }

    return $output;
}, 10, 2 );

This approach is cleaner because it works with WordPress’s built-in robots.txt handling and is compatible with SEO plugins that also modify robots.txt.

The Critical Limitation

robots.txt is a voluntary standard. It is an honor system - crawlers can choose to ignore it entirely. While major companies like OpenAI and Google claim to respect robots.txt, there is no technical enforcement mechanism. Think of it as putting up a “No Trespassing” sign: it establishes your intent, but it does not build a fence.

robots.txt is the “please don’t” of web security. It tells well-behaved bots to stay away, but does nothing to stop the ones that ignore the rules. For real protection, you need server-level blocking.

Solution 2: The Block AI Crawlers WordPress Plugin

For WordPress users who want a one-click solution, the Block AI Crawlers plugin (available on wordpress.org) offers a straightforward approach.

What the Plugin Does

  • Automatically updates your robots.txt with AI crawler blocking directives
  • Adds <meta> tags to your site’s HTML telling AI bots not to index or use content for training
  • Maintains an updated list of known AI crawler user agents
  • Requires no configuration beyond activation

Installation and Setup

  1. Go to Plugins → Add New in your WordPress admin
  2. Search for “Block AI Crawlers”
  3. Install and activate the plugin
  4. Verify by visiting yoursite.com/robots.txt - you should see the new directives

What the Plugin Adds to Your HTML

Beyond robots.txt, the plugin injects meta tags like:

<meta name="robots" content="noai, noimageai">

These meta directives tell AI systems not to use your content or images for training purposes. Like robots.txt, these are voluntary signals - but they strengthen your legal position if an AI company uses your content after you have explicitly opted out.

Limitations

  • The plugin only works at the WordPress level - it cannot block bots that do not reach WordPress (e.g., bots blocked at the server or CDN level)
  • It relies on bots identifying themselves honestly via user agent strings
  • It does not actively block requests - it only signals intent
  • Cannot prevent AI companies from using cached versions of your content from before the plugin was installed
✓ Best For: Non-technical WordPress users who want basic AI crawler blocking without editing configuration files. Layer this with server-level blocking for comprehensive protection.

Solution 3: Server-Level Blocking - The Real Fence

If robots.txt is a polite request and plugins are a WordPress-level solution, server-level blocking is the actual enforcement mechanism. By blocking AI crawlers at the web server level, you prevent them from even reaching WordPress, saving server resources and bandwidth.

Apache (.htaccess) Configuration

If your WordPress site runs on Apache (the most common configuration for shared hosting), add these rules to your .htaccess file in the WordPress root directory. Place them before the WordPress rewrite rules:

# Block AI Crawlers - Add before WordPress rules
<IfModule mod_rewrite.c>
RewriteEngine On

# OpenAI
RewriteCond %{HTTP_USER_AGENT} GPTBot [NC,OR]
RewriteCond %{HTTP_USER_AGENT} ChatGPT-User [NC,OR]

# Meta
RewriteCond %{HTTP_USER_AGENT} meta-externalagent [NC,OR]
RewriteCond %{HTTP_USER_AGENT} FacebookBot [NC,OR]

# Google AI Training
RewriteCond %{HTTP_USER_AGENT} Google-Extended [NC,OR]

# Anthropic
RewriteCond %{HTTP_USER_AGENT} ClaudeBot [NC,OR]
RewriteCond %{HTTP_USER_AGENT} anthropic-ai [NC,OR]

# Common Crawl
RewriteCond %{HTTP_USER_AGENT} CCBot [NC,OR]

# ByteDance
RewriteCond %{HTTP_USER_AGENT} Bytespider [NC,OR]

# Amazon
RewriteCond %{HTTP_USER_AGENT} Amazonbot [NC,OR]

# Apple AI
RewriteCond %{HTTP_USER_AGENT} Applebot-Extended [NC,OR]

# Perplexity
RewriteCond %{HTTP_USER_AGENT} PerplexityBot [NC,OR]

# Cohere
RewriteCond %{HTTP_USER_AGENT} cohere-ai [NC,OR]

# Diffbot
RewriteCond %{HTTP_USER_AGENT} Diffbot [NC,OR]

# You.com
RewriteCond %{HTTP_USER_AGENT} YouBot [NC]

# Return 403 Forbidden
RewriteRule ^ - [F,L]
</IfModule>
⚠️ Warning: Always back up your .htaccess file before editing. A syntax error in .htaccess can take your entire site offline. Test by visiting your site immediately after saving changes.

Nginx Configuration

If your WordPress site runs on Nginx (common with VPS hosting, Cloudways, and performance-optimized setups), add these rules to your server block:

# Block AI Crawlers
map $http_user_agent $is_ai_bot {
    default 0;
    ~*GPTBot 1;
    ~*ChatGPT-User 1;
    ~*meta-externalagent 1;
    ~*FacebookBot 1;
    ~*Google-Extended 1;
    ~*ClaudeBot 1;
    ~*anthropic-ai 1;
    ~*CCBot 1;
    ~*Bytespider 1;
    ~*Amazonbot 1;
    ~*Applebot-Extended 1;
    ~*PerplexityBot 1;
    ~*cohere-ai 1;
    ~*Diffbot 1;
    ~*YouBot 1;
}

server {
    # ... your existing server configuration ...

    if ($is_ai_bot) {
        return 403;
    }
}

The Nginx map directive is more efficient than multiple if statements because it evaluates once and stores the result. This is the recommended approach for Nginx-based WordPress setups.

Why Server-Level Blocking Is Superior

Server-level blocking has several advantages over robots.txt and plugin-based approaches:

  • Enforcement, not suggestion: The request is rejected before it reaches WordPress, PHP, or MySQL
  • Resource savings: No PHP execution, no database queries, no WordPress bootstrapping for blocked requests
  • Bandwidth savings: A 403 response is a few bytes, versus the kilobytes or megabytes of a full page render
  • Speed: Server-level rules evaluate in microseconds, before any application code runs

The limitation, of course, is that it relies on user agent identification. Bots that disguise their user agent as a regular browser will bypass these rules. This is where CDN-level protection becomes valuable.

Solution 4: Cloudflare AI Crawl Control - The CDN Shield

Cloudflare, which sits in front of millions of WordPress sites, has introduced dedicated AI crawler management features. This is arguably the most effective single solution for most WordPress sites.

How to Enable Cloudflare’s AI Bot Protection

  1. Log into your Cloudflare dashboard
  2. Navigate to Security → Bots
  3. Look for the AI Crawlers or AI Scrapers section
  4. Toggle the blocking controls for AI crawlers

Cloudflare’s approach is more sophisticated than user agent blocking because it uses behavioral analysis, IP reputation, and machine learning to identify AI crawlers - even ones that try to disguise their identity.

Advantages of CDN-Level Blocking

  • Blocks before traffic reaches your server: Cloudflare’s edge network stops AI crawlers at the CDN level, meaning zero bandwidth and zero server resources consumed
  • Behavioral detection: Can identify crawlers that disguise their user agent based on request patterns
  • Automatic updates: Cloudflare maintains and updates their bot detection, so new AI crawlers are caught without you needing to update configurations
  • Global edge network: Blocking happens at the edge location closest to the bot, minimizing even the bandwidth used for rejection
  • Bot analytics: Cloudflare provides dashboards showing which bots are hitting your site and how much traffic they generate

Custom Cloudflare WAF Rules

For more granular control, you can create custom WAF (Web Application Firewall) rules in Cloudflare:

  1. Go to Security → WAF → Custom Rules
  2. Create a new rule with conditions matching AI crawler user agents
  3. Set the action to Block or Challenge (Challenge presents a CAPTCHA)

Example expression for Cloudflare WAF:

(http.user_agent contains "GPTBot") or
(http.user_agent contains "ChatGPT-User") or
(http.user_agent contains "ClaudeBot") or
(http.user_agent contains "meta-externalagent") or
(http.user_agent contains "Google-Extended") or
(http.user_agent contains "CCBot") or
(http.user_agent contains "Bytespider") or
(http.user_agent contains "Amazonbot") or
(http.user_agent contains "Applebot-Extended") or
(http.user_agent contains "PerplexityBot") or
(http.user_agent contains "anthropic-ai") or
(http.user_agent contains "cohere-ai")
✓ Recommended: Cloudflare’s free plan includes basic bot management. The Pro plan ($20/month) includes more advanced bot analytics. For most WordPress sites, the free plan plus custom WAF rules provides excellent AI crawler protection.

Solution 5: llms.txt - The Nuanced Approach

While the solutions above are about blocking AI crawlers entirely, llms.txt represents a more nuanced approach: telling AI systems how to use your content rather than simply blocking them.

What Is llms.txt?

llms.txt is a proposed standard (gaining traction in developer communities, particularly on Dev.to) that functions as a companion to robots.txt. While robots.txt tells crawlers where they can and cannot go, llms.txt provides context about your site that helps AI systems use your content appropriately.

Think of it this way:

  • robots.txt = “You can/cannot access these pages”
  • llms.txt = “Here is what my site is about and how you should represent it”

Setting Up llms.txt on WordPress

Create a file named llms.txt in your WordPress root directory (same level as wp-config.php):

# Your Site Name

> A brief, accurate description of your site and its expertise area.

## Primary Content

- [WordPress Tutorials](/category/tutorials): Step-by-step WordPress development guides
- [Plugin Reviews](/category/reviews): In-depth WordPress plugin reviews and comparisons
- [Performance](/category/performance): WordPress optimization and speed guides

## Authoritative Resources

- [Documentation](/docs): Technical reference documentation
- [Case Studies](/case-studies): Real-world implementation examples

## About

- [About Us](/about): Our team, credentials, and expertise
- [Editorial Policy](/editorial-policy): How we research and verify our content

## Usage Guidelines

- Content may be cited with attribution
- Please link to original source when referencing
- Do not reproduce full articles
- Data and statistics should include publication date context

Dynamic llms.txt Generation with WordPress

For sites with frequently changing content, you can generate llms.txt dynamically. Here is a mu-plugin approach:

<?php
/**
 * Dynamic llms.txt generator for WordPress.
 * Place in wp-content/mu-plugins/llms-txt.php
 */
add_action( 'init', function() {
    add_rewrite_rule( '^llms\.txt$', 'index.php?llms_txt=1', 'top' );
});

add_filter( 'query_vars', function( $vars ) {
    $vars[] = 'llms_txt';
    return $vars;
});

add_action( 'template_redirect', function() {
    if ( ! get_query_var( 'llms_txt' ) ) {
        return;
    }

    header( 'Content-Type: text/plain; charset=utf-8' );
    header( 'X-Robots-Tag: noindex' );

    $site_name = get_bloginfo( 'name' );
    $site_desc = get_bloginfo( 'description' );

    echo "# {$site_name}\n\n";
    echo "> {$site_desc}\n\n";

    // List categories with descriptions.
    $categories = get_categories( array(
        'hide_empty' => true,
        'orderby'    => 'count',
        'order'      => 'DESC',
        'number'     => 20,
    ) );

    if ( ! empty( $categories ) ) {
        echo "## Content Areas\n\n";
        foreach ( $categories as $category ) {
            $url  = get_category_link( $category->term_id );
            $path = wp_parse_url( $url, PHP_URL_PATH );
            $desc = $category->description ?: "Posts about {$category->name}";
            echo "- [{$category->name}]({$path}): {$desc}\n";
        }
        echo "\n";
    }

    // List top pages.
    $pages = get_pages( array(
        'sort_column' => 'menu_order',
        'number'      => 10,
    ) );

    if ( ! empty( $pages ) ) {
        echo "## Key Pages\n\n";
        foreach ( $pages as $page ) {
            $path = wp_parse_url( get_permalink( $page ), PHP_URL_PATH );
            echo "- [{$page->post_title}]({$path})\n";
        }
        echo "\n";
    }

    // List recent popular posts.
    $recent = get_posts( array(
        'numberposts' => 10,
        'orderby'     => 'comment_count',
        'order'       => 'DESC',
    ) );

    if ( ! empty( $recent ) ) {
        echo "## Popular Content\n\n";
        foreach ( $recent as $post ) {
            $path = wp_parse_url( get_permalink( $post ), PHP_URL_PATH );
            echo "- [{$post->post_title}]({$path})\n";
        }
    }

    exit;
});

After adding this mu-plugin, flush your rewrite rules by visiting Settings → Permalinks and clicking “Save Changes.” Then verify by visiting yoursite.com/llms.txt.

The Strategic Question: Block Everything or Be Selective?

This is where the AI crawler conversation gets genuinely complex. There is no one-size-fits-all answer, and the right strategy depends on your site’s goals, content type, and business model.

The Case for Blocking All AI Crawlers

  • Bandwidth savings - the most immediate and measurable benefit
  • Content protection - your content is your intellectual property; AI companies profiting from it without compensation is a legitimate concern
  • Server performance - fewer bot requests means more resources for real visitors
  • Legal positioning - explicitly opting out strengthens your position in potential future copyright or data use disputes
  • Privacy - if your site handles user-generated content, AI crawlers could be ingesting your users’ data

The Case for Selective Blocking

Here is the uncomfortable trade-off: if you block ALL AI crawlers, you may be hurting your visibility in the rapidly-growing AI search channel.

  • Blocking Google-Extended prevents Google from using your content for Gemini training, but it does NOT affect AI Overviews (which use regular Googlebot data). However, future changes to Google’s AI features might be affected.
  • Blocking GPTBot prevents OpenAI from training on your content, but blocking ChatGPT-User prevents your content from being cited when ChatGPT users ask questions. These are different decisions with different trade-offs.
  • Blocking PerplexityBot removes you from Perplexity’s AI search results entirely - and Perplexity is a growing traffic source for content sites.

The Dilemma

Block AI crawlers to save bandwidth, or allow them for AI search visibility?

A Recommended Tiered Strategy

Based on current data and trends, here is a tiered approach that balances protection with discoverability:

Tier 1 - Always Block (aggressive crawlers with no search benefit):

  • meta-externalagent - extremely aggressive, no search benefit
  • Bytespider - aggressive crawling, limited benefit for Western markets
  • CCBot - used by many AI companies for training, no direct search benefit
  • cohere-ai - enterprise AI training, no consumer search benefit

Tier 2 - Block for Training, Allow for Search:

  • GPTBot - block (prevents training), but keep ChatGPT-User - allow (enables live citations)
  • Google-Extended - block (prevents Gemini training, does not affect regular search)
  • Applebot-Extended - block (prevents Apple AI training)

Tier 3 - Consider Allowing (if AI search visibility matters to you):

  • ChatGPT-User - allowing this lets your content be cited in ChatGPT conversations
  • PerplexityBot - allowing this keeps you in Perplexity search results
  • ClaudeBot - Anthropic’s crawler; consider based on your traffic from Claude

Monitoring: How to Detect AI Crawler Activity

Before you can block AI crawlers effectively, you need to know which ones are hitting your site and how aggressively. Here are methods for WordPress-specific monitoring.

Server Access Logs

The most reliable method is analyzing your server access logs. Here is a command you can run via SSH:

# Count requests by AI bot user agents in the last 30 days
grep -iE "(GPTBot|ChatGPT-User|ClaudeBot|meta-externalagent|Google-Extended|CCBot|Bytespider|Amazonbot|PerplexityBot)" \
    /var/log/apache2/access.log | \
    awk '{print $NF}' | \
    sort | uniq -c | sort -rn

For Nginx:

grep -iE "(GPTBot|ChatGPT-User|ClaudeBot|meta-externalagent|Google-Extended|CCBot|Bytespider)" \
    /var/log/nginx/access.log | \
    awk -F'"' '{print $6}' | \
    sort | uniq -c | sort -rn

WordPress Plugin for Bot Monitoring

If you do not have SSH access, you can use a simple mu-plugin to log AI crawler activity:

<?php
/**
 * AI Crawler Logger - Logs AI bot visits to a custom database table.
 * Place in wp-content/mu-plugins/ai-crawler-logger.php
 */
add_action( 'init', function() {
    $user_agent = isset( $_SERVER['HTTP_USER_AGENT'] ) ? $_SERVER['HTTP_USER_AGENT'] : '';

    $ai_bots = array(
        'GPTBot',
        'ChatGPT-User',
        'ClaudeBot',
        'anthropic-ai',
        'meta-externalagent',
        'FacebookBot',
        'Google-Extended',
        'CCBot',
        'Bytespider',
        'Amazonbot',
        'PerplexityBot',
        'Applebot-Extended',
        'cohere-ai',
    );

    $detected_bot = '';
    foreach ( $ai_bots as $bot ) {
        if ( stripos( $user_agent, $bot ) !== false ) {
            $detected_bot = $bot;
            break;
        }
    }

    if ( ! $detected_bot ) {
        return;
    }

    global $wpdb;
    $table = $wpdb->prefix . 'ai_crawler_log';

    // Create table if needed.
    if ( $wpdb->get_var( "SHOW TABLES LIKE '{$table}'" ) !== $table ) {
        $wpdb->query(
            "CREATE TABLE {$table} (
                id BIGINT UNSIGNED AUTO_INCREMENT PRIMARY KEY,
                bot_name VARCHAR(100) NOT NULL,
                request_uri TEXT,
                ip_address VARCHAR(45),
                logged_at DATETIME DEFAULT CURRENT_TIMESTAMP,
                INDEX idx_bot (bot_name),
                INDEX idx_date (logged_at)
            ) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4"
        );
    }

    $wpdb->insert(
        $table,
        array(
            'bot_name'    => $detected_bot,
            'request_uri' => isset( $_SERVER['REQUEST_URI'] ) ? $_SERVER['REQUEST_URI'] : '',
            'ip_address'  => isset( $_SERVER['REMOTE_ADDR'] ) ? $_SERVER['REMOTE_ADDR'] : '',
        ),
        array( '%s', '%s', '%s' )
    );
}, 0 );

Then query the data with:

-- Total requests per bot in the last 30 days
SELECT bot_name, COUNT(*) as requests
FROM wp_ai_crawler_log
WHERE logged_at > DATE_SUB(NOW(), INTERVAL 30 DAY)
GROUP BY bot_name
ORDER BY requests DESC;

-- Requests per day for trending
SELECT DATE(logged_at) as day, bot_name, COUNT(*) as requests
FROM wp_ai_crawler_log
WHERE logged_at > DATE_SUB(NOW(), INTERVAL 7 DAY)
GROUP BY day, bot_name
ORDER BY day DESC, requests DESC;

Cloudflare Bot Analytics

If you use Cloudflare, the Security → Bots dashboard provides comprehensive bot traffic analytics without any code or plugins needed. It shows:

  • Total bot requests over time
  • Bot types (verified, unverified, automated)
  • Top bot user agents
  • Bot traffic by country
  • Requests blocked vs. allowed

The Bandwidth vs. Discoverability Trade-Off: A Decision Framework

To help you make the right decision for your specific WordPress site, consider these factors:

Block Aggressively If:

  • You are on shared hosting with bandwidth limits
  • Your site is experiencing performance issues
  • Your content is behind a paywall or membership gate
  • Your site contains proprietary research, original photography, or other high-value IP
  • You run a WooCommerce store where site speed directly impacts revenue
  • Your hosting costs are metered by bandwidth

Be Selective If:

  • You are a content publisher who benefits from broad visibility
  • You are on robust hosting that can handle additional crawler traffic
  • AI search traffic is a meaningful and growing referral source
  • Your business model depends on being discovered (consulting, SaaS, etc.)
  • You are actively implementing an AI search optimization strategy

Allow Most AI Crawlers If:

  • You actively want your content to appear in AI responses
  • Your hosting can handle the load without performance impact
  • Brand visibility in AI platforms is a strategic priority
  • You produce content specifically designed for AI citation (comparisons, guides, reference material)

Combining Strategies: The Defense-in-Depth Approach

The most effective approach combines multiple solutions in layers. Here is the recommended stack, from outermost to innermost:

Layer 1: CDN/WAF (Cloudflare or similar)

First line of defense. Blocks known AI crawlers at the edge before traffic reaches your server. Handles behavioral detection for disguised bots.

Layer 2: Server Configuration (.htaccess/Nginx)

Second line of defense. Catches anything that gets past the CDN layer. Returns 403 before WordPress loads.

Layer 3: WordPress Plugin or mu-plugin

Application-level blocking and meta tag injection. Adds noai/noimageai meta directives for any bot that makes it through the upper layers.

Layer 4: robots.txt

The honor system layer. Tells well-behaved bots your preferences. Creates a documented record of your opt-out intent.

Layer 5: llms.txt

The nuanced layer. For bots you choose to allow, provides context about how your content should be used and attributed.

# Visualization of the defense-in-depth stack:

Request from AI Bot
    |
    v
+---------------------+
|  CDN/WAF (Cloudflare)| <- Blocks ~95% of AI crawlers
|  Behavioral + UA     |
+----------+----------+
           | (if not blocked)
           v
+---------------------+
|  Server (.htaccess)  | <- Blocks by User Agent
|  or Nginx config     |
+----------+----------+
           | (if not blocked)
           v
+---------------------+
|  WordPress Plugin    | <- Meta tags + robots.txt
|  or mu-plugin        |
+----------+----------+
           |
           v
+---------------------+
|  robots.txt          | <- Honor system
|  + llms.txt          | <- Usage guidance
+---------------------+

Legal and Ethical Considerations

The legal landscape around AI crawling is evolving rapidly, and WordPress site owners should be aware of their rights and the current state of the law.

Copyright and Fair Use

Several high-profile lawsuits are currently working through the courts regarding whether AI training on copyrighted web content constitutes fair use. The outcomes of these cases will significantly impact the legal framework around AI crawling. In the meantime:

  • Explicit opt-out strengthens your position. Having robots.txt directives, meta tags, and llms.txt that explicitly deny AI training use creates a documented record of your intent.
  • Terms of service matter. If your site’s ToS prohibits automated scraping for AI training, this provides additional legal grounds.
  • DMCA and takedown requests. If you find your content reproduced by AI systems after explicitly opting out, you may have grounds for DMCA takedown requests.

GDPR and Privacy

If your WordPress site has European visitors or handles EU personal data, AI crawling raises GDPR concerns:

  • AI crawlers scraping user-generated content (comments, forum posts, reviews) may be collecting personal data without consent
  • GDPR gives individuals the right to object to processing of their data for AI training purposes
  • Site owners may have a data controller obligation to prevent unauthorized data collection by AI crawlers

The Compensation Question

There is growing momentum behind the idea that content creators should be compensated when their work is used to train AI models. Several publishing organizations and content creator groups are advocating for licensing frameworks. WordPress site owners should follow these developments, as they could create new revenue opportunities or obligations.

The question is not whether AI companies should pay for content - it is when and how the compensation frameworks will be established. Protect your content now so you are in the strongest possible position when those frameworks arrive.

A Practical Decision Matrix for WordPress Sites

To summarize the entire article into actionable decisions, use this matrix based on your site type:

Personal Blog / Portfolio

  • Recommendation: Block aggressively
  • Implementation: robots.txt + Block AI Crawlers plugin
  • Rationale: Minimal benefit from AI search visibility; protect your original content

Business / Agency Website

  • Recommendation: Selective blocking (Tier 1 + Tier 2 from our strategy above)
  • Implementation: Cloudflare + .htaccess + llms.txt
  • Rationale: Block training crawlers, allow search-oriented crawlers for discoverability

Content Publisher / Media Site

  • Recommendation: Selective blocking with monitoring
  • Implementation: Cloudflare + selective robots.txt + llms.txt + monitoring mu-plugin
  • Rationale: Balance content protection with AI search traffic potential

WooCommerce / E-Commerce

  • Recommendation: Block aggressively
  • Implementation: Full stack (Cloudflare + server + plugin + robots.txt)
  • Rationale: Site performance directly impacts revenue; product data is competitive intelligence

Membership / Gated Content

  • Recommendation: Block everything
  • Implementation: Full stack + ensure login walls prevent crawling gated content
  • Rationale: Gated content leaking into AI responses undermines your business model

Implementation Checklist

Here is a step-by-step checklist you can follow today to implement AI crawler protection on your WordPress site:

Step 1: Assess Current Impact (15 minutes)

  • ☐ Check server logs or Cloudflare analytics for AI crawler activity
  • ☐ Note which bots are hitting your site and how frequently
  • ☐ Calculate approximate bandwidth consumed by AI crawlers

Step 2: Decide Your Strategy (10 minutes)

  • ☐ Review the decision matrix above and determine your site type
  • ☐ Decide which crawlers to block and which to allow
  • ☐ Document your decision for future reference

Step 3: Implement Blocking (30 minutes)

  • ☐ Update robots.txt with appropriate directives
  • ☐ Install and activate the Block AI Crawlers plugin (optional)
  • ☐ Add server-level rules (.htaccess or Nginx)
  • ☐ Configure Cloudflare AI bot protection (if applicable)

Step 4: Set Up Monitoring (20 minutes)

  • ☐ Install the AI crawler logger mu-plugin or configure log analysis
  • ☐ Set up a monthly review of AI crawler activity
  • ☐ Monitor hosting costs and performance metrics for changes

Step 5: Create llms.txt (15 minutes)

  • ☐ Draft your llms.txt content
  • ☐ Place in WordPress root or implement via mu-plugin
  • ☐ Verify it is accessible at yoursite.com/llms.txt

Step 6: Verify Everything (10 minutes)

  • ☐ Test your site loads normally after changes
  • ☐ Verify robots.txt is correct at yoursite.com/robots.txt
  • ☐ Confirm server rules are not blocking legitimate search engines
  • ☐ Check Google Search Console for any new crawl errors

Conclusion: Take Control of Your WordPress Site’s AI Footprint

AI crawlers are here, they are aggressive, and they are not going away. The Reddit story of 7.9 million Meta crawler requests is not an outlier - it is a preview of the default state for unprotected websites. The question is not whether AI bots are crawling your WordPress site. They almost certainly are. The question is whether you are going to do anything about it.

The good news is that effective protection is straightforward to implement. A combination of robots.txt directives, server-level rules, and CDN protection can eliminate virtually all AI crawler impact on your server. The harder question - whether to block selectively or comprehensively - requires understanding your site’s specific goals and making a strategic decision about the trade-off between content protection and AI search visibility.

Whatever you decide, do not ignore the problem. Every day that AI crawlers hit your unprotected WordPress site is a day they are consuming your bandwidth, slowing your server, and training on your content without permission. The tools to stop them are free, the implementation takes less than an hour, and the benefits are immediate and measurable.

Under 1 Hour

Total time to implement comprehensive AI crawler protection on WordPress


Further Reading:

Related Reading

Varun Dubey

Written by

Varun Dubey

Varun Dubey runs Wbcom Designs, the WordPress studio he founded in India in 2009. He has spent sixteen years building on WordPress and BuddyPress, shipping client work and products such as Reign, BuddyX, Jetonomy and MediaVerse, and has been putting Claude and OpenAI workflows into production since 2023. He writes up what the studio learns along the way.

More about Varun

No comments yet