The AI wave has hit WordPress hard. Every hosting provider, plugin developer, and agency is pitching AI features - and developers using AI tools are finding they have more to manage, not less. But most of those features depend on third-party APIs - OpenAI, Anthropic, Google. You send your users’ data to a cloud server, pay per token, and accept the privacy trade-offs that come with it. What if you didn’t have to? What if you could run an AI model on your own server, right alongside WordPress, without a single API call leaving your infrastructure?
This is no longer a pipe dream. Tools like Ollama and LM Studio have made it genuinely practical to run large language models on commodity hardware. And with WordPress’s REST API, connecting those local models to your site is more straightforward than most developers expect. This guide covers exactly how to do it - the hardware reality, the software setup, the REST API integration, and where self-hosted AI actually makes sense for WordPress sites.
Why Self-Hosted AI for WordPress?
Before we dig into the how, it’s worth being honest about the why. Self-hosted AI is not the right call for every WordPress site. The setup has real overhead, the hardware costs are significant, and the model quality still lags behind GPT-4o or Claude 3.5 for most tasks. So when does it make sense?
- Data privacy requirements - Healthcare, legal, financial, or government sites where data cannot leave your infrastructure. HIPAA compliance, GDPR strict interpretations, or client contracts that prohibit third-party data sharing.
- High-volume, cost-sensitive applications - If you’re running thousands of AI completions per day, API costs add up fast. Self-hosted models have a fixed infrastructure cost with zero per-token fees.
- Offline or air-gapped environments - Intranet WordPress installs, internal knowledge bases, or sites that operate in restricted network environments.
- Experimentation and development - Building and testing AI features without racking up API costs during development cycles.
- Custom fine-tuned models - Running models you’ve fine-tuned on your own data, which you wouldn’t want to upload to a cloud API anyway.
For a standard blog or small business site, a cloud API with generous free tiers (like Gemini Flash or GPT-4o Mini) is probably the better path. Self-hosted AI rewards sites with volume, compliance needs, or specific control requirements.
The Hardware Reality
Let’s talk about what you actually need before we go any further. Running LLMs locally requires serious compute resources, and the requirements scale with model size.
RAM Requirements by Model Size
Model Size | Minimum RAM | Recommended RAM | Example Models |
|---|---|---|---|
3B parameters | 4 GB | 8 GB | Llama 3.2 3B, Phi-3 Mini |
7B parameters | 8 GB | 16 GB | Mistral 7B, Llama 3.1 8B |
13B parameters | 16 GB | 32 GB | Llama 2 13B, CodeLlama 13B |
34B parameters | 32 GB | 64 GB | CodeLlama 34B, Yi 34B |
70B parameters | 64 GB | 128 GB | Llama 3.1 70B, Mixtral 8x7B |
These are RAM figures for CPU inference. GPU inference is significantly faster and uses VRAM instead. A consumer GPU with 8GB VRAM (like an RTX 3070) can run 7B models comfortably. For server deployments, NVIDIA A10 or A100 cards are common choices.
CPU vs GPU Inference
CPU inference is slow but accessible. A 7B model on a modern 8-core server CPU might generate 10-20 tokens per second - workable for batch operations or async tasks, painful for real-time chat. GPU inference on the same model typically delivers 50-100+ tokens per second.
For a WordPress plugin that generates post summaries in the background (async), CPU inference is fine. For a live chatbot that users interact with in real time, you need GPU acceleration or a faster small model.
The sweet spot for most WordPress self-hosted AI setups is a 7B parameter model on a server with 16GB RAM and a mid-range GPU. It delivers solid performance at a hardware cost that’s recoverable within months if you’re replacing paid API usage.
Ollama: The Easiest Path to Local LLMs
Ollama is the tool that changed the game for local AI. It wraps model downloads, quantization, and inference into a single CLI tool with a clean REST API. If you’ve run Docker, Ollama will feel familiar.
Installing Ollama on a Linux Server
Installation on Ubuntu or Debian is a single command. Ollama installs as a systemd service and starts automatically. The install script handles CUDA detection for GPU support.
Pulling and Running a Model
Once installed, pulling a model is as simple as pulling a Docker image. Ollama downloads the quantized model weights and stores them in ~/.ollama/models.
Ollama’s REST API
By default, Ollama listens on localhost:11434. It exposes a clean REST API with endpoints that feel familiar if you’ve used the OpenAI API. The /api/generate endpoint handles text completion, and /api/chat handles conversation-style interactions with message history.
The response comes back as a JSON object with the generated text in the response field. Ollama also supports streaming responses, which is useful for chat interfaces where you want to show text as it generates.
Exposing Ollama to Your WordPress Server
If Ollama and WordPress run on the same server, no network configuration is needed - WordPress PHP code calls http://localhost:11434 directly. If they run on separate machines on the same private network, you need to bind Ollama to its network interface.
Never expose Ollama directly to the public internet without authentication. The API has no built-in auth. Use a reverse proxy (nginx) with basic auth or an API key header if you need external access.
LM Studio: Local AI with a GUI
LM Studio is the desktop application counterpart to Ollama. It’s primarily aimed at Mac and Windows users who want a graphical interface for downloading and running models. It also exposes an OpenAI-compatible local server, which makes it dead simple to integrate with WordPress code written for the OpenAI API.
LM Studio’s local server defaults to port 1234 and accepts requests in exactly the OpenAI format - the same /v1/chat/completions endpoint, the same request body structure. This means any WordPress plugin or code that already works with the OpenAI API can point to LM Studio’s local server with a URL change and zero code modifications.
LM Studio vs Ollama - When to Use Which
Factor | Ollama | LM Studio |
|---|---|---|
Best for | Headless Linux servers, production | Mac/Windows dev machines |
API style | Custom Ollama API + OpenAI-compatible mode | OpenAI-compatible only |
GPU support | CUDA, ROCm, Metal | Metal (Mac), CUDA (Windows) |
Model library | Ollama model registry | HuggingFace direct |
CLI tooling | Full CLI interface | GUI only |
Production use | Yes (systemd service) | No (desktop app) |
For production WordPress deployments, Ollama on a Linux server is the right choice. LM Studio is excellent for development and testing on your local machine before deploying to a server.
Connecting Ollama to WordPress via REST API
Now the practical part: actually wiring Ollama into WordPress. There are three main patterns for this integration, depending on what you’re building.
Pattern 1: Server-Side PHP Calls (Best for Background Processing)
The simplest integration - WordPress PHP code makes HTTP requests to the Ollama API directly, server-to-server. This works well for tasks like generating post summaries, creating excerpts, or processing content when a post is saved. The user doesn’t wait for the response because it happens asynchronously.
The class above handles the API call with WordPress’s wp_remote_post(), which respects WordPress’s HTTP timeout settings and works through any proxy configuration. The 120-second timeout is intentional - local model inference can be slow for longer prompts.
Pattern 2: REST API Proxy Endpoint (Best for Frontend Chat)
For real-time chat interfaces, you need a WordPress REST API endpoint that the frontend can call. The endpoint acts as a proxy - it receives the user’s message from the browser, forwards it to Ollama, and streams the response back. This keeps Ollama hidden from the public internet while still enabling real-time interaction.
The nonce verification in the REST endpoint is critical. Without it, anyone could hit your endpoint and generate inference load on your server. The nonce ties requests to logged-in WordPress sessions.
Pattern 3: WP-Cron Batch Processing
For bulk operations - like generating SEO descriptions for all existing posts, or creating summaries for a library of resources - WP-Cron batch processing is the way to go. Queue the posts that need processing, run a cron job that processes them in small batches, and avoid timing out on any single request.
Building a Self-Hosted WordPress Chatbot
A self-hosted chatbot is one of the most compelling use cases for local AI on WordPress. You can give site visitors a conversational interface that knows your content, answers questions, and never sends a word to a third-party server. Here’s how to build a minimal but functional version.
The System Prompt Strategy
The quality of a chatbot comes down to the system prompt. For a WordPress site chatbot, you want to inject relevant context about your site’s purpose, tone, and content into the system prompt. For smaller models (7B and below), keep the system prompt tight - under 500 tokens. Larger context windows degrade performance on small models.
Choosing the Right Model for a Chatbot
Not all models are equally suited for chatbot use on WordPress. You want a model that’s instruction-tuned (follows directions rather than just completing text), has a good conversation format, and runs fast enough for real-time interaction.
- Llama 3.2 3B (Instruct) - Fast, surprisingly capable, good for simple Q&A chatbots. Runs well on CPU-only servers.
- Mistral 7B Instruct - Better reasoning than 3B models. Requires 8GB+ RAM. Good balance of quality and speed.
- Llama 3.1 8B Instruct - Meta’s latest small model. Strong instruction following, good for multi-turn conversations.
- Phi-3 Mini (3.8B) - Microsoft’s model, surprisingly strong for its size. Good for factual Q&A tasks.
For a public-facing chatbot on a shared server, start with Llama 3.2 3B. It’s fast enough for real-time responses even without GPU acceleration and delivers acceptable quality for most customer-facing use cases.
WordPress Plugins That Support Local AI Endpoints
You don’t always have to write custom code. Several WordPress AI plugins support custom API endpoints, which means you can point them at your local Ollama instance instead of OpenAI.
AI Engine (by Jordy Meow)
AI Engine is one of the most flexible AI plugins for WordPress. It supports custom OpenAI-compatible endpoints, which means you can configure it to use your local Ollama server (in OpenAI-compatible mode) or LM Studio server instead of the real OpenAI API. Navigate to AI Engine settings, add a custom model pointing to http://localhost:11434/v1, and your existing AI Engine features work with your local model.
BerriAI / LiteLLM Integration
LiteLLM is a proxy layer that gives any Ollama model an OpenAI-compatible interface. If you’re already running LiteLLM for other purposes, you can add your Ollama models to its config and any plugin that supports OpenAI will automatically have access to your local models through LiteLLM’s unified endpoint.
Writing Your Own Plugin
For serious integrations, a lightweight custom plugin is the cleanest approach. You control the model parameters, the prompt templates, the caching layer, and the admin interface. The code patterns in the previous section give you a solid foundation. Add an options page with add_options_page() to configure the Ollama endpoint URL and model name, and you have a production-ready self-hosted AI plugin.
Performance Tuning for WordPress + Ollama
Out-of-the-box Ollama performance is fine for experimentation but leaves optimization on the table for production. These are the levers that matter most.
Response Caching
Many AI requests on a WordPress site are deterministic enough to cache. If 50 users ask your chatbot the same FAQ question, there’s no reason to run inference 50 times. Use WordPress transients or a Redis object cache to store responses keyed by the prompt hash. Cache expiry of 1-24 hours is appropriate for most FAQ and informational queries.
Model Quantization Choices
Ollama downloads quantized models by default. Quantization reduces model size and memory usage at a small cost to output quality. The Q4_K_M quantization level (4-bit with medium quality) is the standard sweet spot - it halves memory usage compared to full precision with barely noticeable quality degradation for text tasks. Q8_0 is higher quality but needs more RAM. Q2_K is the smallest but noticeably worse for complex tasks.
Concurrency and Queue Management
Ollama handles one request at a time by default (single-threaded inference). If multiple WordPress requests hit Ollama simultaneously, they queue up. For a lightly trafficked site this is fine. For heavier loads, either run multiple Ollama instances on different ports and round-robin between them, or use a proper async job queue (WP Background Processing library or custom WP-Cron implementation) to serialize AI requests rather than letting them pile up.
Security Considerations
Running AI locally introduces security considerations that don’t exist with cloud APIs. Address these before deploying to production.
- Network isolation - Ollama should only listen on localhost or your internal network. Never expose port 11434 to the public internet. Use a firewall rule to enforce this at the network level, not just Ollama’s bind address.
- Input sanitization - All user-provided text going into AI prompts must be sanitized. Prompt injection attacks are real - a malicious user can craft inputs that manipulate the model’s behavior. Strip HTML, limit input length, and validate that inputs match expected formats before including them in prompts.
- Rate limiting - Without rate limiting, a single user can hammer your endpoint and exhaust server resources. Implement rate limiting on your WordPress REST API endpoint using a transient-based counter or a plugin like WP Rate Limiter. Keeping your endpoint responses fast also matters - consider reviewing how third-party integrations affect page load time as a reference for similar API optimizations.
- Output sanitization - AI output is untrusted user content from a security perspective. Always run
wp_kses()orsanitize_text_field()on model output before displaying it to other users or storing it in the database. - Audit logging - Log all AI requests and responses to a database table for audit purposes. Include the user ID, timestamp, prompt hash, and response. This is especially important for compliance use cases.
Practical Use Cases Worth Building
Beyond the chatbot, here are the WordPress self-hosted AI use cases that deliver genuine value:
Automatic Post Summarization
Hook into save_post to automatically generate a 2-3 sentence summary when a post is saved. Store it as post meta. Display it in email newsletters, RSS feeds, or social share previews. Zero API cost, runs on your server, and your users get better content previews.
Content Moderation
For sites with user-generated content (BuddyPress, WooCommerce reviews, bbPress), run comments through a local moderation model before they go to the standard WordPress comment queue. A 7B model can classify content as spam, hate speech, or off-topic with reasonable accuracy - faster and cheaper than a cloud moderation API at volume.
Internal Search Enhancement
WordPress’s default search is keyword-based and often frustrating. A local embedding model (like nomic-embed-text, available via Ollama) can generate semantic embeddings for your posts and enable semantic search - finding relevant content even when the exact keywords don’t match. This is the “AI search” feature that enterprise products charge thousands for, running on your own server.
Translation Drafts
For multilingual sites using WPML or Polylang, a local LLM can generate draft translations that your human translators then review and polish. Not as accurate as DeepL, but free at scale, and the privacy argument is compelling for sensitive content.
When to Stick with Cloud APIs
Self-hosted AI is genuinely powerful, but it’s not always the right answer. Be honest about the trade-offs.
- Small sites with low volume - If you’re running fewer than a few hundred AI requests per day, cloud API costs are likely under $10/month. The infrastructure and maintenance overhead of self-hosted is not worth it.
- Tasks requiring frontier model quality - For complex reasoning, nuanced writing, or tasks where output quality critically matters, GPT-4o and Claude 3.5 Sonnet still outperform the best self-hosted 7B models significantly.
- No dedicated hardware - Shared hosting and basic VPS plans don’t have the RAM for local LLMs. You need a dedicated server or VPS with at least 8GB of free RAM after your normal stack is running.
- Rapid iteration on prompts - Cloud APIs let you switch models and experiment quickly. Self-hosted requires managing model downloads, storage, and version management.
Getting Started: A Minimal Setup
If you want to experiment with self-hosted AI for WordPress without committing to production infrastructure, here’s the minimal path to a working setup:
- Spin up a DigitalOcean or Hetzner VPS with 16GB RAM - This is your Ollama server. Hetzner’s CPX31 (4 vCPU, 8GB RAM) costs about 12 euros/month and handles 7B models reasonably well on CPU.
- Install Ollama and pull Llama 3.2 3B Instruct - Lightweight model, fast CPU inference, good for experimentation.
- Set up a test WordPress install on the same server or a separate VPS - Use wp-cli to set up WordPress quickly.
- Add the OllamaWP class from the code section above - Drop it in a simple custom plugin or your theme’s functions.php.
- Create a simple admin page that sends a test prompt and displays the response - Verify the integration works end to end before building anything more complex.
Start small, verify the integration, measure actual inference latency on your hardware, and then decide whether to invest in better hardware or more sophisticated features. The cost of the experiment is 2-3 weeks of a cheap VPS - easily worth it before committing to a production architecture.
The Bigger Picture
Self-hosted AI for WordPress is not a replacement for cloud APIs - it’s a complement to them. Cloud APIs give you access to frontier models with zero infrastructure overhead. Self-hosted gives you privacy, cost control, and the ability to run on your own infrastructure. The right architecture for most serious WordPress AI applications is probably a hybrid: cloud APIs for tasks where quality matters most, local models for high-volume background processing and compliance-sensitive operations.
The tools have matured enough in 2026 that self-hosted AI is no longer a research project. Ollama in particular has made the setup genuinely accessible for developers who aren’t AI specialists. If your WordPress site has a use case that fits - high volume, privacy requirements, or offline operation - the barrier to entry has never been lower.
Ready to Explore More WordPress Development?
Self-hosted AI is one piece of the modern WordPress stack. If you’re building custom plugins, integrating external APIs, or optimizing performance, AttowP covers the technical WordPress topics that actually matter for developers. Browse the WordPress development category for more in-depth technical guides, or check out our coverage of WordPress plugins for practical plugin development tutorials.




No comments yet