How to Set Up a Free, Uncensored AI Companion Locally (No GPU)
Written by Clement

You want a completely private AI companion that rolls with whatever you throw at it, free and unlimited? You're on the right guide. This installs a 100% local, uncensored chat model on your own hardware in about ten minutes, no GPU required.
Try it now!
This is exactly what you get with this guide. The server behind this demo is an old laptop with 8GB of RAM.
What you need
It runs on CPU alone. On a laptop with only 8GB of RAM I got about 10 tokens per second.
- OS
- Windows, Mac, or Linux
- GPU
- None needed. Runs on CPU only.
- RAM
- Minimum of 4GB of available RAM.
- Disk space
- Ollama itself plus about 3.3 GB for the model file
One-click scripts
Get Ollama and the abliterated model running. The script saves you time; if you'd rather not run one, use the "Do it by hand" card.
Script for Windows
Installs Ollama, pulls the model, and verifies it responds. Just double-click the downloaded file to run it, no right-click steps needed.
Script for Mac / Linux
Installs Ollama, pulls the model, and verifies it responds. After downloading, open a terminal in that folder and run: bash setup-local-uncensored-companion.sh
Stuck anywhere in this guide?
A companion needs a persona on top of the model.
You can test the model right now with ollama commands. But the terminal forgets every message, and what you installed is a blank model with no name and no personality.
So I built a small chat page for you. It runs in your browser, talks straight to your local Ollama, and keeps everything local. You set a name, an avatar, and a personality once.
- 2
Step 2 of 4
Download the companion chat page
A single HTML file.
No build step, no server, no account. Once it's saved you can open it with your internet off.
Download the companion chat page - 3
Step 3 of 4
Open it in your browser
Double-click the saved file.
It opens in a normal browser tab, straight from your disk, to a setup screen: name, avatar image URL, personality prompt.
- 4
Step 4 of 4
Start chatting
Hit Start chatting.
Your companion's config is saved in your browser for next time.
If something breaks
Chat page says "Ollama unreachable"
Make sure Ollama is running in the background, and that the server URL field matches where it listens (http://localhost:11434 by default)
Replies never finish, CPU pinned at 100%
Thinking mode is on. The chat page and both scripts already disable it; in a raw ollama run session, type /set nothink first
If you do have a GPU: there's a middle tier between this and a hosted API
Everything above assumes CPU only, which is deliberately the lowest common denominator: it runs on a decade-old laptop. If you have a GPU sitting idle, a bigger local model becomes realistic, and the appeal is the same as the CPU setup, nothing leaves your machine, plus noticeably better memory and coherence in a long roleplay session than a 4B model can manage.
The model worth pointing at here is Qwen 3.8 27B, one tier up from the abliterated 4B this guide installs by default. At Q4_K_M quantization (about 14GB) it fits a single 16GB-VRAM consumer GPU alongside 32GB of system RAM, running around 25 to 35 tokens per second. Push to Q3_K_M (about 12.5 to 13GB) on that same 16GB card and throughput rises to roughly 41 tokens per second, with a small quality cost. Running it on system RAM instead of a GPU (the BF16 full-precision weights, about 55GB, needing 64GB+ of RAM) drops throughput to around 4 tokens per second, a bandwidth bottleneck that makes it slower than the CPU-only 4B setup above rather than an upgrade from it, so this model is worth running only with a GPU to offload onto.
On cost, running Q3/Q4 on a single RTX 3090/4090-class card at roughly 250W draws about 0.37 US dollars per million output tokens at typical US residential electricity rates, against roughly 3.00 US dollars per million output tokens on a comparable FP8 cloud endpoint at lower throughput. We confirmed this ourselves in a roleplay session on the abliterated build, not just the coding/agentic benchmark it was first measured on: same speed per tier, same cost, nothing about running it as a companion instead of a coding assistant changes the numbers.
When a hosted API makes more sense than local
Local is the right default for privacy, and it is the only setup where you can keep chatting offline, on CPU or GPU. But it is not automatically the better experience. Even a GPU-accelerated local model tops out below what you can rent by the token, and in long conversations that gap shows up as thinner memory and flatter replies.
Pick local when privacy is the point, when you want zero running cost, or when you want it to work offline. Pick a hosted API when you want a larger model than any of your own hardware can hold.
Which hosted models allow uncensored roleplay is its own question, so we wrote up which hosted APIs actually allow uncensored roleplay to give some solutions.
Frequently asked questions
- Why run it locally?
- Nothing from your machine is transferred online, so you can cut your internet and still use it. That is the only way to get full privacy.
- Why does it need to be uncensored?
- Every base model is trained to refuse certain requests. An uncensored (abliterated) version has that refusal behavior removed from the weights, not prompted around. I tested the same base model without abliteration first and it flat-out refused things the abliterated one handles without issue. It's still a general chat model, not a tool for generating explicit media.
- Is it really private?
- Yes. Once the model is downloaded, nothing is sent anywhere. It runs as a local process and the chat page talks to it over your own machine's loopback address.
- Why point to your own chat page instead of just the raw Ollama setup?
- The raw setup gives you a blank model with no character. The chat page adds a name, an avatar, and a persistent personality prompt on top, so you get an actual companion instead of rebuilding that in the terminal every time.
Keep reading

How to Set Up a Local NSFW Image Model: Full Tutorial
Install an adult image-generation workflow on a rented cloud GPU: RunPod, ComfyUI, a real workflow file, and every step in between, screenshot by screenshot.

How to run MiniMax H3 locally in ComfyUI
Install MiniMax H3 in ComfyUI: the right FP8/INT8 files for your GPU, the license restriction to check first, and all three workflows, step by step.

What is a LoRA and when do you need one
A plain-language explanation of LoRAs for AI image generation: what they add to a base model, why they only work on open-weight models, and what weight to load one at.
Get new guides by email
One email when we publish new guides and model breakdowns. No spam, unsubscribe anytime.
