Welcome To The ScrollBuilding My Local AI Brain: A Beginner's Research Journey

The short version

A blog post about the AI bubble sent me down a weeks-long rabbit hole. The mission: stop renting AI and start owning it — a machine on my desk that keeps working even if every AI company evaporates tomorrow.

The original dream was a ₱145,000 Mac Mini running a 141-billion-parameter monster. The research killed that dream — and replaced it with something better. Along the way: what "8x22B" actually means (not what you think), why a hospital is the best metaphor for modern AI, how a response is born one token at a time, and the moment I realised I was speccing a server rack for a laptop's worth of work.

This is the research phase — everything I learned before clicking "Buy." The build comes next.

Read the full journey

A Beginner's Research Journey

Building My Local AI Brain

It started with a blog post that hit me like a truck.

No API keys. No uptime anxiety. No "service temporarily unavailable" at 11 PM the night before a deadline. This is how I got from that thought to a plan.

The Spark: Why I Wanted a Local AI

It started with a blog post that hit me like a truck: "AI Bubble Survival Plan".

The gist? We're living through an AI gold rush, and most of us are renting our picks and shovels from the same few companies. What happens when the bubble bursts, the VCs pull funding, the APIs get shut down, or the subscription prices triple overnight? What happens when the cloud AI you depend on for your daily workflow suddenly goes dark — or worse, gets acquired and dismantled?

The answer: you stop renting. You start owning.

That blog post became my manifesto. This local AI build isn't just a side project or a privacy flex — it's my first step in a survival plan. I want a setup where, even if every AI startup on Earth evaporates tomorrow, I can still code, write, research, and deliver work at full speed. No API keys. No uptime anxiety. No "service temporarily unavailable" at 11 PM the night before a deadline.

So yes, I want privacy. Yes, I want to own my data. But more than anything, I want resilience. I want to be the person who keeps shipping while everyone else is scrambling to find a new AI provider.

Of course, I'm a developer, not an AI researcher. I know how to build web apps, wrangle APIs, and debug JavaScript at 2 AM. But the world of local LLMs? That felt like walking into a hardware store where everyone speaks in tongues — quantization, MoE, VRAM, GGUF, parameter counts, context windows.

So I did what any reasonable beginner does: I went down a rabbit hole. Hard.

Phase 1: The Ambitious Plan (a.k.a. "Go Big or Go Home")

My first thought was simple and naive: I want the best model available.

If I'm building a local AI, why settle? I kept seeing the name Mixtral 8x22B pop up in forums and benchmark charts. It was crushing it on coding tasks, reasoning, and long-context understanding. The name alone sounded powerful — like something out of a sci-fi movie.

The Build — v1

Machine
Mac Mini M4 Pro, 64GB Unified Memory
Model
Mixtral 8x22B (the full thing, no compromises)
The Dream
A local ChatGPT killer running silently on my desk

I started pricing it out. The Mac Mini M4 Pro with 64GB RAM and 1TB SSD sits around ₱125,000–₱145,000 (roughly $2,000–$2,300 USD). Not cheap, but I told myself it was an investment. One machine to rule them all. No more subscription fees. No more "service temporarily unavailable."

But then I started asking why I needed 8x22B. And that's when the rabbit hole got deeper.

Phase 2: Understanding What "8x22B" Actually Means

Here's where my beginner brain got its first reality check. I thought Mixtral 8x22B meant the model had 8 times 22 billion parameters — so, 176 billion total. And I thought all 176 billion parameters woke up every time I asked it a question.

I was wrong.

~141BTotal parameters (not 176B — the naming convention is weird)
~39BActive parameters per token
8Distinct "expert" neural networks inside
~22BParameters per expert
RouterExpert 1Expert 2Expert 3Expert 4Expert 5Expert 6Expert 7Expert 8

8 experts, 2 awake — ~39B of ~141B active at any moment.

When you send a prompt, the model doesn't wake up all 141 billion parameters. Instead, a tiny routing network looks at your input and says: "For this specific word/token, these 2 experts are the most qualified to handle it." It activates only those 2 experts (out of 8), processes the token, and moves to the next one.

Think of it like a hospital. You don't wake up every single doctor for every patient. The triage nurse (the router) directs each case to the right specialist. The cardiologist handles heart stuff. The neurologist handles brain stuff. They don't all crowd around a broken ankle.

This is why MoE models are so efficient — and why the naming is confusing. 8x22B doesn't mean 176B active. It means 8 experts of ~22B each, with only ~39B active at any moment.

Why This Matters for Local Hosting

If I had actually needed to load all 141 billion parameters into memory at full precision (FP32), I'd need ~564 GB of VRAM (4 bytes per parameter × 141B parameters).

That's not a Mac Mini. That's not even a high-end gaming PC. That's a server rack.

But because of the MoE architecture and quantization (compressing the model's precision from 32-bit to 16-bit, 8-bit, or even 4-bit), the actual requirements drop dramatically:

Mac Mini 64GBFP32 (full)564 GBFP16282 GBQ4_K_M (4-bit)80 GBQ4_0 (aggressive 4-bit)~65 GB

Even at aggressive quantization, 60–80 GB is still beyond what a 64GB Mac Mini can handle comfortably. The Mac Mini's unified memory is shared between CPU and GPU, and macOS itself needs a chunk. You'd be swapping to SSD constantly, which turns your AI into a very expensive paperweight.

The Pivot — What Do I Actually Need?

This is where I had to get honest with myself. I was planning a ₱145,000 machine to run a model I didn't actually need.

So I asked myself three questions:

1

What will I use this AI for? Coding assistance (Python, JavaScript, Three.js). Writing and brainstorming blog posts. Learning new concepts through conversation. Occasional document analysis. Not: running a business, serving multiple users, or training models.

2

How big are my context windows? A "context window" is how much text the model can "remember" at once. If you're pasting entire codebases or 50-page PDFs, you need a big window (128K+ tokens). For my use case? I rarely paste more than a few hundred lines of code. I don't need to ingest War and Peace in one prompt.

3

Do I need the absolute best model, or a good enough one? I kept reading that Mixtral 8x7B — the smaller sibling — was already excellent at coding, reasoning, and general conversation. And Mistral Small 24B (a dense model, not MoE) was optimized specifically for coding tasks with a much smaller footprint. The kicker? Mistral Small 24B runs beautifully on a Mac Mini with 32GB of RAM. No quantization gymnastics. No swap-file anxiety. Just download, load, and chat.

The Realization

I was suffering from spec-itis — the beginner's disease of buying more than you need because bigger numbers feel safer. But AI doesn't work like that. The right model for your workload is better than the biggest model you can barely run.

The Build — v2

Machine
Mac Mini M4 Pro, 32GB Unified Memory (or even base M4 with 24GB)
Model
Mistral Small 24B or Mixtral 8x7B (Q4 quantization)
The Savings
~₱50,000–₱70,000 ($800–$1,100 USD)
The Performance
Near-instant responses, no thermal throttling, room to grow

Phase 4: How Local AI Actually Works (The Beginner's Guide)

Before I committed to hardware, I needed to understand the pipeline. Not the marketing fluff — the actual mechanics. Here's what I learned:

The Model Weights (The "Brain")

The model file is essentially a massive matrix of numbers — billions of floating-point values that encode everything the AI "knows." These are static. They don't learn from your conversations (unless you fine-tune, which is a whole different beast). Stored as .gguf or .safetensors files. Downloaded once, loaded into RAM/VRAM on startup. Size depends on parameter count × precision (quantization).

The Inference Engine (The "Runner")

The software that actually executes the model. It takes your prompt, feeds it through the neural network, and generates tokens (words) one at a time. Ollama — dead simple, one command, perfect for beginners. LM Studio — GUI-based, great for experimenting. vLLM / llama.cpp — faster, more configurable, steeper learning curve.

The Frontend (The "Face")

What you actually type into. A terminal window (raw Ollama). A web UI like Open WebUI or ChatGPT-Next-Web. An IDE plugin (Continue.dev for VS Code). Or a custom app you build.

The Token Pipeline (How a Response Is Born)

When you type "Explain quantum computing like I'm five," here's what happens:

×700–800 passes per reply1Tokenization2Embedding3Attention4Feed-Forward(The Experts)5Prediction6Sampling7Detokenization

Tokenization

Your text is split into tokens (words or word-pieces). "Quantum" might be one token. "Computing" might be two.

Embedding

Each token is converted into a high-dimensional vector (a list of numbers) that represents its meaning in the model's "brain space."

Attention

The model looks at all tokens in your prompt and figures out which ones relate to each other. "Quantum" pays attention to "computing." "Explain" pays attention to "like I'm five."

Feed-Forward (The Experts)

In an MoE model like Mixtral, the router sends each token to its assigned experts. Each expert is a specialized neural network that processes the token and updates its representation.

Prediction

The model looks at the final representation and predicts the most likely next token. It generates "Quantum" → "computing" → "is" → "like" → "magic" → ... one token at a time.

Sampling

The model doesn't always pick the #1 most likely token. Temperature and top-p settings add randomness, so responses feel creative rather than robotic.

Detokenization

The generated tokens are converted back into human-readable text and streamed to your screen.

The whole process repeats for every single token. A 500-word response might require 700–800 forward passes through the model. That's why speed matters — and why VRAM bandwidth is just as important as capacity.

Phase 5: The Hardware Reality Check

I spent days comparing setups. Here's the distilled version of what I learned:

Option A: The Mac Mini Route (What I Almost Bought)

Specs: M4 Pro, 32–64GB Unified Memory, 1TB SSD. Pros: silent, tiny, unified memory means no CPU/GPU copy bottleneck, excellent per-watt performance. Cons: no CUDA (NVIDIA's GPU language), limited to Metal-compatible inference engines, memory is non-upgradeable. Best for: developers who want a clean, quiet desk setup and don't need the absolute fastest inference.

Option B: The NVIDIA PC Route (The Power User Path)

Specs: RTX 4090 (24GB VRAM) or dual RTX 3090s, 64GB+ system RAM. Pros: CUDA ecosystem is massive, fastest inference for large models, upgradeable. Cons: loud, power-hungry, expensive, takes up desk space. Best for: people running 70B+ models, doing fine-tuning, or serving multiple users.

Option C: The Hybrid (The "Best of Both Worlds")

Specs: MacBook Air for daily work + remote SSH into a headless Linux/NVIDIA box at home. Pros: portable laptop + desktop-grade AI power. Cons: two machines to maintain, network dependency. Best for: people who need mobility but want home horsepower.

My Decision

I'm going Option A — but the smaller one. A base Mac Mini M4 with 32GB RAM and a 512GB SSD. I'll add external storage later if needed.

Why? Because Mistral Small 24B fits comfortably in 32GB. Mixtral 8x7B (Q4) also fits with room to spare. I don't need to run 70B+ models. I value silence and desk space. And I can always upgrade or add an eGPU later.

Phase 6: The Software Stack (What I'll Actually Run)

Model Runner

Ollama. One-command install, handles downloading, simple API.

Chat UI

Open WebUI. Self-hosted ChatGPT interface, supports multiple models.

IDE Integration

Continue.dev (VS Code). Inline AI suggestions, chat sidebar, uses local Ollama.

Document RAG

AnythingLLM or LangChain. Upload PDFs, ask questions about them.

System Monitor

Activity Monitor / asitop. Watch memory and GPU usage.

The Setup Flow

1

Install Ollama: curl -fsSL https://ollama.com/install.sh | sh

2

Pull a model: ollama pull mistral-small

3

Test it: ollama run mistral-small

4

Install Open WebUI via Docker for a pretty interface

5

Configure Continue.dev in VS Code to point at http://localhost:11434

That's it. No cloud accounts. No API keys. No usage quotas.

Ad placeholder

Final Thoughts: What I Really Need (And Why I'm Doing This)

I'm a solo developer who wants a coding assistant that knows my codebase, a writing partner that doesn't judge my drafts, a research tool that doesn't cost $20/month, and privacy — my prompts stay in my house.

I don't need to run the biggest model on earth. I need a reliable, fast, private model that I can talk to without thinking about token costs. Mistral Small 24B or Mixtral 8x7B will handle 95% of what I throw at them. The remaining 5%? I'll use cloud APIs for specialized tasks. That's not failure — that's smart resource allocation.

What's Next

This was the research phase. The next blog — if I write it — will be the build phase:

Unboxing the Mac Mini. Installing Ollama and pulling my first model. Setting up Open WebUI. Benchmarking inference speed. Integrating with VS Code via Continue.dev. The first "holy shit, this actually works" moment.

Will I actually build it? I think so. The research is done. The plan is solid. The only thing left is to click "Buy" and start typing ollama run mistral.

If you're on the fence about local AI, my advice is this: start smaller than you think. You don't need a $3,000 rig. You need a clear understanding of what you're building for, the humility to admit you don't need the biggest model, and the curiosity to learn how the sausage is made.

The AI revolution doesn't have to live in the cloud. It can live on your desk. Quietly. Privately. Resilient.

When the bubble pops, I'll still be here — typing away, model loaded, no API key required. That's the plan. This is just step one.

— Written by a developer who finally understands what "8x22B" actually means.

Ad placeholder