πŸ’» Run AI Locally

243 articles. Run AI models on your own hardware. Ollama, vLLM, llama.cpp guides. Hardware requirements and self-hosted tutorials.

Running AI models on your own hardware means no per-token bill, no data leaving your machine, and no dependency on a provider staying online. This hub covers what we've written about doing that: which models are practical to self-host, what hardware you need, and how to set up the inference engine that fits your setup. We separate this from general model comparisons because local deployment has its own constraints β€” VRAM budgets, quantization tradeoffs, and inference engine choice matter more here than raw benchmark scores. Guides below include hardware requirements based on hands-on setup where we've run it ourselves, or on vendor documentation and community-reported figures where noted, whether that's a Mac with unified memory, a single consumer GPU, or a self-hosted server for privacy reasons.

Where to start

πŸ¦™ Ollama (88)

The easiest way to run models locally β€” setup guides and model recommendations for Ollama specifically.

Build an AI Documentation Writer β€” Generate Docs From Code Comments

Stop writing docs manually. Build a tool that reads your code and generates comprehensive documentat

Build a Local AI Git Conflict Resolver β€” Auto-Fix Merge Conflicts

Stop manually resolving merge conflicts. Build a tool that reads both sides and uses AI to merge the

Build an AI Docker Compose Generator β€” Describe Your Stack, Get Config

Describe your application stack in plain English and get a complete docker-compose.yml. Uses Ollama

πŸ“– How to Run Locally (46)

Running Bonsai 2 27B Locally: Memory, Runtime and Limits

Deploy Prism ML's Bonsai 2 27B with the correct runtime. Compare weight formats, budget memory separ

Ornith 1.5 35B-A3B: Specs, Benchmarks and How to Run It Locally

Ornith 1.5 is a 35B MoE coding and agent model with about 3B active parameters, 262K context, vision

Running Hermes Agent Locally with LFM2.5-2.6B (2026)

Liquid AI documents Hermes Agent as a supported harness for LFM2.5-2.6B. Full local setup: serve the

πŸ–₯️ Hardware & VRAM (42)

I Used RunPod for a Week β€” The AI-Focused GPU Cloud

Week 24 of my AI tool series. RunPod is a GPU cloud designed for AI workloads. After a week, here's

Meta Muse Glimmer 30B: Running Local Multimodal AI Agents on Consumer Hardware (2026)

Muse Glimmer 30B is Meta's Apache 2.0 multimodal model for local agents. Architecture, hardware need

Best NPU-Powered Mini PCs for Local AI in 2026

The best mini PCs with dedicated NPUs for local AI inference. Intel NPU, AMD Ryzen AI, and Qualcomm

⚑ Inference Engines (3)

🏠 Self-Hosted & Privacy (64)

For teams that need AI infrastructure to stay on their own network, including GDPR-relevant setups.

I Used Open WebUI for a Week β€” The Self-Hosted ChatGPT Alternative

Week 25 of my AI tool series. Open WebUI is a self-hosted ChatGPT alternative. After a week, here's

MiniCPM5-2B Explained: A Compact Model for Local AI Agents

MiniCPM5-2B is a 2.5B-parameter text model for local assistants, coding agents and tool use. See its

K2 Horizon Explained: Six Open Models From 0.9B to 375B

K2 Horizon spans six Apache-2.0 models from 0.9B dense to 375B-A23B MoE. Compare sizes, openness, co