Topic
AI & LLMs
Models you can run on your own hardware, prompt patterns that ship, agent frameworks that don't catch fire, and the awkward questions nobody answers in the breathless launch posts. Ollama, vLLM, llama.cpp, LocalAI, plus the quieter stuff — embeddings, RAG, evals, and figuring out when the cloud API is actually the right answer. If you'd rather understand the trade-offs than chase benchmarks, you'll feel at home here.
107 articles in this topic.
Featured posts
-
Prompt Injection vs Your Coding Agent
Prompt-level defenses against injection are probabilistic and lose to a patient attacker. Here's the capability-level design that holds, with working configs.
14 min read -
Batch Inference: One Model, 100 Users
Loading model weights, not math, is your GPU bottleneck. Continuous batching with llama-server serves many concurrent requests for close to the cost of one.
12 min read -
The Update That Broke Confidence
A 2026-09-22 patch left a self-hosted model's weights untouched but made it more confident on nonsense input. Six systems compared, real numbers only.
· Updated:18 min read -
Docker MCP Toolkit: Trust the Sandbox?
Docker MCP Toolkit sandboxes AI agent tools in containers. Here's exactly what that isolation stops, where it silently breaks, and how to configure it safely.
12 min read -
Jev vs Von vs SemIf: Real Numbers
Jev, Von, and SemIf scored head to head on real commit history: a frozen 4B general model beat a purpose-built 395M decision model on every task tested.
14 min read -
NetAlertX MCP: One Token, Twelve Tools
NetAlertX v26.9.0 ships a built-in MCP server with 12 tools. Learn how to wire it into Claude Code, what it exposes, and when to use Home Assistant instead.
13 min read
All AI & LLMs articles
- Prompt Injection vs Your Coding Agent
- Batch Inference: One Model, 100 Users
- The Update That Broke Confidence
- Docker MCP Toolkit: Trust the Sandbox?
- Jev vs Von vs SemIf: Real Numbers
- NetAlertX MCP: One Token, Twelve Tools
- Skeleton Key MCP vs One Token Per Box
- Jev Retagged 840 Posts For 26 Cents
- Proxmox MCP: Read-Only Beats Root
- Home Assistant MCP: Server and Client
- Komodo MCP vs the Docker Socket
- Nano Banana Is Eating Its Own Images
- Your Coding Agent Can't Draw
- Aider vs Continue: Two Modes of Terminal AI Coding
- Docker Ate 215GB in a Week
- Continue.dev + Ollama: Local Code Assistant for Cheap
- Anubis: Anti-AI-Crawler Proof-of-Work
- pgvector for Local Embeddings
- What Is an AI Harness?
- Multiple Claude Code Accounts, One Config
- The Free Tier Rug Pull
- Real-ESRGAN & Upscaling Tools
- ControlNet & LoRA: Advanced Image Control
- API vs Self-Hosted LLMs: The Real Cost
- Free AI Image Gen vs Your Own GPU
- RAG Beyond Vector Search: BM25, Hybrid, Re-ranking
- How to Stretch a Free LLM Tier
- The Free AI Stack: What $0 Gets You
- Agentic Browsers: Browser-Use & Skyvern Reviewed
- Local Voice Assistant: Whisper + Piper + Home Assistant
- Your Agent Doesn't Need a Shell
- Llamafile: Single-Binary LLMs That Actually Just Run
- DeepSeek V4 Flash: 28 Cents a Million
- Where Should Your Coding Agent Run?
- DIY Perplexity: SearXNG + Local LLM = Private Web Search
- Local Coding Agents Need Less Context
- LM Studio vs Jan vs GPT4All: Desktop LLM Clients
- RAGAS: Evaluating RAG Without Vibes
- KV Cache Quantization: Free LLM Context, Almost
- Aider & Cline: Terminal AI Coding That Actually Ships
- Mixture of Experts (MoE) for Self-Hosters, Demystified
- Python Libraries Worth Your Time in 2026
- Speculative Decoding: Faster LLMs With a Tiny Sidekick
- Karakeep: Self-Hosted Bookmarks With AI Tagging
- Stop Feeding the AI Your Whole Repo
- RTK vs snip vs lean-ctx: Token Killers
- Hailo-8 vs Coral: AI Accelerators for the Edge
- AI Swarm Audited My 840-Post Blog
- Used GPU Buying Guide for Home Lab LLMs
- Claude Code in a Homelab Workflow
- Self-Host a Local AI Coding Workhorse
- Give Your AI Agent a Cheap Intern
- Claude Code + SearXNG: Private Web Search
- Dify: Visual Agent Workflows
- OpenRouter vs LiteLLM
- Function Calling in Local LLMs
- Gemma 4 vs Qwen3.6
- AnythingLLM as Knowledge Base
- Local Vision LLMs Worth Running in 2026
- MCP Servers: Tools for LLMs
- RAG Evaluation with Ragas
- LLM Distillation Explained
- GPU Passthrough on Proxmox: Run LLMs in a VM
- Open WebUI Tools, Functions and Pipelines Explained
- Self-Supervised Learning Explained
- Continue.dev vs Cody vs Tabby: AI Code Help Without the Cloud
- Qdrant vs Weaviate vs Chroma: Vector DB Showdown
- LangChain vs LlamaIndex: RAG Framework Showdown
- Beyond RAG: When a Virtual Filesystem Works Better
- Running Gemma 4 Locally with Ollama
- 1-Bit LLMs: The Quantization Endgame
- AMD Lemonade: Local LLM Serving for AMD GPUs
- When to Use Structured Output (JSON Mode) in LLMs
- Using AI to Find Security Bugs in Your Code
- LLM Temperature and top_p Explained Without the Math
- GPU Memory Math: Will This Model Actually Fit?
- LLM Backends: vLLM vs llama.cpp vs Ollama
- RAG Chunking: Why Chunk Size Is Everything
- LiteLLM & vLLM: One API to Rule All Your Models
- System Prompts: The LLM Feature Most People Ignore
- LLM Quantization: Q4_K_M Isn't Always the Best Choice
- Running Multiple Ollama Models Without Running Out of RAM
- Piper vs Coqui: Text-to-Speech on Your Own Hardware (Because AWS Polly Charges Per Character Like It's 1999 SMS)
- Context Window vs Token Limit: Not the Same Thing
- The Embedding Model Choice Nobody Explains
- Ollama Keep Alive: Unload Models from VRAM
- RAG on a Budget: Building a Knowledge Base with Ollama & ChromaDB
- Stable Diffusion vs ComfyUI vs Fooocus: AI Image Generation at Home
- n8n + LLM: Building Automations That Actually Think
- Text Generation Web UI vs KoboldCpp: Power User LLM Interfaces
- LangGraph vs CrewAI vs AutoGen: AI Agent Frameworks for Mere Mortals
- Open WebUI vs LibreChat: Self-Hosted ChatGPT Alternatives Compared
- LLM Fine-Tuning for Mortals: LoRA, QLoRA, and Your Gaming GPU
- Ollama Beyond the Basics: Model Management, Custom Models, and Optimization
- Whisper & Faster-Whisper: Self-Hosted Speech-to-Text That Actually Works
- Continue.dev vs Cody vs Tabby: AI Code Assists That Live on Your Machine
- CUDA vs ROCm vs CPU: Running AI on Whatever GPU You've Got
- Flowise vs Langflow: Build AI Pipelines Without Writing a Novel
- n8n vs Node-RED: Automate Everything Without Learning to Code (Much)
- Key Parameters of Large Language Models
- Prompts for Image Generation in Stable Diffusion
- Prompt Engineering for Generative AI 101
- Large Language Model Formats and Quantization
- Exploring the Diverse World of LLM Models
- Ollama: Powerful Language Models on Your Own Machine
- LocalAI: The OpenAI API, Self-Hosted
- Machine Learning models (AI)