About
Local Inference covers self-hosting large language models: picking hardware for your VRAM budget, GGUF quantization tradeoffs, backends (llama.cpp, KoboldCpp, Ollama), and frontends, plus what the local-model community actually recommends for different use cases โ coding, general chat, and creative/roleplay writing among them.
No hosted-API affiliate kickbacks driving the recommendations here โ if a model or tool gets mentioned, it’s because it came up repeatedly in the communities that actually run this stuff (r/LocalLLaMA, r/LocalLLM, SillyTavern’s own community) or was tested directly.
Corrections and follow-up questions welcome.