About

Local Inference covers self-hosting large language models: picking hardware for your VRAM budget, GGUF quantization tradeoffs, backends (llama.cpp, KoboldCpp, Ollama), and frontends, plus what the local-model community actually recommends for different use cases โ€” coding, general chat, and creative/roleplay writing among them.

No hosted-API affiliate kickbacks driving the recommendations here โ€” if a model or tool gets mentioned, it’s because it came up repeatedly in the communities that actually run this stuff (r/LocalLLaMA, r/LocalLLM, SillyTavern’s own community) or was tested directly.

Corrections and follow-up questions welcome.