What Qwen3.8's MoE Lineup Means for Sizing Your Local Rig
Qwen3.8-27B and 3.8-Max landed this week. r/LocalLLaMA's reaction says more about active-parameter counts and offload rigs than about the models themselves.
Notes on running your own LLMs at home: hardware, quantization, and the tool stack, without a hosted API in the loop.
Qwen3.8-27B and 3.8-Max landed this week. r/LocalLLaMA's reaction says more about active-parameter counts and offload rigs than about the models themselves.
Why homelabbers run local LLMs for fiction and roleplay writing instead of a hosted API, what 'uncensored' and 'abliterated' actually mean, and which models r/LocalLLaMA, r/LocalLLM, and r/SillyTavernAI actually recommend for it.
A hobbyist ran the full Kimi K3 model across 16 NVIDIA GB10 nodes, and the ensuing r/LocalLLaMA thread digs into the real throughput, hardware costs, and cheaper single-node alternatives commenters proposed.