What Qwen3.8's MoE Lineup Means for Sizing Your Local Rig
Qwen3.8-27B and 3.8-Max landed this week. r/LocalLLaMA's reaction says more about active-parameter counts and offload rigs than about the models themselves.
Qwen3.8-27B and 3.8-Max landed this week. r/LocalLLaMA's reaction says more about active-parameter counts and offload rigs than about the models themselves.
Why homelabbers run local LLMs for fiction and roleplay writing instead of a hosted API, what 'uncensored' and 'abliterated' actually mean, and which models r/LocalLLaMA, r/LocalLLM, and r/SillyTavernAI actually recommend for it.
A hobbyist ran the full Kimi K3 model across 16 NVIDIA GB10 nodes, and the ensuing r/LocalLLaMA thread digs into the real throughput, hardware costs, and cheaper single-node alternatives commenters proposed.
A r/LocalLLaMA thread on Qwen3.8-27B's VRAM footprint turns into a debate over whether 17GB is really news, plus community tips on offloading, quantization, and MoE alternatives for low-VRAM setups.