<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Local Inference</title><link>https://llms.zer0contextlost.net/</link><description>Recent content on Local Inference</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Sun, 09 Aug 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://llms.zer0contextlost.net/index.xml" rel="self" type="application/rss+xml"/><item><title>About</title><link>https://llms.zer0contextlost.net/about/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://llms.zer0contextlost.net/about/</guid><description>&lt;p&gt;Local Inference covers self-hosting large language models: picking
hardware for your VRAM budget, GGUF quantization tradeoffs, backends
(llama.cpp, KoboldCpp, Ollama), and frontends, plus what the local-model
community actually recommends for different use cases — coding, general
chat, and creative/roleplay writing among them.&lt;/p&gt;
&lt;p&gt;No hosted-API affiliate kickbacks driving the recommendations here — if a
model or tool gets mentioned, it&amp;rsquo;s because it came up repeatedly in the
communities that actually run this stuff (r/LocalLLaMA, r/LocalLLM,
SillyTavern&amp;rsquo;s own community) or was tested directly.&lt;/p&gt;</description></item><item><title>Running Kimi K3 Locally: A 16x GB10 Cluster and the Community's Alternatives</title><link>https://llms.zer0contextlost.net/posts/running-kimi-k3-locally-a-16x-gb10-cluster-and-the-community/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://llms.zer0contextlost.net/posts/running-kimi-k3-locally-a-16x-gb10-cluster-and-the-community/</guid><description>&lt;!-- SOURCE_THREAD: https://old.reddit.com/r/LocalLLaMA/comments/1vfl525/kimi_k3_full_model_running_on_16x_gb10_cluster_at/ --&gt;
&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;A recent r/LocalLLaMA post showed the full Kimi K3 model running across a cluster of 16 NVIDIA GB10 nodes, the same Grace Blackwell Superchip found in devices like the DGX Spark. The build runs its monitoring dashboard off a Raspberry Pi, which amused a string of commenters given the price gap between the Pi and the rest of the rig; the joke spun off into a side debate about whether Raspberry Pis are even still worth buying given how well used corporate desktops compare on price and speed.&lt;/p&gt;</description></item><item><title>Self-Hosted LLMs for Fiction and Roleplay Writing: Hardware, Models, and Stack</title><link>https://llms.zer0contextlost.net/posts/self-hosted-llms-for-fiction-and-roleplay-writing/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://llms.zer0contextlost.net/posts/self-hosted-llms-for-fiction-and-roleplay-writing/</guid><description>&lt;p&gt;Hosted chat APIs apply content filters tuned for a general audience, which
means anything even adjacent to mature fiction — violence, adult themes,
morally gray characters — can get flagged, refused, or silently softened
mid-story. That&amp;rsquo;s the main reason r/LocalLLaMA and r/LocalLLM both have a
whole recurring thread genre around &amp;ldquo;best uncensored model for writing.&amp;rdquo;
Run the model on
your own hardware and there&amp;rsquo;s no third-party filter sitting between your
prompt and the output, no chat log leaving the house, and no surprise
policy change breaking a story you&amp;rsquo;re partway through.&lt;/p&gt;</description></item><item><title>What Qwen3.8's MoE Lineup Means for Sizing Your Local Rig</title><link>https://llms.zer0contextlost.net/posts/what-qwen3-8-s-moe-lineup-means-for-sizing-your-local-rig/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://llms.zer0contextlost.net/posts/what-qwen3-8-s-moe-lineup-means-for-sizing-your-local-rig/</guid><description>&lt;!-- SOURCE_THREAD: https://old.reddit.com/r/LocalLLaMA/comments/1ve0psn/qwen3827b_announced_alongside_qwen38max/ --&gt;
&lt;p&gt;Qwen released 3.8-27B alongside a much larger 3.8-Max, and the r/LocalLLaMA thread reacting to it turned into a decent snapshot of what people are actually planning to run on their own hardware. Buried under the celebration (&amp;ldquo;still waiting for 3.8 35b a3b, but excited for the 27b version,&amp;rdquo; one top comment put it) is a pretty clear map of which model shapes this community wants and why.&lt;/p&gt;</description></item></channel></rss>