Qwen3.8-Flash-Next (IQ3\_XXS, 76 GB) at the full 262K context on 31 GB of RAM: 1,800 tok/s prompt, 130 tok/s decode, on a 5090 + a 4070 Ti SUPER on a PCIe x1 slot Strata (github.com/Niko1221/Strata) streams MoE experts from an mmap'd GGUF, and its own sizing rule says RAM >= expert shard + 10…
What it is: an Android app (PixelUnlockGPU) that turns a Pixel into an OpenAI-compatible HTTP server. Standard \`/v1/chat/completions\` with streaming, so it talks to TypingMind or any OpenAI client directly — no cloud, no subscription, model runs entirely on-device via LiteRT-LM. Device: Pixel 10…
Hi! I'm currently working on creating Doom from scratch, in full JS. It is not a WASM project, it is really written in JavaScript, with my own 3D engine. It can be installed as a PWA on computer or mobile, and I think it is better to use it like that. You can play with keyboard + mouse, gamepad, or…
I built **Claude Fables**, a mod for the Claude Code desktop app. While Claude works, it watches each tool call and each line Claude says. Every few seconds it retells the latest moment as a short art scene in the band above the prompt. A bug hunt becomes a nature documentary, a failing test makes…
Last week, I ran out of tokens on my Claude Code subscription. Again. As someone who codes quite a lot using LLMs (like most of us these days, let’s face it), it got me thinking: Why are we spending so many tokens on programming language syntax designed primarily for humans? Braces, closing quotes,…
I like to use Deepseek Harness as my vibe coding harness. It's great and support local models. The issue is that most models I can run locally have context size of about 262K. Therefore, for long coding sessions, I need a memory management tool. DSH comes with a context compaction tool that I can…