

That this makes no sense is the whole point. The goal is to show how little sense this makes
You americans are weird, ngl
But you do you guys. Hell, I would only benefit if that works, big corp mostly looks at USA and less on EU when deciding on their policies/strategy. If USA mandate open hardware, rest will follow


Context limit is not really a problem on local models. Qwen3.6 can do up to ~256k tokens, it’s not that far from what things like cursor uses. grok in cursor have exactly 256k tokens for example. Also, you can use opencode with custom config, where you need to set trimming close to that number.
That being said, quality wise it still kinda shit and looses to paid models. I think you need something like GLM 5.2 to compete, which needs 228gb at lowest 1bit quant (not including context size, which can be up to 1m tokens, so you can round up requirements for RAM up to 256gb), so yeah, the only thing that limits you locally — the fact that “AI” companies bought all supply of focking ram and we can’t afford any.