I Paid for a Top-Tier Model. It Got Dumber and Marked Me
Codex GPT-5.5 clusters reasoning at 516; Claude Code secretly marks Chinese users. When cloud AI silently degrades, open-weight self-hosting is the hedge.
Codex GPT-5.5 clusters reasoning at 516; Claude Code secretly marks Chinese users. When cloud AI silently degrades, open-weight self-hosting is the hedge.
Fable 5 ran a 15-hour task for $420. After June 22 it leaves Coding Plan for API rates, a 10x cost jump: access to top intelligence is a wealth problem.
Claude Opus 4.8 kept failing my research and engineering tasks, so I moved them to GPT 5.5. Top models rotate—NVIDIA, AMD, and Zhipu prove survival wins.
Native Qwen3.6-35B-A3B BF16 inference on AMD Ryzen AI Max+ 395: 8K prefill 2,126 token/s, decode 30.55 token/s, 131K prefix-reuse scenario TTFT 8.59s.
Intel PTL/Arc B390 OpenVINO U4 engine for Qwen3.6-35B-A3B: prefill ≥1.48×, decode ≥1.59× vs stock GPU; 10,752 greedy tokens bit-for-bit identical.
Native BF16 engine for Qwen3.6-35B-A3B on AMD Ryzen AI Max+ 395. No Python/PyTorch/vLLM/system ROCm. 262K ctx, OpenAI API, SSE streaming, function calling.