I Paid for a Top-Tier Model. It Got Dumber and Marked Me
Codex GPT-5.5 clusters reasoning at 516; Claude Code secretly marks Chinese users. When cloud AI silently degrades, open-weight self-hosting is the hedge.
Model labs, open vs. closed, where capability is heading
Codex GPT-5.5 clusters reasoning at 516; Claude Code secretly marks Chinese users. When cloud AI silently degrades, open-weight self-hosting is the hedge.
GPT 5.6 looks strong, but only ~20 U.S.-approved partners can use it; I realize intelligence you can buy is already fair, while permission is the real gap.
I found Fable 5 swapped to Opus 4.8; U.S. Commerce banned it in three days. Closed-source AI can vanish overnight; Chinese open-source models step in.
Fable 5 ran a 15-hour task for $420. After June 22 it leaves Coding Plan for API rates, a 10x cost jump: access to top intelligence is a wealth problem.
Top models tie on GPQA at 92–94%, yet real-world results diverge. Three skills now split: hard problems, rough tasks, open exploration; the last is hardest to measure.
Claude Opus 4.8 kept failing my research and engineering tasks, so I moved them to GPT 5.5. Top models rotate—NVIDIA, AMD, and Zhipu prove survival wins.
I used Codex's /goal for weeks, then tested Claude Code's new /goal on tasks Codex failed; the behavioral difference is striking enough to document.
A friend asked if I'd take an AI transformation role. My take: don't save money, avoid IDEs, keep agents inspectable. GPT-5.5, two weeks old; 10 weeks left.
DeepSeek's 300B round didn't surprise me. Liang Wenfeng investing 20B at the high valuation he set with outside investors did—something I haven't seen before.
Codex APP agents burned a Pro account in two days for little gain; Codex CLI Goal fixed it. I blamed the model; agents are maturing; no screen babysitting.
After a week with GPT-5.5 and Opus 4.7, I learned stronger models don't remove the human burden: Claude paused, GPT-5.5 killed a direction; I still steered.
Two weeks with GPT-5.5 (SPUD): I've mostly dropped Claude Code. The leap isn't a few percent gain—old weaknesses fixed make agent design philosophy clear.
GPT-5.5 arrived six weeks after 5.4; Opus 4.7 followed 4.6 in two months. GPT now sounds human, Opus doesn't, and judgment's half-life keeps shrinking.
DeepSeek V4 matches Opus 4.6, but FP4, 1M-token context, and day-0 chip support stress inference infra. GPT-5.5, Vision Banana, and LPM 1.0 landed too.
LLM pricing is stuck: chip controls cap supply, while three user groups pull demand into different shapes. The once-obvious Coding Plan is now under fire.
Opus 4.7 kept me up. I tested it, merged OpenClaude PRs, and pushed Strix Halo Qwen3-30B prefill to DGX Spark levels. Agents make parallel work real.
HappyHorse topped video charts anonymously, yet squatters took its domains and HuggingFace name. The real issue: the hardware economics of private text-to-video demand.
How do open-source models make money? DeepSeek's numbers show model and cloud companies converging, with open source as the cheapest acquisition channel.
Researching text-to-video, I found Chinese dominance in open-source LLMs hasn't reached every domain. From LLaMA to Qwen to DeepSeek, what changed?
When creation and distribution costs near zero, platforms explode. Short video proved it; software repeats the pattern, and OpenClaw is worth watching.
I switch models daily. Opus is reckless but strong, GPT 5.4 drifts, Gemini misses bugs, domestic models each have quirks. Switching beats prompt tuning.
We open-sourced AIMA, a Go binary with 57 MCP tools and a YAML knowledge base for heterogeneous AI inference, as AI servers halve in value in three months.
Get notified when I publish new posts. No spam, ever.
Only used for blog update notifications. Unsubscribe anytime.