AIMA AMD395 Qwen3.6 35B Windows Engine — Native Windows Inference
Native Qwen3.6-35B-A3B BF16 inference on AMD Ryzen AI Max+ 395: 8K prefill 2,126 token/s, decode 30.55 token/s, 131K prefix-reuse scenario TTFT 8.59s.
Personal projects and works
Native Qwen3.6-35B-A3B BF16 inference on AMD Ryzen AI Max+ 395: 8K prefill 2,126 token/s, decode 30.55 token/s, 131K prefix-reuse scenario TTFT 8.59s.
Intel PTL/Arc B390 OpenVINO U4 engine for Qwen3.6-35B-A3B: prefill ≥1.48×, decode ≥1.59× vs stock GPU; 10,752 greedy tokens bit-for-bit identical.
Native BF16 engine for Qwen3.6-35B-A3B on AMD Ryzen AI Max+ 395. No Python/PyTorch/vLLM/system ROCm. 262K ctx, OpenAI API, SSE streaming, function calling.
Make AI inference deployment as simple as installing an app. AIMA automatically detects hardware, selects optimal configurations, and deploys with one click, turning every machine into an AI inference node.
Say goodbye to SSH and manual troubleshooting. Connect devices with one command, describe requirements in natural language, and let the AI Agent automatically handle diagnostics, installation, repairs, and upgrades.
Compress 1 hour of Feishu app configuration into 5 minutes. One-click automation completes the entire workflow of app creation, permission configuration, and OpenClaw integration—let developers focus on products, not platform integration.