Local LLM · RTX PRO 6000 Blackwell · 192GB VRAM · 2026-09
Go above an RTX 5090 and the field splits into professional and data-center territory. With a budget of ¥6 million (about $41,000), the answer converges on one build: two NVIDIA RTX PRO 6000 Blackwell cards (96GB each) — 192GB of VRAM total, enough to run Qwen3-235B-class MoE models and Llama 405B-class dense models on the desk in your home. This page is the blueprint: parts list, reasoning, and the electrical realities.
For cheaper tiers (dual 5090, dual used 3090), see the used RTX 3090 PC finder. This page covers the ceiling.
The verdict (September 2026 prices)
An RTX PRO 6000 Blackwell 96GB runs about ¥2.3–2.5M ($16,000-ish in the US). Two of them is roughly ¥4.8M; a Threadripper platform fills out the rest for a grand total around ¥5.6–5.9M. That buys 192GB of VRAM and 3.6TB/s of combined memory bandwidth — 3× the VRAM of a dual-5090 build, and unlike data-center cards (H100 etc.) it fits in a normal PC case and runs on ordinary GeForce-style drivers. This is the practical ceiling of what an individual can buy.
Estimated from Japanese retail (authorized and parallel imports), including tax and shipping.
| Part | Pick | Approx. price |
|---|---|---|
| GPU | NVIDIA RTX PRO 6000 Blackwell 96GB ×2 (PNY/Leadtek etc. — dual-slot workstation cards, 400W each, one 12V-2x6 cable per card) | ¥2.3–2.5M ×2 = ¥4.6–5.0M |
| CPU | AMD Ryzen Threadripper 7970X (32 cores) 128 PCIe 5.0 lanes — gives both GPUs x16/x16 | ~¥0.5M |
| Motherboard | TRX50 chipset (e.g. ASUS Pro WS TRX50-SAGE) The SAGE-style 7-slot boards leave a gap between cards | ~¥0.12–0.15M |
| RAM | DDR5 RDIMM 128GB (4×32GB) Spillover target when offloading MoE experts to CPU | ~¥0.12–0.15M |
| PSU | 1600–1700W ATX 3.1 with 12V-2x6 ports Two GPUs plus CPU can peak past 1.2kW; 1600W is the floor for sustained load | ~¥0.07–0.09M |
| Storage | 4TB NVMe SSD (system + models) A 235B-class GGUF is 130GB+ on its own | ~¥0.06–0.08M |
| Case & cooling | Silent full tower (e.g. Fractal Define 7 XL) + 360mm AIO Airflow-first case selection | ~¥0.05–0.07M |
| Total | — | ~¥5.6–5.9M |
The two GPUs are 80% of the body price. Since launch in 2025 (about ¥1.3M/card) they have roughly 1.8×'d in price, so ¥6M is a tight fit; if this tier cools off, the same build lands around ¥4.5M.
Above the 5090, your options split into pro-workstation and data-center cards.
| Candidate | VRAM (each) | Verdict for this build |
|---|---|---|
| RTX PRO 6000 Blackwell | 96GB | The pick. GeForce-style drivers, dual-slot 400W cards that fit a normal case, ECC GDDR7 at 1.79TB/s |
| RTX 6000 Ada / A6000 (used) | 48GB | Runner-up. 96GB for two — but no room for 235B-class |
| H100 80GB (used) | 80GB | Wrong tool. Over ¥4M for two, plus blower noise, exotic power, and awkward form factors |
| A100 80GB (used) | 80GB | Wrong tool. Cheaper per card but everything around it costs you |
| RTX 5090 ×2 | 32GB | Alternate. 64GB for well under ¥1M total — one model tier lower |
Data-center cards may win on VRAM-per-yen, but they lose on cooling, power, drivers, and installation — all four, for an individual. The PRO 6000 is a workstation card, built on the assumption that it works on a desk, and all four of those are as tame as a GeForce. That's the whole argument.
| Model (examples) | Approx. size | On this rig |
|---|---|---|
| Qwen3-235B-A22B-class (MoE) Q4_K_M | ~130–140GB | Fits. Only 22B of experts are active per token, so it's fast for its size — the flagship use of this build |
| Llama-3.1-405B-class Q3 | ~170GB | Just fits. Q4 (~230GB) does not |
| 70B-class dense Q8 / fp16 | ~75GB / ~141GB | Comfortably. Uncensored (abliterated) 70Bs at high precision |
| DeepSeek-R1/V3-class (671B MoE) Q2_K | 200GB+ | Overflows. Ultra-low-bit quants squeeze in, at a quality cost |
| 27–32B-class (mainstream uncensored) | ~16–20GB | Trivially. Even 100k+ token contexts leave headroom |
Token generation speed is bounded by total memory bandwidth ÷ bytes read per token. Two cards give ~3.6TB/s. Qwen3-235B-A22B reads ~13GB per token at Q4 (22B active experts), so the theoretical ceiling is ~270 tokens/s — but only with tensor parallelism (vLLM and similar) keeping both cards busy at once. With llama.cpp's default layer split, the layers run on one card after the other, so a single conversation is capped by one card's 1.79TB/s, about 135 tokens/s. Even 20–40% below those ceilings in practice, it outpaces reading. These are spec-sheet estimates — no hands-on benchmark was done for this page.
| Budget | Core of the build | VRAM | Model ceiling |
|---|---|---|---|
| ~¥5.8M | This page: PRO 6000 96GB ×2 | 192GB | Qwen3-235B Q4 / 405B Q3 |
| ~¥3M | PRO 6000 96GB ×1 | 96GB | 70B fp16 / 235B Q2-class |
| ~¥0.9M | RTX 5090 32GB ×2 | 64GB | 70B Q4 / ~120B MoE |
| ~¥0.5M | Four routes compared: V100 32GB / modded 3080 20GB ×2 / used RTX 3090 ×2 / EVO-X2 | 32–48GB | Fast 27B or 70B Q4 (3090 hunting guide) |
| ~¥0.24M | Used 3090 desktop PC | 24GB | 27–32B Q4 (pc3090) |
At any tier: compute the Q4_K_M size of the model you actually want plus context plus 20% headroom, and buy one tier above that. That's the whole strategy.
With used 3090s holding at ~¥200k each, dual 3090s (48GB) is now just one option — the budget-maxed generalist. Priced September 2026; speeds are a mix of spec-sheet math and published measurements.
| Route | Approx. cost | Memory | 27B Q4 speed | 70B Q4 | Video gen | In one line |
|---|---|---|---|---|---|---|
| ① V100 32GB (SXM2 + blower shroud) | ¥150–350k | 32GB HBM2 900GB/s | ~70 tok/s | No (40GB) | No | Fastest yen-per-token for chat. Conditions: CUDA 12.x only (CUDA 13 dropped Volta), no surrounding ecosystem. SXM2 needs a converter board + fan shroud; the PCIe version costs twice as much |
| ② Modded 3080 20GB ×2 | ¥250–280k | 40GB GDDR6X 760GB/s ×2 | Comfortable | Just fits | Tight (20GB×2) | Often rebuilt with 4090-class coolers, so condition tends to beat tired mining-era 3090s. But domestic supply is a trickle (a few listings/month), no warranty, 320W each |
| ③ Used RTX 3090 ×2 | ¥500–520k | 48GB GDDR6X 936GB/s ×2 | Comfortable | Fits | Yes (+64GB RAM advised) | The generalist: LLM and video. Biggest risk is worn mining-era boards (condition checklist at pc3090) |
| ④ EVO-X2 RAM64 (Ryzen AI Max+ 395) | ~¥350k | 64GB unified 256GB/s | 10–15 tok/s | ~5 tok/s | Weak | A big-MoE appliance: gpt-oss-120b at usable speeds (32 tok/s on the 128GB version, 235B at 11). You give up CUDA (ROCm/Vulkan). Nothing else is this small, quiet, and frugal |
How to choose: fastest uncensored-27B chat → ① V100. LLM and video generation (Wan-class) → ③ 3090×2, no real alternative. Big MoE models idling under RAG/agent workloads → ④ EVO-X2. ② is the fallback when you can't win a V100 — spend the change on storage and RAM.
If you buy a V100
SXM2 cards are passively cooled: a blower-fan shroud conversion is mandatory (kits and 3D-printed shrouds are widely available). Run llama.cpp on CUDA 12.x builds; don't count on Flash-Attention-class optimizations (Volta has fp16 tensor cores only). Power is a plain 250W per card on standard PCIe 8-pin.
Update (2026-09-23): sizing for Qwen4-27B
At the Apsara Conference (Sep 22–24, 2026) Alibaba announced the Qwen4 lineup (Max / Plus / Flash / 27B). The 27B isn't released yet and specs are unannounced, but recent 27B-class models (Qwen3.6-27B: ~19GB at Q4, ~30GB at Q8) make the preparation obvious. Abliterated builds lose a little capability to the pruning, so you want to run them at Q5/Q6 or higher — meaning "Q4 fits" is not the bar. Q6 plus long context is: 32GB (V100 / 5090) is the comfortable Q4 line, and 40–48GB (modded 3080 20GB ×2 / 3090×2) is the right size for high-quant + long context. One more reason to buy one rung up the ladder if 27B-class is your main battlefield. And if Qwen4-27B turns out to be MoE (the Flash-Next architecture is the rumored preview), total size could jump — that lands in ④EVO-X2 (128GB) or flagship-rig territory.