Buy an NVIDIA-powered GPU server you own. We build it by hand in Eugene, Oregon and host it; you choose how it runs, and you keep 85% of what it earns.
You hold title, and you have root on the machine. Monetizing is an option, not an obligation.
Earnings depend on marketplace demand: projections, not guarantees.
Cloud rent is a metered bill that never ends and buys you nothing. A machine you own stays yours, and earns whenever you make it available. Hosting it with us costs a fraction of that rent, and on an earning machine it comes out of what the machine makes rather than out of your pocket.
There's a tax difference too: cloud rent is 100% expense, forever. A machine you own is a depreciable business asset. How that applies to you depends on your situation; consult your tax advisor.
Your models run on hardware you own in a US facility you can name: weights, keys, and data never leave machines you control. Hand-built GPU clusters, US-assembled by Andromeda, with N+1 redundant power. Hosted GPUs for Stanford Blockchain, Penn Blockchain, NYU, and Dartmouth Blockchain.
For every model B3IQ routes: the cheapest machine that serves it, the cheapest that fine-tunes it, and the throughput real operators report.
| Model | Params | Serve VRAM | Cheapest machine |
|---|---|---|---|
| Qwen3 8B (llama.cpp) | 8B | not listed | Not on a machine we sell |
| Qwen3 14B (llama.cpp) | 14B | not listed | Not on a machine we sell |
| Qwen3 8B | 8B | not listed | Not on a machine we sell |
| Mistral 7B | 7B | not listed | Not on a machine we sell |
| Granite 3.3 8B | 8B | not listed | Not on a machine we sell |
| Private Chat 8B | 8B | 8 GB | 1x RTX 5090 |
| Qwen Coder 7B | 7B | 8 GB | 1x RTX 5090 |
| DeepSeek-R1 7B | 7B | 8 GB | 1x RTX 5090 |
| DeepSeek-R1 8B | 8B | 8 GB | 1x RTX 5090 |
| Code Workbench 14B | 14B | 16 GB | 1x RTX 5090 |
| Qwen3 14B | 14B | 16 GB | 1x RTX 5090 |
| Phi-4 14B | 14B | 16 GB | 1x RTX 5090 |
| DeepSeek-R1 14B | 14B | 16 GB | 1x RTX 5090 |
| Gemma 3 12B | 12B | 16 GB | 1x RTX 5090 |
| Gemma 3 27B | 27B | 18 GB | 1x RTX 5090 |
| gpt-oss 20B | 20B | 16 GB | 1x RTX 5090 |
| Mistral Small 24B | 24B | 16 GB | 1x RTX 5090 |
| vLLM Qwen3 14B | 14B | 24 GB | 1x RTX 5090 |
| vLLM MiniMax M2 | 230B | 138 GB | 2x RTX PRO 6000 |
| Qwen Quality 32B | 32B | 22 GB | 1x RTX 5090 |
| Qwen3 30B A3B | 30B | 20 GB | 1x RTX 5090 |
| Qwen3 30B A3B | 30B | 32 GB | 1x RTX 5090 |
| Qwen3 Coder 30B | 30B | 20 GB | 1x RTX 5090 |
| Qwen Coder 32B | 32B | 32 GB | 1x RTX 5090 |
| DeepSeek-R1 32B | 32B | 32 GB | 1x RTX 5090 |
| Llama 3.3 70B | 70B | 42 GB | 2x RTX 5090 |
| gpt-oss 120B | 120B | 72 GB | 1x RTX PRO 6000 |
| Qwen3 Coder 8B | 8B | not listed | Not on a machine we sell |
| Devstral Small 24B | 24B | 16 GB | 1x RTX 5090 |
| Qwen3 Coder 27B | 27B | 32 GB | 1x RTX 5090 |
| DeepSeek V4 Pro | 671B | 400 GB | 4x H200 NVL |
| Kimi K2.6 MoE | 1000B | 600 GB | 8x H200 NVL |
| Qwen2.5 7B Instruct | 7B | 8 GB | 1x RTX 5090 |
| Qwen2.5 14B Instruct | 14B | 16 GB | 1x RTX 5090 |
| Qwen2.5 32B Instruct | 32B | 32 GB | 1x RTX 5090 |
| Llama 3.2 Vision 11B | 11B | 8 GB | 1x RTX 5090 |
| LLaVA 7B (Vision) | 7B | not listed | Not on a machine we sell |
| Qwen3 Embedding 8B | 8B | 18 GB | 1x RTX 5090 |
| gte-Qwen2 7B | 7B | 16 GB | 1x RTX 5090 |
| Qwen2.5-VL 7B | 7B | 7 GB | 1x RTX 5090 |
| Qwen2.5-VL 32B | 32B | 24 GB | 1x RTX 5090 |
| Qwen2.5-VL 72B | 72B | 45 GB | 2x RTX 5090 |
| MiniCPM-V 2.6 8B | 8B | 7 GB | 1x RTX 5090 |
| LLaVA-Llama3 8B | 8B | 7 GB | 1x RTX 5090 |
| LLaVA 13B | 13B | 10 GB | 1x RTX 5090 |
| LLaVA 34B | 34B | 24 GB | 1x RTX 5090 |
| BakLLaVA 7B | 7B | 6 GB | 1x RTX 5090 |
| Mistral Small 3.1 24B (Vision) | 24B | 18 GB | 1x RTX 5090 |
| Llama 3.2 Vision 90B | 90B | 55 GB | 2x RTX 5090 |
| Wan2.1 14B (Xinference) | 14B | 40 GB | 2x RTX 5090 |
| HunyuanVideo (Xinference) | 13B | 48 GB | 2x RTX 5090 |