51 model entries. Compare GPU memory figures and their sources before choosing a machine configuration.
These are catalog sizing thresholds, planning estimates and external reports. Context, concurrency and runtime overhead can increase memory use. A memory match does not establish a tested deployment, throughput or live capacity.
| Model | Params | Serve VRAM sizing | Fine-tune VRAM sizing |
|---|---|---|---|
| Kimi K2.6 MoE | 1000B | 600 GB (4-bit) | not listed |
| DeepSeek V4 Pro | 671B | 400 GB (4-bit) | not listed |
| MiniMax M2 | 230B | 138 GB (4-bit); 140 GB (Runtime default) | 200 GB (QLoRA) |
| gpt-oss 120B | 120B | 72 GB (4-bit); 80 GB (Runtime default) | 90 GB (QLoRA) |
| Llama 3.2 Vision 90B | 90B | 55 GB (4-bit); 64 GB (Runtime default); 216 GB (16-bit) | 68 GB (QLoRA) |
| Qwen2.5-VL 72B | 72B | 45 GB (4-bit); 48 GB (Runtime default); 173 GB (16-bit) | 54 GB (QLoRA) |
| Llama 3.3 70B | 70B | 42 GB (4-bit); 80 GB (Runtime default); 84 GB (8-bit); 168 GB (16-bit) | 53 GB (QLoRA) |
| LLaVA 34B | 34B | 24 GB (Runtime default) | 26 GB (QLoRA) |
| Qwen3 32B | 32B | 22 GB (4-bit); 32 GB (Runtime default); 77 GB (16-bit) | 24 GB (QLoRA) |
| Qwen Coder 32B | 32B | 32 GB (Runtime default) | 24 GB (QLoRA) |
| DeepSeek-R1 32B | 32B | 32 GB (Runtime default) | 24 GB (QLoRA) |
| Qwen2.5 32B Instruct | 32B | 32 GB (Runtime default) | 24 GB (QLoRA) |
| Qwen2.5-VL 32B | 32B | 24 GB (Runtime default) | 24 GB (QLoRA) |
| MOVA 720p + audio (SGLang) | 32B | not listed | not listed |
| Qwen3 30B A3B | 30B | 20 GB (4-bit); 32 GB (Runtime default) | 23 GB (QLoRA) |
| Qwen3 Coder 30B | 30B | 20 GB (4-bit); 32 GB (Runtime default) | 23 GB (QLoRA) |
| Gemma 3 27B | 27B | 18 GB (4-bit); 24 GB (Runtime default); 65 GB (16-bit) | 21 GB (QLoRA) |
| Qwen3 Coder 27B | 27B | 32 GB (Runtime default) | 20 GB (QLoRA) |
| Mistral Small 24B | 24B | 16 GB (4-bit); 24 GB (Runtime default); 29 GB (8-bit) | 18 GB (QLoRA) |
| Devstral Small 24B | 24B | 16 GB (Runtime default) | 18 GB (QLoRA) |
| Mistral Small 3.1 24B (Vision) | 24B | 18 GB (Runtime default) | 18 GB (QLoRA) |
| gpt-oss 20B | 20B | 16 GB (4-bit); 24 GB (Runtime default) | 15 GB (QLoRA) |
| Qwen3 14B | 14B | 16 GB (Runtime default); 24 GB (Runtime default) | 11 GB (QLoRA) |
| Qwen2.5 Coder 14B | 14B | 16 GB (Runtime default) | 11 GB (QLoRA) |
| Phi-4 14B | 14B | 16 GB (Runtime default) | 11 GB (QLoRA) |
| DeepSeek-R1 14B | 14B | 16 GB (Runtime default) | 11 GB (QLoRA) |
| Qwen2.5 14B Instruct | 14B | 16 GB (Runtime default) | 11 GB (QLoRA) |
| Wan2.1 14B | 14B | 40 GB (Runtime default) | not listed |
| LLaVA 13B | 13B | 10 GB (Runtime default) | 10 GB (QLoRA) |
| HunyuanVideo | 13B | 48 GB (Runtime default) | not listed |
| Gemma 3 12B | 12B | 16 GB (Runtime default) | 9 GB (QLoRA) |
| Llama 3.2 Vision 11B | 11B | 8 GB (Runtime default) | 8 GB (QLoRA) |
| Qwen3 8B | 8B | not listed | 6 GB (QLoRA) |
| Granite 3.3 8B | 8B | not listed | 6 GB (QLoRA) |
| Llama 3.1 8B | 8B | 8 GB (Runtime default) | 6 GB (QLoRA) |
| DeepSeek-R1 8B | 8B | 8 GB (Runtime default) | 6 GB (QLoRA) |
| Qwen3 Coder 8B | 8B | not listed | 6 GB (QLoRA) |
| Qwen3 Embedding 8B | 8B | 18 GB (Runtime default) | not listed |
| MiniCPM-V 2.6 8B | 8B | 7 GB (Runtime default) | 6 GB (QLoRA) |
| LLaVA-Llama3 8B | 8B | 7 GB (Runtime default) | 6 GB (QLoRA) |
| Mistral 7B | 7B | not listed | 5 GB (QLoRA) |
| Qwen Coder 7B | 7B | 8 GB (Runtime default) | 5 GB (QLoRA) |
| DeepSeek-R1 7B | 7B | 8 GB (Runtime default) | 5 GB (QLoRA) |
| Qwen2.5 7B Instruct | 7B | 8 GB (Runtime default) | 5 GB (QLoRA) |
| LLaVA 7B (Vision) | 7B | not listed | 5 GB (QLoRA) |
| gte-Qwen2 7B | 7B | 16 GB (Runtime default) | not listed |
| Qwen2.5-VL 7B | 7B | 7 GB (Runtime default) | 5 GB (QLoRA) |
| BakLLaVA 7B | 7B | 6 GB (Runtime default) | 5 GB (QLoRA) |
| BGE-M3 | 0.567B | not listed | not listed |
| Whisper Large v3 STT (LocalAI) | — | not listed | not listed |
| Whisper Large v3 STT | — | 6 GB (Runtime default) | not listed |
kimi-k2-6-moe · Profile runtime: ollama · Serve: 600 GB (4-bit) · System RAM: 768 GB
Model-card-derived planning figure. The row does not record a source URL; not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
deepseek-v4-pro · Profile runtime: ollama · Serve: 400 GB (4-bit) · System RAM: 512 GB
Model-card-derived planning figure. The row does not record a source URL; not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vllm-minimax-m2 · Profile runtime: vllm · Serve: 140 GB (Runtime default) · System RAM: 192 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vllm-minimax-m2 · Profile runtime: vllm · Serve: 138 GB (4-bit (NVFP4)) · System RAM: 64 GB
External report. Not independently verified on B3IQ. The profile runtime may differ from the reported setup. Sizing note: VRAM includes 20% over the weights for context and runtime. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment. View source
vllm-minimax-m2 · Profile runtime: vllm · QLoRA: 200 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
Read gpt-oss 120B requirements and sources
gpt-oss-120b · Profile runtime: ollama · Serve: 80 GB (Runtime default) · System RAM: 128 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
gpt-oss-120b · Profile runtime: ollama · Serve: 72 GB (4-bit (MXFP4)) · System RAM: 64 GB
External report. Not independently verified on B3IQ. The profile runtime may differ from the reported setup. Sizing note: VRAM includes 20% over the weights for context and runtime. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment. View source
gpt-oss-120b · Profile runtime: ollama · QLoRA: 90 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
gpt-oss-120b · Profile runtime: ollama · LoRA: 360 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-llama32-90b · Profile runtime: ollama · Serve: 64 GB (Runtime default) · System RAM: 96 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-llama32-90b · Profile runtime: ollama · Serve: 55 GB (4-bit) · System RAM: 64 GB
Model-card-derived planning figure. The row does not record a source URL; not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-llama32-90b · Profile runtime: ollama · Serve: 216 GB (16-bit) · System RAM: 192 GB
Model-card-derived planning figure. The row does not record a source URL; not a B3IQ measurement. Sizing note: VRAM includes 20% over the weights for context and runtime. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-llama32-90b · Profile runtime: ollama · QLoRA: 68 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-llama32-90b · Profile runtime: ollama · LoRA: 270 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-qwen25vl-72b · Profile runtime: ollama · Serve: 48 GB (Runtime default) · System RAM: 64 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-qwen25vl-72b · Profile runtime: ollama · Serve: 45 GB (4-bit) · System RAM: 64 GB
Model-card-derived planning figure. The row does not record a source URL; not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-qwen25vl-72b · Profile runtime: ollama · Serve: 173 GB (16-bit) · System RAM: 192 GB
Model-card-derived planning figure. The row does not record a source URL; not a B3IQ measurement. Sizing note: VRAM includes 20% over the weights for context and runtime. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-qwen25vl-72b · Profile runtime: ollama · QLoRA: 54 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-qwen25vl-72b · Profile runtime: ollama · LoRA: 216 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
Read Llama 3.3 70B requirements and sources
max-70b · Profile runtime: ollama · Serve: 80 GB (Runtime default) · System RAM: 128 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
max-70b · Profile runtime: ollama · Serve: 42 GB (4-bit (Q4_K_M)) · System RAM: 64 GB
External report. Not independently verified on B3IQ. The profile runtime may differ from the reported setup. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment. View source
max-70b · Profile runtime: ollama · Serve: 84 GB (8-bit (Q6_K_L)) · System RAM: 96 GB
External report. Not independently verified on B3IQ. The profile runtime may differ from the reported setup. Sizing note: VRAM includes 20% over the weights for context and runtime. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment. View source
max-70b · Profile runtime: ollama · Serve: 168 GB (16-bit) · System RAM: 192 GB
Model-card-derived planning figure. The row does not record a source URL; not a B3IQ measurement. Sizing note: VRAM includes 20% over the weights for context and runtime. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
max-70b · Profile runtime: ollama · QLoRA: 53 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
max-70b · Profile runtime: ollama · LoRA: 210 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-llava-34b · Profile runtime: ollama · Serve: 24 GB (Runtime default) · System RAM: 48 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-llava-34b · Profile runtime: ollama · QLoRA: 26 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-llava-34b · Profile runtime: ollama · LoRA: 102 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
Read Qwen3 32B requirements and sources
quality-32b · Profile runtime: ollama · Serve: 32 GB (Runtime default) · System RAM: 64 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
quality-32b · Profile runtime: ollama · Serve: 22 GB (4-bit) · System RAM: 32 GB
Model-card-derived planning figure. The row does not record a source URL; not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
quality-32b · Profile runtime: ollama · Serve: 77 GB (16-bit) · System RAM: 96 GB
Model-card-derived planning figure. The row does not record a source URL; not a B3IQ measurement. Sizing note: VRAM includes 20% over the weights for context and runtime. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
quality-32b · Profile runtime: ollama · QLoRA: 24 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
quality-32b · Profile runtime: ollama · LoRA: 96 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
qwen-coder-32b · Profile runtime: ollama · Serve: 32 GB (Runtime default) · System RAM: 64 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
qwen-coder-32b · Profile runtime: ollama · QLoRA: 24 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
qwen-coder-32b · Profile runtime: ollama · LoRA: 96 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
deepseek-r1-32b · Profile runtime: ollama · Serve: 32 GB (Runtime default) · System RAM: 64 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
deepseek-r1-32b · Profile runtime: ollama · QLoRA: 24 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
deepseek-r1-32b · Profile runtime: ollama · LoRA: 96 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
qwen25-32b · Profile runtime: ollama · Serve: 32 GB (Runtime default) · System RAM: 64 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
qwen25-32b · Profile runtime: ollama · QLoRA: 24 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
qwen25-32b · Profile runtime: ollama · LoRA: 96 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-qwen25vl-32b · Profile runtime: ollama · Serve: 24 GB (Runtime default) · System RAM: 48 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-qwen25vl-32b · Profile runtime: ollama · QLoRA: 24 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-qwen25vl-32b · Profile runtime: ollama · LoRA: 96 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
No sourced sizing row is recorded for this model.
qwen3-30b · Profile runtime: ollama · Serve: 32 GB (Runtime default) · System RAM: 64 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
qwen3-30b · Profile runtime: ollama · Serve: 20 GB (4-bit (Q4_K_M)) · System RAM: 32 GB
External report. Not independently verified on B3IQ. The profile runtime may differ from the reported setup. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment. View source
qwen3-30b · Profile runtime: ollama · QLoRA: 23 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
qwen3-30b · Profile runtime: ollama · LoRA: 90 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
qwen3-30b-a3b · Profile runtime: ollama · Serve: 32 GB (Runtime default) · System RAM: 64 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
qwen3-30b-a3b · Profile runtime: ollama · QLoRA: 23 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
qwen3-30b-a3b · Profile runtime: ollama · LoRA: 90 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
Read Qwen3 Coder 30B requirements and sources
qwen3-coder-30b · Profile runtime: ollama · Serve: 32 GB (Runtime default) · System RAM: 64 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
qwen3-coder-30b · Profile runtime: ollama · Serve: 20 GB (4-bit (Q4_K_M)) · System RAM: 32 GB
External report. Not independently verified on B3IQ. The profile runtime may differ from the reported setup. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment. View source
qwen3-coder-30b · Profile runtime: ollama · QLoRA: 23 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
qwen3-coder-30b · Profile runtime: ollama · LoRA: 90 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
Read Gemma 3 27B requirements and sources
gemma3-27b · Profile runtime: ollama · Serve: 24 GB (Runtime default) · System RAM: 48 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
gemma3-27b · Profile runtime: ollama · Serve: 18 GB (4-bit (Q4_K_M)) · System RAM: 32 GB
External report. Not independently verified on B3IQ. The profile runtime may differ from the reported setup. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment. View source
gemma3-27b · Profile runtime: ollama · Serve: 65 GB (16-bit) · System RAM: 64 GB
Model-card-derived planning figure. The row does not record a source URL; not a B3IQ measurement. Sizing note: VRAM includes 20% over the weights for context and runtime. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
gemma3-27b · Profile runtime: ollama · QLoRA: 21 GB (QLoRA)
External report. Not independently verified on B3IQ. The profile runtime may differ from the reported setup. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment. View source
gemma3-27b · Profile runtime: ollama · LoRA: 81 GB (LoRA)
External report. Not independently verified on B3IQ. The profile runtime may differ from the reported setup. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment. View source
qwen3-coder-27b · Profile runtime: ollama · Serve: 32 GB (Runtime default) · System RAM: 48 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
qwen3-coder-27b · Profile runtime: ollama · QLoRA: 20 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
qwen3-coder-27b · Profile runtime: ollama · LoRA: 81 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
mistral-small-24b · Profile runtime: ollama · Serve: 24 GB (Runtime default) · System RAM: 32 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
mistral-small-24b · Profile runtime: ollama · Serve: 16 GB (4-bit) · System RAM: 32 GB
Model-card-derived planning figure. The row does not record a source URL; not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
mistral-small-24b · Profile runtime: ollama · Serve: 29 GB (8-bit) · System RAM: 48 GB
Model-card-derived planning figure. The row does not record a source URL; not a B3IQ measurement. Sizing note: VRAM includes 20% over the weights for context and runtime. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
mistral-small-24b · Profile runtime: ollama · QLoRA: 18 GB (QLoRA)
External report. Not independently verified on B3IQ. The profile runtime may differ from the reported setup. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment. View source
mistral-small-24b · Profile runtime: ollama · LoRA: 72 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
devstral-small-24b · Profile runtime: ollama · Serve: 16 GB (Runtime default) · System RAM: 32 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
devstral-small-24b · Profile runtime: ollama · QLoRA: 18 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
devstral-small-24b · Profile runtime: ollama · LoRA: 72 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-mistral-small31 · Profile runtime: ollama · Serve: 18 GB (Runtime default) · System RAM: 32 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-mistral-small31 · Profile runtime: ollama · QLoRA: 18 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-mistral-small31 · Profile runtime: ollama · LoRA: 72 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
Read gpt-oss 20B requirements and sources
gpt-oss-20b · Profile runtime: ollama · Serve: 24 GB (Runtime default) · System RAM: 32 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
gpt-oss-20b · Profile runtime: ollama · Serve: 16 GB (4-bit) · System RAM: 32 GB
External report. Not independently verified on B3IQ. The profile runtime may differ from the reported setup. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment. View source
gpt-oss-20b · Profile runtime: ollama · QLoRA: 15 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
gpt-oss-20b · Profile runtime: ollama · LoRA: 60 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
Read Qwen3 14B requirements and sources
llamacpp-qwen3-14b · Profile runtime: llama.cpp · QLoRA: 11 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
llamacpp-qwen3-14b · Profile runtime: llama.cpp · LoRA: 42 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
qwen3-14b · Profile runtime: ollama · Serve: 16 GB (Runtime default) · System RAM: 48 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
qwen3-14b · Profile runtime: ollama · QLoRA: 11 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
qwen3-14b · Profile runtime: ollama · LoRA: 42 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vllm-qwen3-14b · Profile runtime: vllm · Serve: 24 GB (Runtime default) · System RAM: 64 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vllm-qwen3-14b · Profile runtime: vllm · QLoRA: 11 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vllm-qwen3-14b · Profile runtime: vllm · LoRA: 42 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
code-14b · Profile runtime: ollama · Serve: 16 GB (Runtime default) · System RAM: 32 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
code-14b · Profile runtime: ollama · QLoRA: 11 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
code-14b · Profile runtime: ollama · LoRA: 42 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
phi4-14b · Profile runtime: ollama · Serve: 16 GB (Runtime default) · System RAM: 48 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
phi4-14b · Profile runtime: ollama · QLoRA: 11 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
phi4-14b · Profile runtime: ollama · LoRA: 42 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
deepseek-r1-14b · Profile runtime: ollama · Serve: 16 GB (Runtime default) · System RAM: 48 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
deepseek-r1-14b · Profile runtime: ollama · QLoRA: 11 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
deepseek-r1-14b · Profile runtime: ollama · LoRA: 42 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
qwen25-14b · Profile runtime: ollama · Serve: 16 GB (Runtime default) · System RAM: 48 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
qwen25-14b · Profile runtime: ollama · QLoRA: 11 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
qwen25-14b · Profile runtime: ollama · LoRA: 42 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
video-wan21-14b-xinference · Profile runtime: xinference · Serve: 40 GB (Runtime default) · System RAM: 48 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-llava-13b · Profile runtime: ollama · Serve: 10 GB (Runtime default) · System RAM: 16 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-llava-13b · Profile runtime: ollama · QLoRA: 10 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-llava-13b · Profile runtime: ollama · LoRA: 39 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
video-hunyuan-xinference · Profile runtime: xinference · Serve: 48 GB (Runtime default) · System RAM: 96 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
gemma3-12b · Profile runtime: ollama · Serve: 16 GB (Runtime default) · System RAM: 48 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
gemma3-12b · Profile runtime: ollama · QLoRA: 9 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
gemma3-12b · Profile runtime: ollama · LoRA: 36 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-llama32-11b · Profile runtime: ollama · Serve: 8 GB (Runtime default) · System RAM: 16 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-llama32-11b · Profile runtime: ollama · QLoRA: 8 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-llama32-11b · Profile runtime: ollama · LoRA: 33 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
Read Qwen3 8B requirements and sources
llamacpp-qwen3-8b · Profile runtime: llama.cpp · QLoRA: 6 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
llamacpp-qwen3-8b · Profile runtime: llama.cpp · LoRA: 24 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
qwen3-8b · Profile runtime: ollama · QLoRA: 6 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
qwen3-8b · Profile runtime: ollama · LoRA: 24 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
granite33-8b · Profile runtime: ollama · QLoRA: 6 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
granite33-8b · Profile runtime: ollama · LoRA: 24 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
Read Llama 3.1 8B requirements and sources
starter-8b · Profile runtime: ollama · Serve: 8 GB (Runtime default) · System RAM: 16 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
starter-8b · Profile runtime: ollama · QLoRA: 6 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
starter-8b · Profile runtime: ollama · LoRA: 24 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
deepseek-r1-8b · Profile runtime: ollama · Serve: 8 GB (Runtime default) · System RAM: 24 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
deepseek-r1-8b · Profile runtime: ollama · QLoRA: 6 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
deepseek-r1-8b · Profile runtime: ollama · LoRA: 24 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
qwen3-coder-8b · Profile runtime: ollama · QLoRA: 6 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
qwen3-coder-8b · Profile runtime: ollama · LoRA: 24 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
embed-qwen3-8b-xinference · Profile runtime: xinference · Serve: 18 GB (Runtime default) · System RAM: 32 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-minicpm-v · Profile runtime: ollama · Serve: 7 GB (Runtime default) · System RAM: 10 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-minicpm-v · Profile runtime: ollama · QLoRA: 6 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-minicpm-v · Profile runtime: ollama · LoRA: 24 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-llava-llama3 · Profile runtime: ollama · Serve: 7 GB (Runtime default) · System RAM: 10 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-llava-llama3 · Profile runtime: ollama · QLoRA: 6 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-llava-llama3 · Profile runtime: ollama · LoRA: 24 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
mistral-7b · Profile runtime: ollama · QLoRA: 5 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
mistral-7b · Profile runtime: ollama · LoRA: 21 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
qwen-coder-7b · Profile runtime: ollama · Serve: 8 GB (Runtime default) · System RAM: 24 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
qwen-coder-7b · Profile runtime: ollama · QLoRA: 5 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
qwen-coder-7b · Profile runtime: ollama · LoRA: 21 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
deepseek-r1-7b · Profile runtime: ollama · Serve: 8 GB (Runtime default) · System RAM: 24 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
deepseek-r1-7b · Profile runtime: ollama · QLoRA: 5 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
deepseek-r1-7b · Profile runtime: ollama · LoRA: 21 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
qwen25-7b · Profile runtime: ollama · Serve: 8 GB (Runtime default) · System RAM: 24 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
qwen25-7b · Profile runtime: ollama · QLoRA: 5 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
qwen25-7b · Profile runtime: ollama · LoRA: 21 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-llava-7b · Profile runtime: ollama · QLoRA: 5 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-llava-7b · Profile runtime: ollama · LoRA: 21 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
embed-gte-qwen2-xinference · Profile runtime: xinference · Serve: 16 GB (Runtime default) · System RAM: 32 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
Read Qwen2.5-VL 7B requirements and sources
vision-qwen25vl-7b · Profile runtime: ollama · Serve: 7 GB (Runtime default) · System RAM: 10 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-qwen25vl-7b · Profile runtime: ollama · QLoRA: 5 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-qwen25vl-7b · Profile runtime: ollama · LoRA: 21 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-bakllava · Profile runtime: ollama · Serve: 6 GB (Runtime default) · System RAM: 8 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-bakllava · Profile runtime: ollama · QLoRA: 5 GB (QLoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
vision-bakllava · Profile runtime: ollama · LoRA: 21 GB (LoRA)
Planning estimate. Not a B3IQ measurement. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.
Read BGE-M3 requirements and sources
No sourced sizing row is recorded for this model.
Read Whisper large-v3 requirements and sources
No sourced sizing row is recorded for this model.
Read Whisper large-v3 requirements and sources
audio-whisper-v3-xinference · Profile runtime: xinference · Serve: 6 GB (Runtime default) · System RAM: 8 GB
Catalog threshold. Configuration guidance, not measured peak memory. No row-specific sizing note is recorded. Test date, measured context length, concurrency, runtime version and test hardware are not recorded for this row. Confirm these before deployment.