Start with the workload. Then choose the machine.
Compare model profiles and memory requirements before you commit to hardware. Each guide separates what the catalog says from what still needs testing.
Sources reviewed 2026-09-30
Find hardware by memory · How we assess fit
Llama 3.1 8B
Chat / reasoning · Llama 3.1 Community License · 1 runtime profile(s)
A text assistant checkpoint for chat, summarization and retrieval-backed answers. Start here when you want a smaller Llama deployment to evaluate with your own prompts.
Read requirements for Llama 3.1 8BLlama 3.3 70B
Chat / reasoning · Llama 3.3 Community License · 1 runtime profile(s)
Meta's instruction-tuned text model for multilingual dialogue. Evaluate it when a smaller assistant misses the quality bar and you can budget more GPU memory.
Read requirements for Llama 3.3 70BQwen3 8B
Chat / reasoning · Apache-2.0 · 2 runtime profile(s)
A compact Qwen text model with thinking and non-thinking modes upstream. It is a useful starting point for comparing reasoning behavior with a modest weight footprint.
Read requirements for Qwen3 8BQwen3 14B
Chat / reasoning · Apache-2.0 · 3 runtime profile(s)
A Qwen text model between the smaller and larger dense checkpoints. Compare it on your own assistant or coding prompts before committing to a larger machine.
Read requirements for Qwen3 14BQwen3 32B
Chat / reasoning · Apache-2.0 · 1 runtime profile(s)
A larger dense Qwen model for text generation and reasoning. It gives you another quality point to test before moving to a much larger checkpoint.
Read requirements for Qwen3 32BQwen3 Coder 30B
Chat / reasoning · Apache-2.0 · 1 runtime profile(s)
A mixture-of-experts checkpoint built for coding and tool use. Evaluate it against repository tasks with checkable outcomes, such as passing tests or a correct patch.
Read requirements for Qwen3 Coder 30Bgpt-oss 20B
Chat / reasoning · Apache-2.0 · 1 runtime profile(s)
OpenAI's smaller open-weight reasoning model, with tool use and structured output support documented upstream. Consider it for a private reasoning or agent evaluation.
Read requirements for gpt-oss 20Bgpt-oss 120B
Chat / reasoning · Apache-2.0 · 1 runtime profile(s)
The larger gpt-oss reasoning checkpoint. Evaluate it when task quality warrants a larger memory budget, especially for tool-using workflows with verifiable results.
Read requirements for gpt-oss 120BGemma 3 27B
Chat / reasoning · Gemma Terms of Use · 1 runtime profile(s)
Google's instruction-tuned Gemma checkpoint supports text and image understanding upstream. The current catalog profile is text chat; image serving needs separate validation.
Read requirements for Gemma 3 27BQwen2.5-VL 7B
Chat / reasoning · Apache-2.0 · 1 runtime profile(s)
A vision-language model for questions about images and visual documents. Use your real image types when assessing extraction accuracy and grounding.
Read requirements for Qwen2.5-VL 7BBGE-M3
Embeddings · MIT · 1 runtime profile(s)
An embedding model for retrieval. It turns documents and queries into searchable representations; pair it with a text model when building retrieval-backed answers.
Read requirements for BGE-M3Whisper large-v3
Speech to text · Apache-2.0 · 2 runtime profile(s)
A speech recognition checkpoint for transcription and speech translation. It serves an audio workflow rather than a text chat endpoint.
Read requirements for Whisper large-v3A guide is the start of an evaluation.
These pages cover reviewed upstream models and catalog profiles. They do not promise live availability or a tested deployment. Use the guide to scope your workload, then confirm a configuration with us.