Loading
Taking longer than expected.
Reload page

Start with the workload. Then choose the machine.

Compare model profiles and memory requirements before you commit to hardware. Each guide separates what the catalog says from what still needs testing.

Sources reviewed 2026-09-30

Find hardware by memory · How we assess fit

Llama 3.1 8B

Chat / reasoning · Llama 3.1 Community License · 1 runtime profile(s)

A text assistant checkpoint for chat, summarization and retrieval-backed answers. Start here when you want a smaller Llama deployment to evaluate with your own prompts.

Read requirements for Llama 3.1 8B

Llama 3.3 70B

Chat / reasoning · Llama 3.3 Community License · 1 runtime profile(s)

Meta's instruction-tuned text model for multilingual dialogue. Evaluate it when a smaller assistant misses the quality bar and you can budget more GPU memory.

Read requirements for Llama 3.3 70B

Qwen3 8B

Chat / reasoning · Apache-2.0 · 2 runtime profile(s)

A compact Qwen text model with thinking and non-thinking modes upstream. It is a useful starting point for comparing reasoning behavior with a modest weight footprint.

Read requirements for Qwen3 8B

Qwen3 14B

Chat / reasoning · Apache-2.0 · 3 runtime profile(s)

A Qwen text model between the smaller and larger dense checkpoints. Compare it on your own assistant or coding prompts before committing to a larger machine.

Read requirements for Qwen3 14B

Qwen3 32B

Chat / reasoning · Apache-2.0 · 1 runtime profile(s)

A larger dense Qwen model for text generation and reasoning. It gives you another quality point to test before moving to a much larger checkpoint.

Read requirements for Qwen3 32B

Qwen3 Coder 30B

Chat / reasoning · Apache-2.0 · 1 runtime profile(s)

A mixture-of-experts checkpoint built for coding and tool use. Evaluate it against repository tasks with checkable outcomes, such as passing tests or a correct patch.

Read requirements for Qwen3 Coder 30B

gpt-oss 20B

Chat / reasoning · Apache-2.0 · 1 runtime profile(s)

OpenAI's smaller open-weight reasoning model, with tool use and structured output support documented upstream. Consider it for a private reasoning or agent evaluation.

Read requirements for gpt-oss 20B

gpt-oss 120B

Chat / reasoning · Apache-2.0 · 1 runtime profile(s)

The larger gpt-oss reasoning checkpoint. Evaluate it when task quality warrants a larger memory budget, especially for tool-using workflows with verifiable results.

Read requirements for gpt-oss 120B

Gemma 3 27B

Chat / reasoning · Gemma Terms of Use · 1 runtime profile(s)

Google's instruction-tuned Gemma checkpoint supports text and image understanding upstream. The current catalog profile is text chat; image serving needs separate validation.

Read requirements for Gemma 3 27B

Qwen2.5-VL 7B

Chat / reasoning · Apache-2.0 · 1 runtime profile(s)

A vision-language model for questions about images and visual documents. Use your real image types when assessing extraction accuracy and grounding.

Read requirements for Qwen2.5-VL 7B

BGE-M3

Embeddings · MIT · 1 runtime profile(s)

An embedding model for retrieval. It turns documents and queries into searchable representations; pair it with a text model when building retrieval-backed answers.

Read requirements for BGE-M3

Whisper large-v3

Speech to text · Apache-2.0 · 2 runtime profile(s)

A speech recognition checkpoint for transcription and speech translation. It serves an audio workflow rather than a text chat endpoint.

Read requirements for Whisper large-v3

A guide is the start of an evaluation.

These pages cover reviewed upstream models and catalog profiles. They do not promise live availability or a tested deployment. Use the guide to scope your workload, then confirm a configuration with us.