Loading
Taking longer than expected.
Reload page

What fits. What we can prove.

A model name and a GPU memory total are not a deployment test. We separate the source facts, configuration choices and measurements so you can see what supports a recommendation.

Updated 2026-09-30 · B3IQ engineering

Four kinds of evidence.

01

Upstream documented

The publisher's checkpoint identity, license metadata and intended inputs or outputs. Each guide links the reviewed model-card revision. This proves what the source says, not that B3IQ has deployed it.

02

Catalog configured

The exact runtime profile, artifact identifier, configured context, concurrency and memory threshold. These values come from the shared model catalog and its hardware companion. A tag can change; an unpinned artifact is labeled.

03

Estimated

Arithmetic for early planning, with its formula and exclusions. Estimates cannot establish installation success, full-context fit, answer quality or a latency target.

04

Measured by B3IQ

A reproducible run on identified hardware and software, with workload, results and date. The current model guides do not publish B3IQ benchmark results. A catalog flag or an upstream score cannot substitute for a run record.

Memory has several parts.

Weights are only the beginning. Serving also needs activations, runtime buffers and, for autoregressive text models, a KV cache that grows with sequence length and concurrent requests.

The same checkpoint can need different memory under different quantization, runtime and workload settings. GPU memory split across cards is not automatically one usable pool.

Weight-only arithmetic

billions of parameters × bits ÷ 8 = decimal GB

At 16 bits, this is two bytes per parameter. It excludes quantization metadata, non-quantized layers, buffers and KV cache. Catalog parameter counts may be rounded.

The existing planning helper adds 20% and rounds to a whole GB. That reserve is a heuristic, not a context- or concurrency-specific measurement. Small-model rounding can make it unsuitable; use the weight arithmetic and measure the full process.

A machine candidate is a starting point.

We match a profile's positive catalog memory threshold against visible store configurations through the same hardware and pricing engine used by the configurator. A machine may be sold with more GPUs than the model's minimum memory calculation requires.

The detail guide links the resulting GPU count and configuration. Current pricing and availability stay on the machine page. A missing threshold, unknown store roster or unsupported fit produces no automatic recommendation.

Before purchase, confirm runtime support, precision, GPU topology, CPU/RAM, storage and the workload's performance target. No guide offers a deployment action without a verified customer path.

What a published benchmark needs.

A result must be tied to a configuration and a workload. We require the following before calling it a B3IQ measurement:

  • Model revision and artifact digest; runtime, driver and relevant kernel versions.
  • Exact GPU edition and count, interconnect, CPU, system RAM and precision.
  • Input/output lengths or media dimensions/duration, batch size, concurrency, reasoning settings and warm-up.
  • Repeated runs with peak memory, errors and out-of-memory failures retained.
  • Time to first token, inter-token latency and aggregate throughput for text; task-appropriate quality and latency for embeddings, vision or audio.
  • Run date, reproduction steps, reviewer and explicit limits on what the result applies to.

Community reports can help choose a test, but their throughput is not presented as a B3IQ result. A historical result also does not establish that capacity is available now.

How the guides stay current.

Model identity and runtime settings remain in the shared catalog. Store configuration and price remain in the store. The public publishing record holds the reviewed source, original explanation and route; it does not maintain another specification table.

Catalog discovery can propose a model. Publication needs a reviewed identity, source, access terms and useful guidance. Runtime changes, artifact changes and license changes trigger review; an old test stays attached to the version tested.

We publish only allowlisted profiles. Tenant model names, private repositories, prompts, node addresses and customer data do not feed the public pages.

Technical references.