Loading
Taking longer than expected.
Reload page

Model library / Speech to text

Whisper large-v3: GPU requirements and hosting

A speech recognition checkpoint for transcription and speech translation. It serves an audio workflow rather than a text chat endpoint.

Sources reviewed 2026-09-30 · B3IQ engineering

Checkpoint overview.

Model creator

OpenAI
Task in catalog
Speech to text
Approx. total parameters (catalog)
Not specified
Runtime families
xinference · localai
Model license
Apache-2.0 ↗
Upstream access at review
No repository access gate reported
Canonical checkpoint
openai/whisper-large-v3 ↗

The creator identifies the original checkpoint. Runtime profiles may use separately packaged artifacts; their publishers and configuration can differ. Check the selected profile below.

Runtime profiles and requirements

The Xinference profile explicitly names large-v3. LocalAI uses a generic whisper-large artifact name, so confirm its resolved checkpoint before treating the profiles as equivalent.

xinference · audio-whisper-v3-xinference

Catalog GPU threshold
6 GB
Configured context
4,096 tokens
Catalog system RAM
8 GB
Runtime
xinference
Artifact identifier
whisper-large-v3
Precision
Not pinned in catalog
Configured concurrency
1 request(s)
Profile inputs / outputs
audio → text
Deployed artifact revision
Not verified

Catalog thresholds are configuration guidance, not measured peak memory. A positive GPU requirement is not recorded for every profile; unspecified is not zero. Context, concurrency and runtime overhead can increase memory use. Configured context is a profile setting, not the upstream maximum. Concurrency is configuration, not a load-test result.

localai · audio-whisper-large-localai

Catalog GPU threshold
Not specified
Configured context
4,096 tokens
Catalog system RAM
8 GB
Runtime
localai
Artifact identifier
whisper-large
Precision
Not pinned in catalog
Configured concurrency
1 request(s)
Profile inputs / outputs
audio → text
Deployed artifact revision
Not verified

Catalog thresholds are configuration guidance, not measured peak memory. A positive GPU requirement is not recorded for every profile; unspecified is not zero. Context, concurrency and runtime overhead can increase memory use. Configured context is a profile setting, not the upstream maximum. Concurrency is configuration, not a load-test result.

Weight-only arithmetic and a planning estimate

The shared catalog does not record a parameter count for this profile. No weight-memory estimate is published.

This arithmetic does not describe the selected quantized artifact. It excludes a workload-specific KV-cache calculation and cannot guarantee fit at the configured context or concurrency. A mixture-of-experts model still stores its full weights.

See the assumptions →

Plan the machine.

Memory-based machine candidates depend on the current visible store configurations. Confirm runtime compatibility, GPU count, interconnect and workload before purchase. Pricing and availability are shown on the machine page.

Explore machines

Before deployment

Test the workload you need.

Measure word error rate and processing time on recordings with your accents, noise and language mix. Report audio duration and batching with any speed result.

These are catalog profiles and planning figures. No B3IQ performance measurement or live capacity is claimed. Confirm the artifact, runtime, workload and machine before deployment.

Evidence available

Upstream identity / license
Source reviewed
Runtime settings
Catalog configuration
Memory fit
Guidance; not a measured peak
B3IQ performance
No published benchmark
Installation / live capacity
Confirm for your machine
Read our evidence standard →

Source and access.

Upstream checkpoint
openai/whisper-large-v3 ↗
Upstream license metadata
Apache-2.0 ↗
Access
No access gate reported by the upstream repository at review.

The reviewed revision identifies the source used for this guide. It is not a claim that this revision is installed on a B3IQ machine. Review the publisher's current license and acceptable-use terms for your application.

Guide updates and corrections.

These are changes to the guide. Source-review dates and workload-test evidence are recorded separately.

· Update

Added creator attribution and a canonical checkpoint overview. The guide now explains how the original creator differs from a runtime artifact publisher.

To report a correction, send the affected claim and a source URL to hello@b3iq.org.

Continue your evaluation.

Compare another checkpoint or explore a complementary task. Each guide explains the workload it is intended for.