All Templates

whisper-turbo.c

LLM

CPU-only Whisper large-v3-turbo speech-to-text behind an OpenAI-compatible API

Deploy Now

README

CPU-only Whisper large-v3-turbo speech-to-text behind an OpenAI-compatible API.

Overview

whisper-turbo.c is a from-scratch C implementation of OpenAI's Whisper large-v3-turbo. There is no Python, no PyTorch and no whisper.cpp underneath it: the encoder, decoder, mel frontend and INT8 kernels are all C, with AVX2/AVX-512/VNNI paths selected at runtime. It runs on CPU and needs no GPU.

It serves a documented subset of the OpenAI transcription API, so an existing OpenAI client can be pointed at it by changing the base URL. It is an API, not an app: the server has exactly two routes, GET /health and POST /v1/audio/transcriptions, and ships no web UI of any kind. Nothing in this template adds one.

Upstream publishes no image, no release and no git tag, so ./Dockerfile builds the server from a pinned commit. It also converts the Whisper checkpoint to upstream's packed INT8 format at build time and bakes the result into the image, which is the one real departure from upstream's own Dockerfile.cloud. See Why the model is in the image.

What you get by hosting it

  • Transcription with automatic language detection, on your own machine, over HTTPS.
  • Audio that never leaves your deployment. whisper-1 and the other model aliases the API accepts all run against the local model; nothing is forwarded to OpenAI.
  • A drop-in base URL for OpenAI transcription clients, authenticated with a bearer token you choose.
  • No cold model download. The converted model is inside the image, so a restart or a redeployment comes back with no network fetch and no warm-up step.
  • Optional timestamped speaker labels, once you supply a Hugging Face token for upstream's gated diarization weights. See Speaker labels.

What you need before deploying

  • A bearer token of your own choosing for WHISPER_API_KEY. Keep a copy before you submit the form: template variables are stored write-only and this one cannot be read back afterwards.
  • Audio as mono 16-bit PCM WAV at 16 kHz, up to 300 seconds per request. The server does no format conversion and rejects anything else, so an MP3 or a 44.1 kHz stereo WAV has to be converted first, for example with ffmpeg -i input.mp3 -ar 16000 -ac 1 -c:a pcm_s16le output.wav.
  • An amd64 box. The server compiles against x86 kernels and there is no arm64 image.
  • Only for speaker labels: a Hugging Face account that has accepted the conditions on pyannote/speaker-diarization-community-1, and an access token from it.

Configuration

VariableRequiredWhat it does
WHISPER_API_KEYyesThe bearer token every request must carry. You choose the value; there is no default and nothing is generated for you, because you need to be able to send it back. Upstream refuses to start on a non-loopback bind without it
OMP_NUM_THREADSfixed, 8Inference threads. Matches the 8 shared vCPUs a compute machine is configured with; the server clamps it to 1 to 8
WHISPER_ACTIVATIONSfixed, int8INT8 encoder activations. See Why INT8 is on by default; clearing it makes transcription time out
WHISPER_DECODER_ACTIVATIONSfixed, int8INT8 decoder projections and vocabulary head. Same reason, same section
HF_TOKENnoA Hugging Face access token, for speaker labels only. See Speaker labels. Blank means transcription only
WHISPER_DIAR_SIMDnoavx512 selects the FP32 AVX-512 diarization kernels where the CPU has them, with fallback. No effect without HF_TOKEN
WHISPER_DIAR_BATCH_INPUTno1 batches the diarizer's LSTM input projections. No effect without HF_TOKEN
WHISPER_SIMDnoPin the kernel path to scalar, avx2 or avx512. Detected automatically when blank
WHISPER_REQUEST_TIMEOUTnoSeconds allowed for one transcription before the server gives up. Default 300, maximum 3600. The edge cuts the connection at 60 seconds regardless

A persistent volume is mounted at /data, and the only thing that ever lands on it is the diarization checkpoints below. Transcription itself keeps no state.

After deploy

Check the service is up. /health is the one route that answers before the token check:

curl https://<your-service-url>/health
# {"status":"ready"}

Then transcribe. Everything else requires the bearer token:

curl https://<your-service-url>/v1/audio/transcriptions \
  -H "Authorization: Bearer $WHISPER_API_KEY" \
  -F file=@speech.wav -F model=whisper-1
# {"text":"..."}

Add -F response_format=text for a bare transcript instead of JSON, or -F language=en to skip language detection. The server holds one transcription at a time and answers a second concurrent request with 429; queue on the client side if you need throughput.

Expect the first request after a start to be slower than later ones. The model is mmapped, so the server is listening immediately but the 808 MiB of weights page in from disk during that first transcription. Measured with the 11 second JFK sample: 15 seconds on the first request after a restart, 9.5 seconds warm.

There are no CORS headers. A cross-origin preflight is answered 401 with no Access-Control-Allow-Origin, so browser JavaScript cannot call this API directly. Call it from your own backend, or put a proxy in front of it.

Speaker labels

Upstream ships a C port of the pyannote Community-1 diarization pipeline, which labels who spoke when. Its four checkpoints are gated on Hugging Face (CC BY 4.0, with access conditions accepted per account), so they are not in this image and CI cannot fetch them either.

Transcription is verified end to end on this platform. This path is not, because the gating is per account and nobody has deployed the template with a token that has accepted the conditions. What follows is what upstream documents and what this image wires up.

To turn it on: accept the conditions on pyannote/speaker-diarization-community-1, create an access token, and set it as HF_TOKEN. On the next start the entrypoint downloads the four files (33 MB), checks each one's SHA-256 against upstream's published values, and leaves them on the volume, so later restarts find them already there.

curl https://<your-service-url>/v1/audio/transcriptions \
  -H "Authorization: Bearer $WHISPER_API_KEY" \
  -F file=@conversation.wav -F model=gpt-4o-transcribe-diarize \
  -F response_format=diarized_json -F chunking_strategy=auto

The response carries text, timestamped segments each with a speaker label (A, B, ...), and duration usage. Labels distinguish voices, not people, and overlapping speech is not separated.

A token that does not work is not fatal. The failure is logged, the service still starts, and diarized requests keep answering 503 with Configure WHISPER_DIARIZATION_MODELS to enable diarization. while transcription carries on unaffected. That is deliberate: making a bad token fatal would produce a restart loop that never reaches the health gate and reads as a broken template. The same 60 second edge limit applies, and diarization is slower than transcription alone, so keep diarized clips short.

Why INT8 is on by default

WHISPER_ACTIVATIONS and WHISPER_DECODER_ACTIVATIONS are int8 in the manifest, and that is not a tuning preference. The edge in front of a compute service cuts a request off at 60 seconds, and upstream's documented-safe FP32-activation default does not finish an 11 second clip inside that. Measured on this platform, same machine, same audio:

ActivationsResult
FP32 (upstream default)more than 118 seconds, never returned, edge 502 at 60 seconds
INT8 (what this ships)15.2 seconds cold, 9.5 seconds warm, 200 with the correct transcript

Upstream marks INT8 activations experimental and documents accuracy limitations, and its own published benchmarks are measured with both of them enabled. The machines this lands on carry AVX-512 VNNI, which is the kernel path that produces the difference. If you would rather have FP32 and drive the API from a client that tolerates a long request, clear both variables on the service after deploy; be aware that the public URL will then time out.

Why the model is in the image

Upstream's Dockerfile.cloud downloads the 1.6 GB GGML checkpoint and converts it on first boot into a mounted volume, deliberately, to keep the image small and the build fast. That shape does not survive this platform's deploy gate: the entrypoint does all of that work before it execs the server, so nothing is listening on the port for minutes, and the gate sends a real HTTP request and fails long before the port opens. There is no manifest field that extends it.

So ./Dockerfile does the download, the SHA-256 check and the conversion in the build stage and copies the 808 MiB result into the final image. The trade is a large image and a slow CI build, against a deploy that passes its health check on the first try and is byte-identical after every restart. The volume this template does mount carries only the optional diarization checkpoints; the Whisper model never touches it.

Links

Services & Specs

whisper
Web service
Image
ghcr.io/insforge/insta-oss/templates/whisper-turbo:1.1.1
Port
8080
Healthcheck
/health

Variables

You supply 1 variable before the first deploy.

Required

WHISPER_API_KEY

Bearer token clients send as `Authorization: Bearer <key>` on every transcription request. Pick your own strong random string and keep a copy: it cannot be read back after deploy, and the server will not start without it

Optional (5)
HF_TOKEN

Hugging Face access token, needed only for speaker labels. Accept the conditions on pyannote/speaker-diarization-community-1 with the same account first. Leave blank for transcription only, in which case diarized requests answer 503

WHISPER_DIAR_SIMD

Diarization kernel path. `avx512` selects the FP32 AVX-512 diarization kernels where the CPU has them, with fallback. No effect without HF_TOKEN

WHISPER_DIAR_BATCH_INPUT

Set to `1` to batch the diarizer's LSTM input projections. No effect without HF_TOKEN

WHISPER_SIMD

Force a CPU kernel path: `scalar`, `avx2` or `avx512`. Detected automatically when blank; set it only to work around a placement whose CPU features differ from what you benchmarked

WHISPER_REQUEST_TIMEOUT

Seconds the server will spend on one transcription before giving up. Default 300, maximum 3600. The edge in front of the service cuts the connection at 60 seconds regardless, so raising this only helps a client that can tolerate that