Part of a larger toolkit
Build your Self-Hosted AI Stack
ScribeCrate is also available as part of Self-Hosted AI Stack. Deploy local models, chat, document processing, voice, and AI tools together with a single command.
Explore Self-Hosted AI StackPrivate, local speech-to-text
Open-source, self-hosted transcription API.
Transcribe recordings, generate subtitles, and translate speech into English on your own hardware. An OpenAI-compatible API, powered by faster-whisper. Deploy with Docker on CPU or an NVIDIA GPU.
Open source · MIT licensed
The Self-Hosted AI Builder's Guide Practical guidance for deployment, security, backups, and everyday operation.
Previously docker-whisper · Maintained by Lin Song (hwdsl2)
Recordings → useful text
Let's review the deployment plan.
We'll start on CPU, then add a GPU when we need it.
A focused API for your audio
Use ScribeCrate for meeting recordings, interviews, podcasts, subtitles, and applications that support the OpenAI Whisper API. Audio is processed on your server.
Transcription and English translation through
/v1/audio/transcriptions and /v1/audio/translations.
Integrate with clients that support the Whisper API.
Audio stays on your server for transcription. Run offline with models
cached in advance and WHISPER_LOCAL_ONLY enabled.
Identify who is speaking in each segment with the local sherpa-onnx extension. Useful for interviews and meeting transcripts.
Choose JSON, text, verbose JSON, SRT, or WebVTT. Receive decoded segments
from uploaded files over SSE with stream=true.
Run on amd64 or arm64 CPUs, or use the
:cuda image for NVIDIA acceleration on amd64.
Persistent caching keeps downloaded models across restarts.
Choose Whisper models from tiny through large-v3 and large-v3-turbo. Accept MP3, M4A, WAV, WebM, OGG, FLAC, and other formats supported by FFmpeg.
Images are automatically built and published through public GitHub Actions workflows. Powered by faster-whisper using Whisper models.
Start with one recording
This CPU example binds the API to your own computer. It uses the established
hwdsl2/whisper-server image and a persistent Docker volume.
Need NVIDIA acceleration, another model, diarization, or remote access? Follow the full deployment instructions. Use HTTPS and authentication when exposing the API to other computers.
docker run \
--name whisper \
--restart=always \
-v whisper-data:/var/lib/whisper \
-p 127.0.0.1:9000:9000 \
-d hwdsl2/whisper-server
The base model downloads on first start. Run docker logs whisper
and wait for “ScribeCrate transcription server is ready”.
scribecrate_api_key="$(docker exec whisper whisper_manage --getkey)"
A fresh install with this volume generates and persists an API key.
Replace audio.mp3 with your file, then run in the same shell:
curl http://127.0.0.1:9000/v1/audio/transcriptions \
-H "Authorization: Bearer $scribecrate_api_key" \
-F file=@audio.mp3 \
-F model=whisper-1
Example JSON response:
{"text": "Your transcribed text appears here."}
Part of a larger toolkit
ScribeCrate is also available as part of Self-Hosted AI Stack. Deploy local models, chat, document processing, voice, and AI tools together with a single command.
Explore Self-Hosted AI StackNeed live microphone transcription?
ScribeCrate focuses on uploaded audio files. The separately deployed WhisperLive project handles live microphone and audio streams over WebSocket.
Explore the WhisperLive serverBefore you deploy
The ScribeCrate project is open source under the MIT License. You provide the hardware or cloud server and pay any associated costs. Models and bundled dependencies retain their own licenses.
No. This website provides information and setup instructions. Deploy ScribeCrate on your own hardware to process audio. An optional Colab demo runs in Google's environment and is subject to its data practices.
ScribeCrate provides transcription and translation endpoints for clients that support the OpenAI Whisper API interface. It runs locally using faster-whisper; an OpenAI account or key is not needed. It does not implement every option of every OpenAI audio model. See the API documentation.
ScribeCrate streams decoded segments from an uploaded audio file over SSE. For live microphone or WebSocket audio, use the separate WhisperLive server. Speaker diarization requires full audio analysis and is not applied in SSE streaming mode.
Yes, with models cached in advance and WHISPER_LOCAL_ONLY=true.
See offline configuration.
The image also has optional aggregate usage counts; disable those separately
with WHISPER_DISABLE_USAGE_COUNTS=1. See the
usage-count documentation.
Bring your audio. Keep control of where it runs.
View on GitHub