ScribeCrate

Private, local speech-to-text

ScribeCrate

Open-source, self-hosted transcription API.

Transcribe recordings, generate subtitles, and translate speech into English on your own hardware. An OpenAI-compatible API, powered by faster-whisper. Deploy with Docker on CPU or an NVIDIA GPU.

View on GitHub

Open source · MIT licensed

Get the book on Amazon

The Self-Hosted AI Builder's Guide Practical guidance for deployment, security, backups, and everyday operation.

Previously docker-whisper · Maintained by Lin Song (hwdsl2)

Recordings → useful text

Ready to read.
Ready to build with.

00:00 – 00:04 · SPEAKER_00

Let's review the deployment plan.

00:04 – 00:08 · SPEAKER_01

We'll start on CPU, then add a GPU when we need it.

  • JSON
  • text
  • verbose JSON
  • SRT
  • WebVTT
Illustrative transcript with optional speaker diarization.

A focused API for your audio

From recordings to transcripts and subtitles.

Use ScribeCrate for meeting recordings, interviews, podcasts, subtitles, and applications that support the OpenAI Whisper API. Audio is processed on your server.

Images are automatically built and published through public GitHub Actions workflows. Powered by faster-whisper using Whisper models.

Start with one recording

Run locally.
Make your first request.

This CPU example binds the API to your own computer. It uses the established hwdsl2/whisper-server image and a persistent Docker volume.

  • Docker on a supported Linux host
  • At least 700 MB of available RAM for the default base model
  • Internet access for the first model download, unless models are pre-cached

Need NVIDIA acceleration, another model, diarization, or remote access? Follow the full deployment instructions. Use HTTPS and authentication when exposing the API to other computers.

  1. Start ScribeCrate

    docker run \
      --name whisper \
      --restart=always \
      -v whisper-data:/var/lib/whisper \
      -p 127.0.0.1:9000:9000 \
      -d hwdsl2/whisper-server

    The base model downloads on first start. Run docker logs whisper and wait for “ScribeCrate transcription server is ready”.

  2. Retrieve your API key

    scribecrate_api_key="$(docker exec whisper whisper_manage --getkey)"

    A fresh install with this volume generates and persists an API key.

  3. Transcribe an audio file

    Replace audio.mp3 with your file, then run in the same shell:

    curl http://127.0.0.1:9000/v1/audio/transcriptions \
      -H "Authorization: Bearer $scribecrate_api_key" \
      -F file=@audio.mp3 \
      -F model=whisper-1

    Example JSON response:

    {"text": "Your transcribed text appears here."}

Part of a larger toolkit

Build your Self-Hosted AI Stack

ScribeCrate is also available as part of Self-Hosted AI Stack. Deploy local models, chat, document processing, voice, and AI tools together with a single command.

Explore Self-Hosted AI Stack

Need live microphone transcription?

A companion for live audio

ScribeCrate focuses on uploaded audio files. The separately deployed WhisperLive project handles live microphone and audio streams over WebSocket.

Explore the WhisperLive server

Before you deploy

A few useful details.

Is ScribeCrate free to use?

The ScribeCrate project is open source under the MIT License. You provide the hardware or cloud server and pay any associated costs. Models and bundled dependencies retain their own licenses.

Does this website accept audio uploads?

No. This website provides information and setup instructions. Deploy ScribeCrate on your own hardware to process audio. An optional Colab demo runs in Google's environment and is subject to its data practices.

What does OpenAI-compatible mean here?

ScribeCrate provides transcription and translation endpoints for clients that support the OpenAI Whisper API interface. It runs locally using faster-whisper; an OpenAI account or key is not needed. It does not implement every option of every OpenAI audio model. See the API documentation.

Does streaming mean live microphone transcription?

ScribeCrate streams decoded segments from an uploaded audio file over SSE. For live microphone or WebSocket audio, use the separate WhisperLive server. Speaker diarization requires full audio analysis and is not applied in SSE streaming mode.

Can it run entirely offline?

Yes, with models cached in advance and WHISPER_LOCAL_ONLY=true. See offline configuration. The image also has optional aggregate usage counts; disable those separately with WHISPER_DISABLE_USAGE_COUNTS=1. See the usage-count documentation.

Bring your audio. Keep control of where it runs.

View on GitHub