Skip to content

Machine specs for local Whisper

Reki note local Whisper performs transcription entirely on your own PC. This page summarizes how much machine power you need for comfortable use, and how the models differ.

Last updated: September 30, 2026

Local Whisper runs OpenAI Whisper transcription on your own PC instead of in the cloud. Reki note bundles a whisper.cpp-based runtime, so audio and transcripts never leave your machine.

All processing happens on your CPU or GPU, so your hardware determines how fast transcription runs and how large a model you can use.

These are guidelines for running local Whisper at a practical speed. A GPU is not required, and CPU-only setups work well with the right model.

  • Memory (RAM): 8GB minimum, 16GB recommended. Processing uses roughly 2-4GB; on an 8GB machine, aim to leave 3-4GB free.

  • CPU: a modern 4-core or better CPU is recommended. AVX2 support helps it stay fast.

  • GPU: optional. An NVIDIA GPU with 6-8GB VRAM handles larger models faster. Apple Silicon (M-series) integrated GPUs are also supported.

  • Storage: free space to store downloaded models. Models range from about 77MB to 3.1GB.

Whisper models come in different sizes: larger means higher accuracy but more disk, memory, and processing time. The main models available in Reki note are:

  1. tiny (~77MB): lightest and fastest, lower accuracy. About 1GB VRAM.
  2. base (~148MB): good balance of speed and accuracy. About 1GB VRAM.
  3. small (~488MB): higher accuracy. About 2GB VRAM.
  4. medium (~1.5GB): high accuracy. About 5GB VRAM.
  5. large-v3-turbo (~1.6GB): fast and accurate. About 6GB VRAM.
  6. large-v3-turbo-q5_0 (~574MB): a compact quantized model, fast and accurate. About 3GB VRAM.
  7. large-v3 (~3.1GB): highest accuracy. About 10GB VRAM; slow on CPU only.

whisper.cpp model list

On a PC with no GPU or about 8GB of memory, large-v3-turbo-q5_0 or small is the easiest choice. Quantized models (q5_0) use less memory and load faster, and the accuracy difference is usually negligible.

If you have 16GB or more and want the best experience, try large-v3-turbo. Test it on a sample meeting, and if it feels heavy, step down to a smaller model.

These specs are guidelines, and real-world speed depends on your CPU, GPU, OS, and model. The easiest and most reliable check is to tell an AI chat about your own PC.

  1. Check your specs. On macOS, use About This Mac from the Apple menu; on Windows, use Settings > System > About.
  2. Paste the details (OS, CPU, memory, GPU) into an AI chat such as ChatGPT, Claude, or Gemini.
  3. Ask something like: “Can this PC run whisper.cpp large-v3-turbo at a practical speed? Please cover both CPU-only and with a GPU.”
  4. If the answer says it is too heavy, switch to a lighter model such as small or large-v3-turbo-q5_0.

For more detail, see the official repositories.

whisper.cpp (GitHub)

openai/whisper (GitHub)