known·good

Notes / cudapythonfaster-whisperctranslate2whispernvidiaarch

faster-whisper: Library libcublas.so.12 is not found or cannot be loaded (system has CUDA 13)

· by shaun, written up with Claude

faster-whisper fails at the first GPU compute with "Library libcublas.so.12 is not found" because the ctranslate2 wheel is built for CUDA 12 and the system has CUDA 13. Install nvidia-cublas-cu12 and nvidia-cudnn-cu12 into the venv and set LD_LIBRARY_PATH in a wrapper before Python starts.

Tested onArch Linux; CUDA 13.3.1 (system libcublas.so.13); NVIDIA RTX 5080; Python 3.11.14 venv; faster-whisper 1.2.1; ctranslate2 4.7.1; large-v3 model

Symptoms #

faster-whisper with device="cuda" dies at the first GPU computation with:

textLibrary libcublas.so.12 is not found or cannot be loaded

The system CUDA is version 13 and only has libcublas.so.13 (on Arch, in /opt/cuda/lib64).

A quick smoke test can hide the problem. A test on silence or a synthetic sine wave passes even with cuBLAS missing: voice activity detection (VAD) finds no speech, so the GPU encoder never runs. Always smoke-test speech-to-text on real speech.

Root cause #

The ctranslate2 4.7.1 wheel from PyPI is linked against CUDA 12. It loads libcublas.so.12 (and cuDNN for CUDA 12) at run time. A CUDA 13 system does not provide those libraries, and the loader will not substitute libcublas.so.13 for a .so.12 dependency.

Fix #

Put the CUDA 12 runtime libraries inside the venv, so it does not depend on what the system or another application happens to ship:

bash<venv>/bin/pip install nvidia-cublas-cu12 nvidia-cudnn-cu12

These wheels put the libraries under site-packages/nvidia/cublas/lib and site-packages/nvidia/cudnn/lib. The dynamic loader still has to find them, so start Python through a wrapper that sets LD_LIBRARY_PATH:

bash#!/usr/bin/env bash
VENV=<venv>
SP="$VENV/lib/python3.11/site-packages"
export LD_LIBRARY_PATH="$SP/nvidia/cublas/lib:$SP/nvidia/cudnn/lib:${LD_LIBRARY_PATH:-}"
exec "$VENV/bin/python" transcribe.py "$@"

Setting os.environ["LD_LIBRARY_PATH"] inside the Python script does not work: the dynamic loader reads LD_LIBRARY_PATH when the process starts, so it has to be exported before exec. Adjust python3.11 in the path to your venv's Python version.

An earlier workaround pointed LD_LIBRARY_PATH at the CUDA 12 libraries bundled with another application (ollama's cuda_v12 directory). That works, but it breaks if that application is removed or its bundled CUDA moves to version 13. The venv-local wheels avoid that.

With this in place, sequential transcription with large-v3 ran at about 31x realtime on an RTX 5080 (5.35 hours of audio in about 11 minutes of GPU time).

In the same bulk job, BatchedInferencePipeline crashed on a file that was music only:

textRuntimeError: No clip timestamps found. Set 'vad_filter' to True or provide 'clip_timestamps'.

The batched pipeline needs VAD to return at least one speech clip and raises when it returns none. One music-only or silent file stops a whole batch. The sequential WhisperModel.transcribe() handles such files without raising, and at ~31x realtime it was fast enough that the batched path was not worth it for bulk work:

pythonfrom faster_whisper import WhisperModel

model = WhisperModel("large-v3", device="cuda", compute_type="float16")
segments, info = model.transcribe(path, language="ja", vad_filter=True,
                                  beam_size=5, condition_on_previous_text=False)
segments = list(segments)   # it is a generator; consume it before writing output

condition_on_previous_text=False stops one hallucinated line from seeding the next.

On music-only audio, large-v3 may still "transcribe" a stock filler phrase (in Japanese, ご視聴ありがとうございました, "thank you for watching") once or twice and nothing else. It looks like a transcript but is hallucination. A cheap check that worked for flagging such files:

pythonuniq  = {s.text.strip() for s in segments if s.text.strip()}
chars = sum(len(s.text.strip()) for s in segments)
no_speech = len(uniq) <= 2 and chars < max(40, duration / 20)

For checking a whole batch without reading every file, characters per second of audio was a useful health metric: conversational Japanese interviews landed in a tight 5.1–5.8 chars/s band, and the outliers were real silence or music, not failures.