Skip to main content

OpenTranscribe v0.5.0 - Native Rust Diarization, Chat With Your Transcripts, and a Full Identity Overhaul

· 15 min read
OpenTranscribe Team
OpenTranscribe Development Team

This is the biggest release OpenTranscribe has shipped, and it took longer than we wanted. OpenTranscribe is a side project on top of a full-time job, and this cycle kept growing every time we got close to calling it done. A feature would surface a bug, the bug would surface a design question, and the question would turn into three more issues. We'd rather ship it right than ship it on a schedule.

This release contains breaking changes and action-required upgrade steps. Read Upgrade Notes before pulling.

A Native Diarization Engine, Written in Rust​

Speaker diarization has run on PyAnnote since day one. It's accurate, but it's also a lot of Python for something that runs on every upload. Rather than guess where the overhead was, we profiled it, and once we understood exactly where the time went, we built diar-server: a Rust sidecar that wraps speakrs to run pyannote-equivalent processing natively (issue #538), now the default on-box diarization engine. PyAnnote is still there as an automatic fallback if the sidecar isn't reachable. We didn't want a rewrite to become a single point of failure.

None of this happens without speakrs existing in the first place, and we're genuinely grateful for the work that went into it. We also contributed performance work back upstream, getting close to a 4x speedup over the package's original processing speed, so the improvement runs in both directions.

The real win is that transcription and diarization now run concurrently instead of back-to-back: measured 87.8s to 50.3s, 43% faster, on a 66.5-minute test clip, with byte-identical output verified against the old path.

Alongside the engine swap, we also fixed what diarization itself gets wrong: it's strongest mid-turn and weakest at the seams, so the exact word where one speaker stops and another starts gets misattributed more than you'd expect. A word-boundary smoother (default on) collapses these "wrong-speaker islands": measured 32% lower word-level speaker error, islands down from 82 to 15 on our test set. An experimental acoustic re-check recovers absorbed backchannels like "mm-hmm" for another 15% on top.

The engine also adapts to hardware it didn't fit on before. A new hybrid mode runs transcription on CPU and diarization on GPU or MPS automatically when the GPU is too small for both, which unlocks 4 to 6 GB NVIDIA cards and Apple Silicon for local diarization for the first time. On multi-GPU hosts, an optional split routes transcription and diarization to separate cards for more throughput.

Chat With Your Transcripts​

The feature we're most excited about: chat with your transcripts, right in the app. Ask a question about any meeting, lecture, or interview you've transcribed and get a streamed, cited answer: click a citation and it deep-links to the exact moment in the recording.

We did not ship the first version that worked. At real library scale, naive retrieval either misses what's relevant or floods the model with too much text to find the answer in. What shipped is a real pipeline: conversational query rewriting, hybrid keyword + vector search over speaker-turn chunks, cross-encoder reranking, and, for libraries too large for one retrieval pass, a query router and map-reduce leg over per-recording digests (issue #403, routing measured at 0.104% leakage). A Speakers scope filter is exact rather than fuzzy: because transcripts are indexed by speaker turn, filtering to one person retrieves only their words, so "what did Dana commit to?" can't be answered from someone else's sentence about Dana.

Trust mattered as much as capability here. A live query-execution trace panel (issue #514) shows what was actually retrieved before the answer arrives. The model's context window is now discovered from the provider instead of a hardcoded default that was silently truncating long transcripts (issue #533). Reasoning/thinking display is measured per model rather than assumed to work (issue #64). And an answer built from zero retrieved excerpts is flagged as such, not presented with false confidence.

Chat also picked up the interaction details that make a chat product usable day to day: edit-and-regenerate, stop mid-stream, conversation export, and projects that pin a default scope for a client or recurring meeting. Plus a new Amazon Bedrock provider and usage tracking so you can see what you're actually spending, in tokens, on every model.

Search got real attention too. Tag search, which was quietly broken, now works, and tags gained proper management and sharing. Switching to a multilingual embedding model is a one-click change in Settings now, not a redeploy decision, and it's measured rather than assumed: real Spanish queries scored nDCG@10 0.76 with the multilingual model against 0.66 with the English default.

Authentication, Identity, and Security​

The single biggest pain point we heard about was OIDC: it only ever really worked with Keycloak, because the login flow hardcoded Keycloak's URL shape. Every other provider (Authentik, Okta, Entra ID, Auth0, Zitadel) got a 404 on the redirect. OIDC is now driven by the provider's own discovery document, with provider presets that fill in the roles-claim path and known quirks for you. KEYCLOAK_* environment variables and existing configuration keep working with zero reconfiguration; nothing needs to change on your identity provider.

That sits alongside a genuinely large identity overhaul: LDAP, SAML, PKI/mTLS, header-based proxy auth, MFA, and SCIM provisioning under one privilege model. Plus GDPR erasure-ledger hardening, FedRAMP account-inactivity enforcement, and a security pass that closed roughly twenty issues (SSRF gaps, a quarantined-file data leak, session and tenant-isolation defects). Most of these came from us adversarially reviewing our own code, not from bug reports, which is the outcome we want, even if it's slow, unglamorous work that doesn't produce a screenshot.

Two more access-control fixes landed in final release-readiness auditing, and we're calling both out by name because they affect any already-deployed instance:

  • Revoking a share could still serve the revoked file's content from the chat retrieval cache. We reproduced it live before writing a fix: grant access → ask a question (chunks from the file come back) → revoke the share → ask the identical question again → the cache still returned the revoked file's chunks. This was TTL-bounded and self-healing — the cache entry expires on its own — but bounded is not the same as fine, so the cache is now invalidated at the moment the access-control index finishes rewriting after a share change.
  • SCIM and OIDC/IdP group membership changes never updated file search/chat visibility at all. Removing a user from a group via SCIM or an identity provider's group mapping updated the group membership row, but nothing re-indexed which files that user could still see. Unlike the cache issue, this one was not time-bounded — a removed user kept search and chat visibility of every file shared with that group indefinitely, until some unrelated write happened to touch the same index rows. For any deployment that provisions access from an IdP (which is exactly the deployment that uses SCIM), deprovisioning a user did not actually deprovision them from OpenTranscribe. See Upgrade Notes for what this means for existing group memberships.

Content Redaction​

A new redaction subsystem detects PII, profanity, and toxicity and masks it wherever a transcript is shown: the transcript view, search, summaries, subtitles, exports. It's built on real NLP models (Presidio, spaCy NER, toxicity classifiers), not a word blocklist, so it understands context instead of over-masking ordinary conversation. Detection runs once per transcript on a dedicated worker and is cached; masking is applied at read time, so changing what you want redacted never means reprocessing anything. Users can opt out of their own redaction; an admin can set a floor nobody can drop below.

Watch Sources: Automatic Ingestion​

New this release: point OpenTranscribe at a local folder, an S3-compatible bucket (AWS, MinIO, Backblaze, Wasabi), or an SMB/CIFS network share, and it polls on its own schedule, copies in new media, and runs the full transcription, diarization, and embedding pipeline with no manual upload. Originals on a remote source are never moved or deleted; local sources can optionally clean up after import.

This cycle also added per-file visibility into what a source is actually doing (issue #489), email notifications scoped to a specific source (issue #490), and a fix for a source that could import the same recording twice under two different names.

Also In This Release​

Cloud ASR (AWS Transcribe, Speechmatics, AssemblyAI, Gladia, pyannote.ai) is now production-verified end to end rather than experimental, and worth saying plainly: our own local pipeline hit 0.27% word-level speaker error at 41× realtime, tying the best commercial engine we benchmarked. There's also a new CrisperWhisper model option with notably precise word-level timestamps. Downloads and playback now go through short-lived presigned MinIO URLs instead of the backend proxying every byte, which removes a real bandwidth bottleneck and makes video seeking faster too. Scheduled backups are now fully admin-UI configured, with GPG encryption and retention policy, no host cron required. The UI now supports 12 languages, four added this release (Italian, Arabic, Korean, and Dutch) — and switching between them actually works now. Final hands-on testing caught that the UI rendered the previous language after every switch: pick a second language and it only partly changed, pick a third and you got the second. The store updated synchronously while the translations loaded asynchronously, and the re-render that should have followed was silently skipped because it depended on a value that had already changed. Both a unit test and a browser test now cover it, and neither asserts on the lang attribute — that was correct the whole time the UI was visibly wrong. And a good share of the smaller fixes in this release (the LDAP full-DN group filtering bug, several security findings, watch-source rough edges) started as a GitHub issue from someone running this in production. Thank you for those.

The API also picked up a committed OpenAPI schema snapshot with a CI diff gate, so an accidental wire-contract change now fails a pull request instead of shipping quietly. And a release-readiness documentation audit went through the production-deployment docs, README, and this project's contributor guide line by line against the running code. It found, among other things, that our production-deployment guide referenced a SECRET_KEY environment variable that the backend has never read — the real name is JWT_SECRET_KEY — and that a documented "verify your secrets are set" command was built around that same wrong name, so it could never actually confirm your secret was set. Both are fixed. If you configured a production instance from those docs, it's worth a look at your .env to confirm JWT_SECRET_KEY is set to a real value rather than a placeholder (the fail-closed startup check described below will refuse to boot on a placeholder value regardless, so a misconfigured instance won't come up silently — but re-checking costs you nothing).

Two supply-chain fixes landed in the final release hours, both found by the release scan rather than by review, and both worth naming because they were invisible from inside the code:

The lite image had no mechanism to receive OS security patches at all. Dockerfile.prod has always run apt-get upgrade -y; Dockerfile.lite never did, so every lite image we published shipped whatever its base image contained at the pinned tag, indefinitely. That's the worst possible file for it to be missing from, because arm64 hosts default to lite — meaning the one backend an arm64 user can install was the one with no patch path. Worse, Dockerfile.prod's upgrade wasn't running either: Docker caches a RUN by its literal command string, so once that layer existed it was reused verbatim and apt never executed again. The comment above it claimed it refreshed "on every image build", and our own changelog repeated that claim. Both were wrong. All three backend Dockerfiles now gate that layer on a date-keyed build argument so it genuinely re-runs, which took measured CRITICAL counts from 16/16/19 down to 1/1/4 across the backend and lite images. What's left has no upstream fix to take. A test now fails the build if any production image loses its upgrade or its cache-busting argument again.

MinIO removed its images from Docker Hub. minio/minio and minio/mc are simply gone — not rate-limited, deleted — which breaks every new install while existing hosts with a warm image cache keep working, so it's invisible in day-to-day development and fatal on a fresh one. The object store now pulls from quay.io/minio/minio at the same pinned tag and a byte-identical digest, verified by pulling both and comparing rather than trusting that two registries agree. No action is needed on your side; the compose file already points at the new location.

Upgrade Notes​

⚠️ Action required: the backend now fails closed on missing production secrets. ENVIRONMENT defaults to production, and startup now refuses to boot without a real REDIS_PASSWORD, non-placeholder JWT_SECRET_KEY/ENCRYPTION_KEY, and DEBUG off. Previously an unset ENVIRONMENT (the documented normal case) silently skipped all of these checks. Set them before upgrading, or the stack will not come back up.

A few more to know about:

  • The ASR/Engine/Backup/Media-Mirror/Watch-source/Redaction admin panels now require super_admin, not admin. Promote the relevant accounts first (now possible from the UI).
  • An OIDC login will no longer silently take over a local account sharing its email address on providers that don't assert email_verified (Authentik, Entra ID). Link affected accounts deliberately via oidc_subject; there's no flag to restore the old, riskier behavior.
  • Breaking: the legacy /video, /simple-video, /content, /download, and /download-with-token media endpoints are removed in favor of presigned-URL streaming (stream-url / prepare-download). bulk-export is now async + SSE. File fingerprints regenerate automatically on first startup. No action is needed, but cross-pipeline dedup is unreliable until it finishes.
  • ⚠️ Action recommended if you use SCIM or IdP group provisioning: the fix described above only prevents new group-membership changes from leaving stale search/chat access behind. It does not retroactively re-check group memberships that changed before you upgraded. If any user was ever removed from a group via SCIM or an IdP's group mapping on a pre-upgrade instance, they may still have search/chat visibility into files shared with that group. There's no automatic full re-check of existing memberships on startup, so after upgrading, an admin should trigger a re-index (POST /api/search/reindex from an admin session, or the equivalent action in Settings → Search) to force those access lists to be recomputed. If you've never used SCIM or IdP group provisioning, this doesn't apply to you.
  • All migrations apply automatically on backend startup.
  • Stopping a GPU worker mid-transcription is still not safe. Workers now release the GPU on shutdown and get a measured 30-second grace period instead of being killed after 10, so an idle or between-stages stop is clean. But an in-flight transcription cannot be interrupted: the worker pool waits for the running task to return, so it still runs to the ceiling and is killed. Releasing sooner would free VRAM out from under a live CUDA kernel, which is worse. Let a transcription finish before stopping the stack. Tracked as #809.

Full detail is in the CHANGELOG.

What's Next​

This release took too long, and the fix for that is smaller releases on a fixed cadence rather than a promise to hurry. From here the cadence is monthly, and the scope floats: a release ships with whatever is finished on its date, and the rest rolls to the next one. Each one has a theme and stated exit criteria rather than a pile of issues.

Next up is v0.6.0, focused on retrieval and chat quality — measuring what the RAG pipeline actually returns rather than assuming, and turning on the parts that measurement justifies. After that, document ingestion gets a release to itself; tying documents into transcripts is a big enough feature that we'd rather not bury it in a grab bag.

The live view is the roadmap, generated from GitHub issues rather than hand-maintained, so it can't drift from what we're actually doing. The reasoning behind each release — what it's for and when it's done — is in Release Themes. If there's something specific you want, open an issue; that's genuinely how most of this release got decided.

Thanks​

OpenTranscribe is AGPL-3.0 and self-hosted. Issues and discussions are on GitHub. Real bug reports from real deployments drove most of what's in this release, and we're grateful for every one of them.