CaYaScribe

CaYaScribe

Published:

Local, offline speaker-aware transcription for Windows. Audio and video never leave your machine.

CaYaScribe Local, offline speaker-aware transcription for Windows. Audio and video never leave your machine. Languages: English (default) · Türkçe The app UI is English or Turkish. The OS/browser language is used on first launch; anything other than Turkish falls back to English. You can switch in the title bar. Transcription itself covers 99+ languages (Whisper); an unknown language code is treated as English. What it does Extracts audio from MP3, MP4 and other media (LGPL FFmpeg) Transcribes in 99+ languages (Whisper; Turkish and English are first-class) Optional speaker diarization: empty count = auto-detect, 1 = single speaker Speakers labeled A, B, C… — rename globally in one action Export TXT / SRT / VTT / JSON; timestamps are optional On first launch, asks before downloading missing models; you can also download later from Models Releases Windows installers are published on GitHub Releases. Push a tag vX.Y.Z to build NSIS (.exe) and attach it to the release Models are not inside the installer; the app downloads them after you consent The installer includes the local Python engine; models download after you consent See docs/releasing.md. Development setup Requires: Windows 10+, Python 3.12+, Node 20+, Rust (Tauri). On first run the app offers FFmpeg, Whisper turbo, the Turkish Whisper large-v3 fine-tune, pyannote segmentation, and WeSpeaker ResNet293-LM. Uncheck what you do not need. Quality profiles Profile Engine (v0.1) Fast Whisper small Balanced (default) Whisper large-v3-turbo High Turkish Whisper large-v3 fine-tune when downloaded; else turbo Maximum Same TR fine-tune, else multilingual large-v3 Speaker diarization is language-independent (sherpa-onnx + TitaNet). No Hugging Face token is required for the default path. Privacy Speech-to-text never goes to the cloud Downloads happen only after you confirm, from Hugging Face / GitHub No telemetry Architecture: docs/design.md License App code: MIT. Third-party: THIRDPARTYNOTICES.md

Özellikler

  • Extracts audio from MP3, MP4 and other media (LGPL FFmpeg)
  • Transcribes in 99+ languages (Whisper; Turkish and English are first-class)
  • Optional speaker diarization: empty count = auto-detect, 1 = single speaker
  • Speakers labeled A, B, C… — rename globally in one action
  • Export TXT / SRT / VTT / JSON; timestamps are optional
  • On first launch, asks before downloading missing models; you can also download later from Models
  • Push a tag vX.Y.Z to build NSIS (.exe) and attach it to the release
  • Models are not inside the installer; the app downloads them after you consent
  • The installer includes the local Python engine; models download after you consent
  • Speech-to-text never goes to the cloud
  • Downloads happen only after you confirm, from Hugging Face / GitHub
  • No telemetry