Skip to content

YazSes

Offline, on-device voice dictation for Linux, macOS & Windows. Hold a key, speak, release — your words are transcribed locally with faster-whisper and typed into any focused app. No cloud. No API key. No subscription. Nothing leaves your machine.

Get started Install from PyPI Star on GitHub

hold→ speak→ release→ text appears

YazSes — hold a key, speak, release; the text is typed into the focused app

40-second tour: the core loop, the command line, and the system tray. Terminal output is real; the command-line typing is re-enacted for legibility. Watch it with narration and chapters on YouTube

yazses doctor — all green, fully offline by default

Why YazSes

  • Fully offline by default & private


    Speech is transcribed on-device with CPU faster-whisper (int8). No GPU, no network, no account. No audio, no text — nothing leaves the machine by default.

    Privacy statement

  • Types into any app


    Hold the hotkey, speak, release — the text lands in whatever window has focus: editor, browser, terminal, chat. Works on X11 and Wayland.

    Install on Linux

  • Works over SSH and Remote-SSH


    Because text is injected at the OS level, not inside an app, it lands in VS Code / Cursor Remote-SSH panes, integrated terminals, tmux and container shells — where in-app dictation usually can't reach.

    Dictation over SSH

  • Voice commands & macros


    A regex grammar maps "undo that", "save file", "go to line 42" to real key sequences.

    Command index

  • Transcribe recordings


    yazses transcribe meeting.m4a turns any audio/video file into text — offline, with --diarize speaker labels and subtitle export.

    Transcription guide

  • Built for accessibility


    VAD calibration, mic-level tuning, dysfluency-friendly mode, and an EMG muscle-sensor trigger for fully hands-free use.

    Features

  • Self-improving, on your terms


    An opt-in, encrypted, on-device learning corpus lets yazses tune propose accuracy fixes from your own corrections.

    Performance tuning

Install

pipx install yazses
bash <(curl -fsSL https://raw.githubusercontent.com/MSKazemi/yazses/main/install-apt.sh)

The snap on Wayland

The strictly confined snap dictates on X11, and on GNOME/KDE Wayland through the xdg-desktop-portal RemoteDesktop API. It cannot configure or use the host ydotoold service, so the portal is the route it takes; the first dictation asks once for permission and the answer is remembered. The unconfined installs above remain the most-tested Wayland path.

On X11, connect both required interfaces after installing:

sudo snap install yazses
sudo snap connect yazses:audio-record
sudo snap connect yazses:raw-input
yazses doctor

yazses setup is not required for the snap. Confinement prevents it from installing host packages, changing groups, or configuring host services, so it can only print the manual permission checklist; the snap already bundles its X11 dependencies.

Non-Snap Linux installs — provision the system in one command (the install-apt.sh / APT path does it automatically):

yazses setup        # installs audio + injection deps, joins the input group, sets up ydotoold
# then log out and back in (the input-group change needs a fresh login)

This installs libportaudio2 (audio), the X11/Wayland injection tools, adds you to the input group, and — on GNOME/KDE Wayland, where wtype is blocked — sets up ydotoold (the only way to inject keystrokes there). Re-run it anytime; it only fixes what's missing.

Then:

yazses doctor     # check mic, injection backend, permissions (want [OK] Keyboard capture)
yazses enroll     # calibrate your microphone (~30 s)
yazses start      # start the dictation daemon

Hold the hotkey (Right Alt on Linux, Right Option on macOS, Right Ctrl on Windows), speak, release — the text appears in the focused app within about a second.

What it does

  • Offline dictation — type into any focused app with on-device faster-whisper (CPU, int8). No GPU needed.
  • Transcribe recordingsyazses transcribe meeting.m4a turns any audio/video file into text, fully offline. Add --diarize to tag who said what and --format srt for subtitles. See the transcription guide.
  • Voice commands — a regex grammar maps phrases to editor/terminal key sequences: "undo that", "save file", "go to line 42", "run the tests", "rename this to user_id".
  • Macros & personal vocabulary — define multi-step commands and teach YazSes your mis-heard words.
  • Dysfluency-Friendly Mode — opt-in collapse of stutters/repeats for stuttered or dysarthric speech.
  • Self-improving — opt-in, encrypted on-device learning corpus; yazses tune proposes accuracy fixes from your own corrections.
  • Accessibility — VAD calibration, mic-level tuning, and EMG (muscle-sensor) trigger support.

What people use it for

Browse all use cases

When not to use it

YazSes is not an LLM agent — it dictates text and runs editor/terminal commands; it does not browse, reason over your files, or hold a conversation. It uses CPU faster-whisper (a cloud service may still win on raw accuracy for a noisy mic), ships English-tuned *.en models by default, and is desktop-only.

How it works

Hold hotkey → record audio → VAD gate → faster-whisper (CPU)
            → clean + disfluency filter → command grammar (Tier 1 regex)
            → dictate? type it · command? send keys

Two at-a-glance signals show you what YazSes is doing: a top-bar "Y" tray icon whose colour is a live state indicator (🔵 idle · 🟢 dictating · 🟡 no text target → clipboard · 🟣 command mode · 🔴 problem) and an optional sonar overlay that pulses near your cursor while it's listening. See Tray icon & overlay for what each colour means.

Want the CLI itself, not a description of it? Watch it play right here — a real asciinema recording of -haboutquickstartfeaturesstatus, with selectable, copy-pasteable text. Or in your own terminal: asciinema play docs/demo/yazses-cli.cast.

Documentation

Curious why it works this way?

The research section is a public, fully-cited notebook on post-keyboard input: how accurate webcam eye tracking really is, how offline speech recognition overtook the cloud, and why a $50 muscle sensor beats a $1,000 EEG headset. Every design decision in YazSes traces back to a measurement there — and the open questions are open to anyone.

FAQ

Does it work without internet? Yes — transcription runs locally; nothing is sent anywhere by default.

What GPU do I need? None. It runs on CPU; 4 GB RAM minimum, 8 GB comfortable.

Does it work on Wayland? Yes via the APT or pipx install (uses wtype/ydotool). Use one of those, not the snap — strict confinement prevents the Snap from using the host injection service, so it cannot type the result into Wayland applications.

Is it a replacement for Talon? YazSes focuses on offline dictation plus a practical command grammar. Talon has far more advanced scripting. They can coexist.

More answers in the full FAQ, and a side-by-side in Comparison & alternatives.


Apache-2.0 licensed. If YazSes is useful to you, a ⭐ on GitHub helps others find it.

Want to help? See Contributing — several tasks need no code at all, like translating the README, sharing a config you already use, or adding the microphone that worked for you.