Antoine Chatry profile picture, web developer

Antoine Chatry

codewars score

How does Dictata work?

I'm proud to share Dictata, an application I developed in Rust to perform speech-to-text transcription directly on your computer.

The application lets you start a transcription with a keyboard shortcut, speak, and then automatically insert the text into the application you're using.

The goal is to make voice dictation easy to use while keeping your data local. Here's how it works:


1. Audio capture

First, Dictata captures audio from the selected source.

You can use your microphone, system audio, or both at the same time. Once the recording is finished, the audio is sent to the transcription engine.

2. Transcription

To handle transcription, Dictata uses another one of my projects: DictataEngine.

I created this engine to provide a common interface for several speech recognition engines and make them easy to use from a Rust application.

It currently supports Whisper, Parakeet, SenseVoice, Moonshine, and Zipformer.

The advantage is that Dictata can switch between transcription engines without having to modify the rest of the transcription logic. Everything runs offline, directly on your machine.

3. Text insertion

Once the transcription is finished, Dictata retrieves the text and automatically inserts it wherever the cursor is. You simply start Dictata with the keyboard shortcut, speak, and then continue using your application as usual.

A continuous mode also allows you to transcribe progressively while speaking.

4. Text processing

Dictata can also use a local LLM to process the transcribed text.

This can be used, for example, to clean up a transcription or directly turn it into a message, email, or another format.

5. Model management

Dictata also lets you manage the models used for transcription directly from the application.

You can view the available models, download the ones you need, and choose which one you want to use.

This functionality is directly connected to DictataEngine, keeping model downloading and management separate from the application itself.

The project is still under development, but this architecture now allows me to easily experiment with different transcription models and improve Dictata without being tied to a single engine.

You can find both projects on GitHub:

Dictata
Dictata Engine