Speech recognition, turn detection and text-to-speech all run on your own machine. Ask a question out loud, get an explanation spoken back, and keep going — with the Wi-Fi off and no audio leaving the room.
Four models cooperating on-device: speech recognition, turn detection, the tutor itself, and a voice.
Speech-to-text runs on-device with Parakeet models. Your voice is transcribed on your own machine and no audio is streamed to a server.
A dedicated end-of-utterance model works out when you have finished speaking, so a conversation flows instead of waiting for a fixed pause or a button press.
Silero VAD separates speech from room noise, which is what makes hands-free use workable in an actual classroom rather than a quiet studio.
The tutor speaks back using on-device voices. Pick a voice, and it works with the network off like everything else.
Build an agent per subject with its own system prompt and model. A physics tutor that works through derivations behaves differently from a language partner.
Agents are set up to work a problem through with you. That is both better pedagogy and the use of AI your school policy almost certainly permits.
Speak in a language you are learning and get corrections explained rather than silently applied, with multilingual models handling both sides.
The same speech engine transcribes a recorded lesson or an interview into text you can then summarise or search.
Typing a question is a small tax, but it is charged every time. It is enough that a student who is stuck on step four of a derivation will usually just stare at it rather than describe the problem in a text box. Speaking removes that tax, and the number of questions asked goes up sharply.
It also changes what kind of question gets asked. Typed questions tend to be well-formed and final. Spoken questions are half-formed, which is exactly the state a student is in when help would be most useful.
Text is one thing. A continuous audio stream from a child’s desk is another, and it is the part parents and data-protection officers object to hardest — with reason. A cloud voice assistant means a microphone feed leaving the building.
On-device speech recognition removes the stream. The audio is transcribed on the machine it was captured on and is never transmitted, which turns a difficult conversation into a short one. It also means no per-minute transcription bill, so a student can talk for an hour without anyone watching a meter.
Yes. Grout runs speech-to-text, turn detection and text-to-speech on-device, alongside a local language model, so you can hold a spoken conversation with a subject tutor with the network disconnected. No audio is streamed anywhere.
Yes. On-device text-to-speech reads the tutor’s replies aloud with a voice you choose. Because the voices are local, this works offline and has no per-minute cost.
Yes, anywhere in the studio. Voice activity detection and end-of-utterance detection mean you can talk naturally rather than holding a button, which matters when your hands are on a worksheet.
The specific risk with voice assistants is that a child’s speech is recorded and sent to a company. On-device processing removes that transfer — the audio is transcribed locally and never uploaded. Beyond that, guardians can see activity and set screen time limits.
Yes. Each agent has its own system prompt and model, so you can configure a chemistry tutor that insists on balanced equations and a separate essay coach that never writes prose for you.
The studio ships multilingual speech and language models, and the interface itself runs in several languages. Quality varies by language — widely spoken languages are noticeably stronger than rare ones.
Free download for Windows and macOS. Speech in, speech out, nothing uploaded.
Download Grout free