LEFTEROS / PERSONAL ARCHIVE
RecMeets' Laya: which word was misheard
29 September 2026 / Lefteris Iliadis / SomniusX
somniusx/recmeets-laya-words is a model I trained and published on Hugging Face. It answers one question: given a word in its sentence, as a speech recognizer wrote it, how likely is it a mistake? It works in Greek and English, fits in 343 MB and runs on the processor, with no graphics card and no cloud. It is open (Apache 2.0), so anyone who needs it can use it in their own projects, and I keep improving it.
It is a fine-tune of Laya (laya-multilingual, Apache 2.0), a multilingual decision model. I built it for Check words in RecMeets, the step that shows which words of a transcript were probably misheard, so you can listen and fix them. Its Hugging Face page holds the description, the data, the results and the three files you need: the model in ONNX, the tokenizer and a small settings file.
Why it was needed
A speech recognizer such as Whisper gives every word a “confidence”. In Greek, though, that alone barely separates the mistakes: on the words the engine is unsure of, just 0.63 on a scale where 0.5 is chance and 1.0 is perfect. A spelling dictionary helps, but it misses words that exist yet are wrong in their place, such as «είμαι» (“I am”) for «είναι» (“is”).
So it needed something that reads the word in its sentence, knows what the engine and the dictionary said, and gives an honest probability. It had to run on the user’s own computer, quickly, and send nothing anywhere, because meetings are private.
There is a ready-made option in the cloud, TypeSafe’s Jev, with a key and per-use pricing. The goal was a local model that comes close to it.
Laya in brief
Laya doesn’t write text like a chatbot. You give it a state (text or JSON) and typed questions (yes/no, choice, score), and in a single pass it answers with calibrated probabilities, in more than 100 languages. It was trained with reinforcement learning against strictly proper scoring rules (RLCD), so the only way for it to score well is to report honest probabilities. Because it generates no text, it has nothing to make up.
Its makers say it plainly: the base checkpoints are near chance without training on the task, and the value is in fine-tuning. I confirmed it: on twelve Greek words, six right and six misheard, no wording of the question separated them. On the test set, stock laya-multilingual scored 0.41, below chance.
The data: real mistakes, not guesses
The model learns from real mistakes, the ones made by the very engine RecMeets users run:
- FLEURS by Google (CC BY 4.0): sentences read aloud, with exact reference text, in Greek and English.
- I transcribed them with RecMeets itself (whisper.cpp, the large-v3-turbo model, with RecMeets’ settings), joining the sentences into 10-minute files with 1.2 seconds of silence between them.
- Every word the engine wrote was labelled right or wrong by aligning it with the reference text.
- The dictionary’s answer for every word came from LibreOffice’s hunspell dictionaries (el_GR, en_US).
- Training: 71,906 Greek and 54,845 English words (9,880 and 2,547 mistakes), balanced so that 35% are mistakes.
- Testing: FLEURS’ test split, with speakers the model never heard in training.
The question I ask it
For every word, Laya gets a yes/no question (noul) and this state:
Sentence: <about ten words either side, the word marked ⟦ ⟧>
Word: <the word>
Engine confidence: <0.00–1.00, or unknown>
In the <Greek|English> dictionary: <yes|no|can't tell>
It answers with the probability that the word is a mistake, corrected by a temperature measured on a separate slice of the data. It is plain text on purpose, so that RecMeets, written in Rust, builds it byte for byte the same as Laya’s Python.
Training
I followed Laya’s own recipe: RLCD (policy gradient with rewards from proper scoring rules) plus soft cross-entropy, a slice of the data held out for calibration, and a temperature per question type at the end. Everything ran locally, in a container with PyTorch and CUDA, with the processor use capped so the computer stays usable. One epoch takes about 8 minutes.
| Run | Data, epochs | Greek, all words | Greek, unsure | English, unsure |
|---|---|---|---|---|
| 1 | Greek, 3 (21 minutes) | 0.85 | 0.78 | |
| 2 | Greek, 3, with the dictionary in the input | 0.90 | 0.89 | |
| 3 | Greek + English, 4 | 0.88 | ||
| 4 (published) | Greek + English, 2 | 0.90 | 0.91 | 0.86 |
Run 3 overfit: in its fourth epoch the loss went below zero. Run 4, with two epochs, is the one I published.
The computer
| Graphics card | NVIDIA GeForce RTX 4080 SUPER, 16 GB |
| Processor | AMD Ryzen 9 5950X, 16 cores / 32 threads |
| Memory | 64 GB |
| Storage | Samsung 990 PRO NVMe, 2 TB and 1 TB |
| Motherboard | Gigabyte X570 AORUS ELITE |
| System | Omarchy (Arch Linux with Hyprland) |
The graphics card was needed only for training. The finished model runs on the processor.
Results

AUROC on the FLEURS test split: 1.0 means every mistake scores above every right word, 0.5 is chance. “Unsure” are the words with an engine confidence under 0.5 and at least four letters, the ones Check words lists.
| Greek, all | Greek, unsure | English, unsure | |
|---|---|---|---|
| The engine’s confidence alone | 0.82 | 0.63 | 0.74 |
| The dictionary alone | 0.85 | 0.70 | |
| Stock laya-multilingual, same question | 0.41 | ||
| This model | 0.90 | 0.91 | 0.86 |
| TypeSafe Jev (cloud), same input | 0.93 | 0.89 |
On the words that matter, the local model comes within a whisker of the cloud’s Jev: free, on your computer, with nothing leaving it.
Size and speed
Exported to ONNX, the model is 1.29 GB at 32 bits, mostly mmBERT’s embedding table with its 256,000 tokens. There was a trap there. ONNX Runtime’s usual 8-bit quantizer with per-channel scales, the one Laya’s script uses, broke the fine-tuned model: it gave every sentence the same answer. With per-tensor scales it lost about 0.03.
The fix was ONNX Runtime’s 8-bit block format (MatMulNBits, blocks of 64) for the matrix products, and 8 bits for the embedding table: 343 MB, as accurate as 32 bits. With 8-bit arithmetic on the processor (accuracy_level 4), a sentence takes about 110 ms on 4 threads, instead of 140.

| Packaging | Size | ms per sentence | AUROC Greek / English |
|---|---|---|---|
| 32-bit | 1.29 GB | 79 | 0.905 / 0.859 |
| 8-bit blocks | 343 MB | 140 | 0.905 / 0.861 |
| 8-bit blocks, 8-bit arithmetic (published) | 343 MB | 109 | 0.905 / 0.860 |
| 4-bit blocks | 290 MB | 77–80 | 0.889 / 0.854 |
The 32-bit model is faster, but nearly four times the download. The 4-bit one is fast too, but loses accuracy. I kept the sweet spot.
Inside RecMeets
The model is built into RecMeets, with no Python and nothing to install. It downloads once, 378 MB with the tokenizer, pinned to a specific revision and checked by SHA-256. RecMeets builds the input exactly as Laya’s Python does, and a test compares it token for token. It runs through ONNX Runtime on 4 processor threads, so the rest of the computer stays free.
To save time, when a word appears in two places, the second is checked only if the first answer is unsure (between 0.3 and 0.7). On a real reading, 146 words were checked in 23.4 seconds. Users can also choose Jev in the cloud if they have a key.
What got fixed along the way
To build the data, RecMeets transcribed hours of read sentences, and that showed four problems in RecMeets itself:
- Loops: the same sentence written 40 times in a row, where 30 different ones had been read. Whisper continues each 30-second window from the text of the one before, and once a window repeats, the next carries the repetition on.
- Skipped sentences: when two people read the same sentence one after the other, Whisper skipped the second as “already said”.
- Quiet voices: some recordings were at −62 dB, under the speech finder’s floor, and were never transcribed.
- A crash: whisper.cpp’s word timing stopped the program when a window moved on by only a few frames.

With the volume levelled per window and no text carried over from the previous window, word errors in English fell from 93.4% to 5.4%, in Greek from 29.0% to 21.7%, and the loops from 36 to none. Two-track recordings, such as calls, keep their tracks without levelling, because there the other side’s echo would get louder. The crash was fixed with a single change in the library. Laya’s data was then rebuilt with the fixed engine, so it learns from the mistakes users see today.
How to use it
The files are on Hugging Face and work with ONNX Runtime in any language:
model.onnx(343 MB): inputsinput_ids,attention_mask,marker_pos,marker_mask,qtype(2), outputlogits. One sentence at a time.tokenizer.json: mmBERT’s tokenizer, unchanged.laya.json: the question, the token budgets, the special tokens and the temperature.
Build the state as above, run the model and apply the temperature to its output: that is the probability that the word is a mistake. It suits anything that turns speech into text, from subtitles to notes, where you want to show people what is worth checking.
What comes next
The model isn’t finished. The next decisions of the same kind that could share it, each with its own data:
- whether two points in the minutes say the same thing and can be merged,
- whether a date, a number or an owner in the minutes is backed by the lines of the transcript,
- whether subtitles next to an imported file belong to it.
When something changes, this page will be updated.
Licences and sources
- somniusx/recmeets-laya-words: the model, Apache 2.0.
- convaiinnovations/laya: the base model, Apache 2.0.
- google/fleurs: the speech data, CC BY 4.0.
- LibreOffice’s hunspell dictionaries for Greek and English.
- RecMeets: the app it was made for.
Published on 29 September 2026.