lefteros_by SomniusX
EL ↗Say hello

LEFTEROS / PERSONAL ARCHIVE

RecMeets' Laya: which word was misheard

29 September 2026 / Lefteris Iliadis / SomniusX

somniusx/recmeets-laya-words is a model I trained and published on Hugging Face. It answers one question: given a word in its sentence, as a speech recognizer wrote it, how likely is it a mistake? It works in Greek and English, fits in 343 MB and runs on the processor, with no graphics card and no cloud. It is open (Apache 2.0), so anyone who needs it can use it in their own projects, and I keep improving it.

The somniusx/recmeets-laya-words page on Hugging Face, with its tags, description and resultsIt is a fine-tune of Laya (laya-multilingual, Apache 2.0), a multilingual decision model. I built it for Check words in RecMeets, the step that shows which words of a transcript were probably misheard, so you can listen and fix them. Its Hugging Face page holds the description, the data, the results and the three files you need: the model in ONNX, the tokenizer and a small settings file.

Why it was needed

A speech recognizer such as Whisper gives every word a “confidence”. In Greek, though, that alone barely separates the mistakes: on the words the engine is unsure of, just 0.63 on a scale where 0.5 is chance and 1.0 is perfect. A spelling dictionary helps, but it misses words that exist yet are wrong in their place, such as «είμαι» (“I am”) for «είναι» (“is”).

RecMeets' Check words: on the left the words to check with the dictionary’s and Laya’s view, on the right the word’s place with its audioSo it needed something that reads the word in its sentence, knows what the engine and the dictionary said, and gives an honest probability. It had to run on the user’s own computer, quickly, and send nothing anywhere, because meetings are private.

There is a ready-made option in the cloud, TypeSafe’s Jev, with a key and per-use pricing. The goal was a local model that comes close to it.

Laya in brief

Laya doesn’t write text like a chatbot. You give it a state (text or JSON) and typed questions (yes/no, choice, score), and in a single pass it answers with calibrated probabilities, in more than 100 languages. It was trained with reinforcement learning against strictly proper scoring rules (RLCD), so the only way for it to score well is to report honest probabilities. Because it generates no text, it has nothing to make up.

Its makers say it plainly: the base checkpoints are near chance without training on the task, and the value is in fine-tuning. I confirmed it: on twelve Greek words, six right and six misheard, no wording of the question separated them. On the test set, stock laya-multilingual scored 0.41, below chance.

The data: real mistakes, not guesses

The model learns from real mistakes, the ones made by the very engine RecMeets users run:

  • FLEURS by Google (CC BY 4.0): sentences read aloud, with exact reference text, in Greek and English.
  • I transcribed them with RecMeets itself (whisper.cpp, the large-v3-turbo model, with RecMeets’ settings), joining the sentences into 10-minute files with 1.2 seconds of silence between them.
  • Every word the engine wrote was labelled right or wrong by aligning it with the reference text.
  • The dictionary’s answer for every word came from LibreOffice’s hunspell dictionaries (el_GR, en_US).
  • Training: 71,906 Greek and 54,845 English words (9,880 and 2,547 mistakes), balanced so that 35% are mistakes.
  • Testing: FLEURS’ test split, with speakers the model never heard in training.

The question I ask it

For every word, Laya gets a yes/no question (noul) and this state:

Sentence: <about ten words either side, the word marked ⟦ ⟧>
Word: <the word>
Engine confidence: <0.00–1.00, or unknown>
In the <Greek|English> dictionary: <yes|no|can't tell>

It answers with the probability that the word is a mistake, corrected by a temperature measured on a separate slice of the data. It is plain text on purpose, so that RecMeets, written in Rust, builds it byte for byte the same as Laya’s Python.

Training

I followed Laya’s own recipe: RLCD (policy gradient with rewards from proper scoring rules) plus soft cross-entropy, a slice of the data held out for calibration, and a temperature per question type at the end. Everything ran locally, in a container with PyTorch and CUDA, with the processor use capped so the computer stays usable. One epoch takes about 8 minutes.

Run Data, epochs Greek, all words Greek, unsure English, unsure
1 Greek, 3 (21 minutes) 0.85 0.78
2 Greek, 3, with the dictionary in the input 0.90 0.89
3 Greek + English, 4 0.88
4 (published) Greek + English, 2 0.90 0.91 0.86

Run 3 overfit: in its fourth epoch the loss went below zero. Run 4, with two epochs, is the one I published.

The computer

Graphics card NVIDIA GeForce RTX 4080 SUPER, 16 GB
Processor AMD Ryzen 9 5950X, 16 cores / 32 threads
Memory 64 GB
Storage Samsung 990 PRO NVMe, 2 TB and 1 TB
Motherboard Gigabyte X570 AORUS ELITE
System Omarchy (Arch Linux with Hyprland)

The graphics card was needed only for training. The finished model runs on the processor.

Results

AUROC chart: the engine's confidence, the dictionary, this model and TypeSafe Jev, in Greek and English

AUROC on the FLEURS test split: 1.0 means every mistake scores above every right word, 0.5 is chance. “Unsure” are the words with an engine confidence under 0.5 and at least four letters, the ones Check words lists.

Greek, all Greek, unsure English, unsure
The engine’s confidence alone 0.82 0.63 0.74
The dictionary alone 0.85 0.70
Stock laya-multilingual, same question 0.41
This model 0.90 0.91 0.86
TypeSafe Jev (cloud), same input 0.93 0.89

On the words that matter, the local model comes within a whisker of the cloud’s Jev: free, on your computer, with nothing leaving it.

Size and speed

The model's files on Hugging Face: model.onnx 343 MB, tokenizer.json, laya.json and READMEExported to ONNX, the model is 1.29 GB at 32 bits, mostly mmBERT’s embedding table with its 256,000 tokens. There was a trap there. ONNX Runtime’s usual 8-bit quantizer with per-channel scales, the one Laya’s script uses, broke the fine-tuned model: it gave every sentence the same answer. With per-tensor scales it lost about 0.03.

The fix was ONNX Runtime’s 8-bit block format (MatMulNBits, blocks of 64) for the matrix products, and 8 bits for the embedding table: 343 MB, as accurate as 32 bits. With 8-bit arithmetic on the processor (accuracy_level 4), a sentence takes about 110 ms on 4 threads, instead of 140.

Size and speed chart for 32-bit, 8-bit, 8-bit with 8-bit arithmetic, and 4-bit

Packaging Size ms per sentence AUROC Greek / English
32-bit 1.29 GB 79 0.905 / 0.859
8-bit blocks 343 MB 140 0.905 / 0.861
8-bit blocks, 8-bit arithmetic (published) 343 MB 109 0.905 / 0.860
4-bit blocks 290 MB 77–80 0.889 / 0.854

The 32-bit model is faster, but nearly four times the download. The 4-bit one is fast too, but loses accuracy. I kept the sweet spot.

Inside RecMeets

RecMeets' settings: a second opinion in Check words, Laya on this computer or Jev in the cloudThe model is built into RecMeets, with no Python and nothing to install. It downloads once, 378 MB with the tokenizer, pinned to a specific revision and checked by SHA-256. RecMeets builds the input exactly as Laya’s Python does, and a test compares it token for token. It runs through ONNX Runtime on 4 processor threads, so the rest of the computer stays free.

To save time, when a word appears in two places, the second is checked only if the first answer is unsure (between 0.3 and 0.7). On a real reading, 146 words were checked in 23.4 seconds. Users can also choose Jev in the cloud if they have a key.

What got fixed along the way

To build the data, RecMeets transcribed hours of read sentences, and that showed four problems in RecMeets itself:

  • Loops: the same sentence written 40 times in a row, where 30 different ones had been read. Whisper continues each 30-second window from the text of the one before, and once a window repeats, the next carries the repetition on.
  • Skipped sentences: when two people read the same sentence one after the other, Whisper skipped the second as “already said”.
  • Quiet voices: some recordings were at −62 dB, under the speech finder’s floor, and were never transcribed.
  • A crash: whisper.cpp’s word timing stopped the program when a window moved on by only a few frames.

Chart: word errors before and after the fixes, English from 93.4% to 5.4%, Greek from 29.0% to 21.7%, loops from 36 to 0

With the volume levelled per window and no text carried over from the previous window, word errors in English fell from 93.4% to 5.4%, in Greek from 29.0% to 21.7%, and the loops from 36 to none. Two-track recordings, such as calls, keep their tracks without levelling, because there the other side’s echo would get louder. The crash was fixed with a single change in the library. Laya’s data was then rebuilt with the fixed engine, so it learns from the mistakes users see today.

How to use it

The files are on Hugging Face and work with ONNX Runtime in any language:

  • model.onnx (343 MB): inputs input_ids, attention_mask, marker_pos, marker_mask, qtype (2), output logits. One sentence at a time.
  • tokenizer.json: mmBERT’s tokenizer, unchanged.
  • laya.json: the question, the token budgets, the special tokens and the temperature.

Build the state as above, run the model and apply the temperature to its output: that is the probability that the word is a mistake. It suits anything that turns speech into text, from subtitles to notes, where you want to show people what is worth checking.

What comes next

The model isn’t finished. The next decisions of the same kind that could share it, each with its own data:

  • whether two points in the minutes say the same thing and can be merged,
  • whether a date, a number or an owner in the minutes is backed by the lines of the transcript,
  • whether subtitles next to an imported file belong to it.

When something changes, this page will be updated.

Licences and sources

Published on 29 September 2026.