RuneSight: the neural engine that tells text from code
Version 1.0.6 retires the hand-written regex list that decided which extracted lines were text and which were paths, identifiers and control codes. RuneSight, a 400 KB neural classifier trained on nine hundred thousand lines from real games, real sentences in sixteen languages and every documented engine control code, now makes that call in microseconds on your CPU, names its reason on every row, and never touches a line that carries real text.
Every game RuneTranslate opens is mostly text a player will read, plus a long tail of things that only look like text: asset paths, variable names, control codes, plugin keys, debug labels, a base64 blob someone left in a data file. Translate one of those and the best case is a wasted request. The worst case is a scene that never loads because the engine now looks up a name that no longer exists.
Since the first release, the line between the two was drawn by hand: hundreds of regular expressions, one per shape of junk we had met, engine by engine. Version 1.0.6 replaces that list with RuneSight, a small neural network trained on the real thing. It decides, for every extracted line, whether the line is text or plumbing, and it does so in a few microseconds, on your CPU, inside the app.
Why a list of rules stopped being enough
A regular expression catches exactly the shape it was written for. bg_office_01 is an identifier, so a rule for snake_case catches it. Then a Wolf RPG game arrives with <<GET_FILE_EXIST>>Data/\s[1], a Bakin game with Lv.\partystatus[0], a Unity game with "CubismTaskQueue.OnTask" already set., and a folder name written in Japanese breaks the path rule that assumed ASCII. Each one became a new rule, scoped carefully so it would not swallow real dialogue, and each new engine started the cycle again.
Rules also cannot weigh evidence. A line that is mostly a control code with three words of dialogue at the end is text, and must stay translatable. A line that is a full sentence about a .png file is text. A line that is nothing but \f[\cself[1]]\space[4] is not. Telling those apart needs a sense of the whole line, not a match on part of it.
What RuneSight is
RuneSight is a classifier of about 400 kilobytes. It reads a line as a bag of character fragments, one to four characters long, hashed into a fixed table, and pairs that with a few dozen measured properties of the line: how much of it is letters, digits, punctuation or spaces, which scripts it uses, whether it looks like a path, whether it has balanced brackets, whether it ends the way a sentence ends. One small hidden layer turns that into two answers: the probability that the line is not text, and, when it is not, which kind of not-text it is.
That second answer is what you see in the editor. An excluded row still carries a reason you can read: code pattern, asset reference, system label, numeric, comment, plugin key, control codes only, binary or plugin command. The row gets a small RuneSight mark and the confidence behind the verdict, so a decision is never a black box, and one click on the row turns it back on.
The whole thing runs in about four microseconds per line. A hundred-thousand-line JRPG is classified in well under a second, on any Windows 10 or 11 machine, with no GPU and nothing sent anywhere.
What it learned from
There was no labelled data for this, so it was built. The training set is a little over nine hundred thousand lines:
- Real games. Full extractions of thirteen games across ten engines, from RPG Maker and Ren’Py to Wolf RPG, Kirikiri, Artemis, Unity, Godot, Bakin and Unreal, labelled by the rules RuneSight was replacing. That is the floor: it had to catch everything the rules caught.
- Real sentences. A quarter of a million sentences in sixteen languages, Japanese, Chinese, Korean, Thai, Arabic, Russian and the Latin-script languages among them, so that ordinary text in any source language is never mistaken for code.
- Real plumbing. A hundred and sixty thousand paths, identifiers, dotted names and comment lines harvested from source code and packages, the shapes engines share.
- The blind spots, on purpose. Almost three hundred thousand generated lines built from the official control-code references of Wolf RPG, RPG Maker and Bakin, each paired with the near-miss dialogue it must not swallow: a sentence that mentions a file, a name followed by a number, a line of text with a code in the middle.
- A hand-checked gold set. Every line the old rules were pinned to, plus real lines reported from played games, adjudicated one by one.
What it had to prove before shipping
A model that is merely better on average would still be a regression for whoever hits its new mistakes, so RuneSight is held to fixed gates, re-run on every retrain:
- On every real game in the set, it catches at least 99.5% of what the hand-written rules caught.
- On the frozen set of lines that must stay translatable, the ones the old rules have always been tested against, it makes zero exclusions.
- On real prose, fewer than 0.2% of lines are ever excluded.
- The reason it names is right at least 97% of the time.
- A line is scored in under 5 microseconds.
The rule that mattered most is the one about mixed lines. A line that contains real text alongside codes is text, and stays translatable. RuneSight only removes a line when the whole line is plumbing.
What did not change
RuneSight decides the shape of a line. Decisions that need context, not shape, still come from the engine adapters: a Godot node name is excluded because of where it sits in the scene tree, an RPG Maker event comment because of the command that carries it, a plugin parameter key because of the file it lives in. Those rules stay, and so do the policy ones you can see in the reason column, such as a line that has no source-language characters at all.
Your own regex filters run after RuneSight and behave exactly as before, and every exclusion, whoever made it, is still a red opt-in row you can switch back on. Nothing RuneSight decides is final.
Where to see it
The RuneSight badge sits at the bottom right of the app, above the version line. Its dot is green when the model is loaded. Open it and it tells you the model size, the confidence threshold it ships with, and what the last extraction did: how many lines it scored, how many it excluded, and where it disagreed with the old rules. After every extraction a short toast repeats those numbers.
If the model ever cannot load, the badge says so, and extraction simply falls back to the previous rules. The app never refuses to work because of it.
Every place the model disagrees with the old rules is written to a small local log and never uploaded. If a line you consider real dialogue shows up excluded, turn it back on in the editor and tell us on Discord. Those reports are exactly what the next retrain learns from.
Ready to try RuneTranslate?
Free tier unlocks every engine + every translation provider. Supporter ($3.99/mo) unlocks full speed.
Download for Windows
