Desert Ant Labs

Tongue

On-device language identification for short text across 84 languages.

Tongue model page
Platforms
iOS, macOS, tvOS, visionOS, Android, Linux, Windows, Browser, Node
Languages
84
Weights
Bundled with the SDK (v1.0.0)

Install

requirements

Swift
.package(url: "https://github.com/Desert-Ant-Labs/desert-ant-core.git", from: "3.5.0")

Then add the Tongue product to your target.

requirements

Kotlin
implementation("ai.desertant:tongue:3.5.0")

requirements

Terminal
npm i @desert-ant-labs/tongue

Usage

Nothing to download and nothing async: the 2 MB model ships inside the package, and a detection is pure arithmetic.

Swift
import Tongue

let tongue = try Tongue()                      // loads the bundled 2 MB model
let detection = tongue.detect("kann ich das haben")

detection.language          // "de"
detection.reliability       // .confident
detection.candidates        // [Prediction(language: "de", probability: 0.999…), …]
detection.isTooCloseToCall  // false

Tongue is a plain jar rather than an AAR, a pure Kotlin port with no native libraries, so it also runs on a bare JVM (17+).

Kotlin
import ai.desertant.tongue.Tongue

// Android: pass the Context. On a bare JVM pass null: Tongue.bundled(null).
val tongue = Tongue.bundled(context)
val detection = tongue.detect("kann ich das haben")
detection.language                               // "de"
detection.isTooCloseToCall                       // false

One import everywhere: no wasm, no LiteRT.js, no native core.

TypeScript
import { Tongue } from "@desert-ant-labs/tongue";

const tongue = await Tongue.load();                    // Node: reads the bundled model
const detection = tongue.detect("kann ich das haben");
detection.language;                                    // "de"
detection.isTooCloseToCall;                            // false

In a browser, serve tongue's two model files yourself (a bundler does not serve files out of node_modules) and pass from; both files are exported subpaths, so a copy script can require.resolve them under pnpm and Yarn PnP too:

TypeScript
const tongue = await Tongue.load({ from: "/models/tongue" });

Files

FileFormatSizeContents
tongue_int8.binRaw int8 + fp322.01 MiBThe shipped artifact. Byte-identical to what the live demo runs.
tongue_int4.binRaw int4 + fp321.01 MiBHalf-size alternative for tight bundles. Costs roughly 0.4pp on short text, see Sizes.
tongue.onnxONNX (fp32, opset 17)8.4 MBPortable graph for onnxruntime / onnxruntime-web.
tongue_meta.jsonJSONtinyRuntime tables: label order, hashing constants, script routing.
labels.jsonJSONtinyThe 59 model labels plus the script-decided languages, with English and native names.

There is no tokenizer file, so nothing has to be shipped or version-matched alongside the weights.

Inputs and outputs

Input: a short UTF-8 string, up to 512 characters. Output: ranked ISO 639-1/639-3 codes with probabilities, plus a reliability signal (confident / likely / tentative). Script-decided inputs return a single confident answer.

The ONNX graph carries only the head: values (int64 hashed bucket ids) and offsets (int64 per-sample starts) in, logits out. Normalization, hashing and script routing run in the host before the graph; tongue_meta.json documents them.

Coverage

59 languages are learned by the lexical model and a further 25 are decided by script alone, across 31 scripts (Latin, Cyrillic, Arabic, Greek and the CJK and Indic families among them) for 84 languages in total. One of the 84, Mongolian, is detected only in the traditional Mongolian script; see failure mode 3.

Sizes

Two quantisations of the same weights ship side by side. Pick on bundle budget, not on principle.

tongue_int8.bintongue_int4.bin
Size2.01 MiB1.01 MiB
FLORES 2-word0.8690.866
FLORES 5-word0.9740.973
Held-out single words0.7590.752
Held-out sentences0.9710.970

The int4 loss lands almost entirely on one- and two-word input; full sentences are unaffected within measurement noise. Since short text is what this model is for, int8 stays the default and int4 is the option when a megabyte matters more than the last half point.

Failure modes (read before deploying)

Publishing these is part of the product.

1. One or two words is often genuinely undecidable, and no model size fixes it. A single common word frequently belongs to several languages at once ("sale" is English, French and Italian; "la casa" is equally Italian and Spanish). tongue reports a tie or a tentative answer in these cases. It does not catch every one: a phrase mixing languages, like "un garage sale", can still draw a confident-looking single answer. Mitigation: treat low-reliability output as "unknown", not as an answer, and ask for more text where the product allows it.

2. Malay and Indonesian are not reliably separable. They share vocabulary and orthography to the point where short samples carry no distinguishing signal. This is a structural limit, not a tuning gap, and it is not cheaply closable at this size, and every detector we measured struggles with it. Mitigation: if you need the distinction, treat ms/id as one bucket or disambiguate from user locale.

3. Mongolian is detected only in the traditional Mongolian script. Mongolian written in Cyrillic, the dominant modern orthography, is not distinguished from the other Cyrillic languages and will usually come back as Russian. The language count includes Mongolian because the traditional script works; Cyrillic Mongolian does not. Mitigation: do not rely on tongue for Cyrillic Mongolian.

4. Brand names, numbers and code are not language. "Samsung Galaxy", "v1.2.3" and "2024 annual report" have no correct answer; the model will still return its best guess for anything with letters in it. Mitigation: filter non-prose input before detection.

5. Single-word scores are vocabulary recognition, not generalization. The frequent words of a language appear in everyone's training data, so any detector's single-word accuracy partly measures memorized vocabulary. Read the word-pair and sentence numbers as the generalization signal.

Measured quality

Every number below is measured on the shipped int8 weights, on three public benchmarks, with other detectors run on the identical rows and language subsets. Higher is better.

FLORES-200

Sentences from FLORES-200 truncated to their first 2, 3 and 5 words. Accuracy over the 20 languages the three detectors share.

DetectorSize2 words3 words5 words
tongue2 MB0.8690.9330.974
lingua293 MB0.8000.8870.956
eld~1 MB0.7800.8560.912

The lingua test set, the benchmark that library publishes

1,000 single words, word pairs and sentences per language, drawn from the same collection lingua trains on. Accuracy over the languages we share.

DetectorSizeSingle wordsWord pairsSentences
tongue2 MB0.7460.9090.988
lingua293 MB0.7520.9150.985

eld, an independent benchmark

Accuracy over the languages tongue supports (53,035 single-word rows, 53,613 word pairs, 53,141 sentences, 9,066 tweets). Apple is the built-in system detector; HeLI-OTS is a 51 MB JVM model. lingua 2.2.0 installs as a single 293 MB compiled extension with its language models embedded.

DetectorSizeTweetsSingle wordsWord pairsSentences
tongue2 MB0.9920.7590.8870.971
lingua293 MB0.9840.7560.8940.950
HeLI-OTS51 MB0.9860.6830.8430.967
Applesystem0.9970.6410.7190.748

Latency

Measured per single detection in JavaScript on an Apple-silicon laptop: 0.013 ms for one word, 0.028 ms for a short sentence, 0.10 ms at 193 characters (p99 0.24 ms). On-device budgets on phone-class hardware will be higher; the design target is under 1 ms.