Desert Ant Labs

Uhm

On-device filler-word detection: frame-precise "uh"/"um"/"hmm" spans.

Uhm model page
Platforms
iOS, macOS, tvOS, visionOS
Languages
5
Weights
v1.1.0

Install

Swift (requirements)

Swift
.package(url: "https://github.com/Desert-Ant-Labs/desert-ant-core.git", from: "3.5.0")

Then add the Uhm product to your target.

Usage

Apple only today. Create one detector and reuse it: the model loads on first use, or earlier if you call download.

Swift

Swift
import Uhm

let uhm = Uhm()
let result = try await uhm.analyze(audioPath: "interview.m4a")

for filler in result.fillers {
    print(filler.start, filler.end, filler.confidence)   // seconds, seconds, 0...1
}
result.audioDuration
result.phaseTimings.inferenceSec                         // where the time went

Any format the platform decoder can read is accepted; audio is decoded to 16 kHz mono internally. analyze(audioURL:), analyze(bytes:) for in-memory audio, and analyze(samples:sampleRate:) for raw PCM are the other entry points.

Options trades recall against precision and drops spans that are too short. bias is the threshold preset: .precision (0.75) for automatic cuts, .balanced (0.65) by default, .recall (0.50) when you would rather review and confirm than miss one.

Swift
let options = Uhm.Options(bias: .precision, minDurationSec: 0.08)
let result = try await uhm.analyze(audioPath: "interview.m4a", options: options)

result.fillers.first?.type      // .uh, .um, .hmm, .and, .other

The type labeler is on by default and Apple-only; type stays nil elsewhere. Pass includeTypes: false to skip it when filler-vs-not spans are enough.

Pass a progressHandler to follow a long file, and cancel the enclosing task to stop the run:

Swift
let result = try await uhm.analyze(audioPath: path) { fraction in
    print("\(Int(fraction * 100))%")
}

Loading the model

The weights are fetched from the Hub on first use and cached. See model downloads and caching.

Files

FileFormatSizeUse
uhm.mlmodelc/Core ML fp16 (compiled)~45 MBiOS / macOS on-device
uhm-web-fp16.onnxONNX fp16~51 MBBrowser, server, Python (onnxruntime)
uhm.onnxONNX fp32~98 MBQuantization-free reference

uhm.mlmodelc/ is a compiled Core ML model directory. The Swift SDK downloads it with the Hugging Face Hub snapshot API, so only changed files are re-fetched on model updates.

The shipped model is a DistilHuBERT fine-tune. It is the smaller and more precise Uhm runtime model; the older HuBERT-base tier is no longer published.

Inputs and outputs

  • Input: 16 kHz mono audio, up to 30-second windows.
  • Output: per-frame softmax over 6 classes, one prediction every 20 ms.
  • Class indices: 0 = not_filler, 1 = uh, 2 = um, 3 = hmm, 4 = and, 5 = other.

Core ML input shape (30, 1, 1, 16080) float16 — the 30-second window pre-cut into 30 overlapping tiles — and output (1, 6, 1, 1499) float16. The SDK builds that layout for you; it exists because the Neural Engine caps every tensor axis at 16384, and it is what lets the whole model run there. Requires iOS 17 / macOS 14 or newer.

The ONNX artifacts keep the plain (1, 480000) float32 in, (1, 1499, 6) out shape: the tiled layout is an Apple-silicon optimization and is slower on a GPU.

Performance

Warm on-device runs on the published fp16 Core ML model:

DeviceRealtime factor
iPhone 17 Pro~296×
iPhone 15 Pro~169×
iPad Pro M4~279×

Realtime factor = audio duration ÷ analyze time; model load excluded.

Those are measured on the previous export. The current one runs every operation on the Neural Engine and is 1.6× faster where it has been measured (M1: 115× → 188× end to end), so these numbers are conservative until they are re-measured on the devices themselves.

Limitations

  • Trained on English; non-English performance is by acoustic transfer and has not been measured against per-language ground truth.
  • Best on podcast / meeting / talking-head audio. Heavy background music, laughter, or multi-speaker overlap degrades quality.
  • Type labels (uh / um / hmm / and / other) are secondary. Trust filler vs. not-filler more than the specific subtype.

Built on