Desert Ant Labs

Clips

Short clips and highlights from talking video and audio: podcasts, interviews, meetings. On-device.

Clips model page
Platforms
iOS, macOS, tvOS, visionOS, Linux, Windows
Weights
v0.1.0

Install

Swift (requirements)

Swift
.package(url: "https://github.com/Desert-Ant-Labs/desert-ant-core.git", from: "3.5.0")

Then add the Clips product to your target.

Usage

Apple, Linux and Windows. Give it a transcript as one sentence per element, in spoken order, and it returns the best non-overlapping moments, ranked.

Swift

Swift
import Clips

let clips = Clips()
let moments = try await clips.clips(in: sentences)      // [Clip], best first

for clip in moments {
    print(clip.text, clip.start, clip.end, clip.score)
    print(clip.sentenceIDs)                             // which sentences it spans
    print(clip.estimatedDurationSec, clip.percentile)
}

A transcript under three sentences returns [].

The limit sizes the work, it does not trim the result. It sets the selection budget, and the candidate pool is four anchors wide per unit of budget, so a smaller limit means fewer scorer passes, which is 60-85% of the runtime:

Swift
let ten = try await clips.clips(in: sentences, limit: 10)   // default is 10
let auto = try await clips.clips(in: sentences, limit: nil) // model decides from duration

One consequence worth knowing: the result at limit: 10 is not generally the first ten of the result at limit: 14. Weighted interval scheduling returns the highest-total non-overlapping set of at most k, and the best set of ten is not the best set of fourteen with four removed.

Writing titles for the clips

Clip carries no title, deliberately. Pair it with Title, which takes a Clip directly:

Swift
let cards = try await titles.cards(for: moments)   // index-aligned with moments

Loading the model

The weights are fetched from the Hub on first use and cached. See model downloads and caching.

Maximum video length

There's no context window. Clips never reads a transcript whole: sentences run through the selector in batches of 16 and each candidate clip is scored on its own, so length is bounded by time rather than by a token limit. The longest we've run is an 835 sentence podcast.

transcriptiPhone 17 ProiPhone 15 Pro
404 sentences, 25 minutes of video, 12 clips9.19s10.22s
57 sentences2.23s
per candidate, encoder only2.78ms3.13ms

Measured on device at batch 16, pinned to .cpuAndNeuralEngine.

Files

FileFormatContents
clips.mlmodelc/Compiled Core ML, int8Multifunction package. Function select: ids, mask, disc → saliency, start_p, end_p. Function score: ids, mask → score
clips-selector.tfliteLiteRT, int8 weight-onlyThe selector, for Android, Linux and Windows
clips-scorer.tfliteLiteRT, int8 weight-onlyThe scorer, for Android, Linux and Windows
clip_tokenizer.binUnigram tokenizerTokenizer pieces and scores, in the compact binary the runtime reads
clips_meta.jsonJSONGraph widths, input roles and feature order a runtime needs

Reaching a function in the Core ML package

A file path names the package, not the graph. Loading needs MLModelConfiguration.functionName set to select or score. Without it Core ML loads the package's default function and reports nothing, so both halves of the pipeline end up being the selector.

Windows

The selector runs at 128 tokens, the scorer at 256, both at a fixed batch of 16 sentences.

A single sentence is truncated to 64 tokens before it reaches the selector, which is not the same thing as the graph width.

Status

Internal testing. This page carries no quality figures yet.

⚠️ The .tflite files do not work with the Desert Ant SDK yet. Do not build on them.

The SDK's LiteRT backend cannot drive them:

  • the graphs name their inputs args_0, args_1, args_2 and their outputs output_0…2; the SDK asks for ids, mask, disc and saliency, start_p, end_p. There is no mapping layer, so the first call fails.
  • the graphs take int64 ids and mask; the SDK builds int32.
  • the scorer is 256 wide and the SDK's LiteRT backend cannot report a width, so it falls back to 128.

They are left published because they are honest artifacts and someone driving LiteRT directly can use them: the shapes and output order are in clips_meta.json. They are not a working Android/Linux/Windows path today.

Also unresolved: this export was measured at ~2.1 GB peak RSS against a 1.6 GB Android budget.

The two platforms are not equally evidenced. The Core ML package has clips that were generated from it and judged. No clip has been read from the LiteRT files, on any platform.

Requirements

The Core ML package is specification version 9: it requires iOS 18 / macOS 15 / tvOS 18 / visionOS 2 / watchOS 11, read off the compiled artifact. Reaching either graph needs MLModelConfiguration.functionName set to select or score.

Limits

Behaviour worth knowing before you build on it, stated without figures for the reason above:

  • It under-emits on short video. Given a short transcript it returns markedly fewer clips than expected. If your product needs a guaranteed number of clips from a two-minute video, measure before relying on it.
  • It emits some dross, most on podcast-length input. There is no confidence score to filter on yet: Clip.score ranks within one video and is not calibrated across videos.
  • The clip limit is a cap, not a quota. Asking for 10 does not mean receiving 10.
  • Selection is sensitive to small score changes. Candidate spans around one moment score very close together, so a different runtime, compute unit or quantization can return a different-but-comparable set rather than the same set. Do not treat exact span equality between two builds as a correctness check.
  • Non-Latin scripts are under-tested. The evaluation corpus is overwhelmingly Latin-script.
  • Duration is a soft prior, not a rule. Clips may come back shorter or longer than a typical Short.

Example app

Clipper Generate short clips from a video podcast or a long recording, fully on device. A macOS app and a command-line tool over the same core.

Install with Homebrew
brew tap desert-ant-labs/tap
brew install --cask clipper

View sourceOr download for macOS