Desert Ant Labs

Voz

语音识别可用

设备端语音转文字:在 iPhone 上 2 秒转写 10 分钟音频,比 Whisper 快 14 倍,逐词时间戳精准。

2 秒转写 10 分钟音频。

用 Voz 转写一个音频或视频文件,为每个词拿到精准的起止时间。在 iPhone 上,10 分钟只要 2 秒,全程在设备上,什么都不上传,也不按分钟计费。

在 Apple 芯片上,模型运行在神经网络引擎上,因此转写快如闪电,也更省电。在 Mac 上,一整库播客、访谈和讲座一遍就能过完,什么都不上传。

时间戳精准到可以据此剪辑。词的起点落在强制对齐器的 83 毫秒以内,词的终点在 95 毫秒以内,足以做基于转写的精确剪辑。Voz 是 NVIDIA Parakeet TDT 0.6B v3 面向 Apple 神经网络引擎优化的版本,在大多数音频上准确度与 Whisper large-v3-turbo 相当,下载体积仅 467 MB。

比 Whisper 快 14 倍。

10 分钟旁白在 iPhone 17 Pro 上 2 秒、iPhone 18 Pro 上 1.2 秒、iPhone 15 Pro 上 2.1 秒。在 M3 Ultra 上,处理同一段 10 分钟播客比 whisper.cpp large-v3-turbo 快 14 倍,在六个公开英语集合上词错误率为 7.40%,Whisper 为 7.00%。

0.0s
0:00 / 10:00

音频:Ant Ventures,一段 LibriVox 录音,公有领域。

2s
10 分钟音频
iPhone 17 Pro,在神经网络引擎上
14x
比 Whisper 快
通过 whisper.cpp 的 large-v3-turbo,10 分钟,M3 Ultra
7.40%
词错误率
Open ASR Leaderboard 的六个集合;Whisper large-v3-turbo 为 7.00%
467MB
磁盘占用
Whisper large-v3-turbo:fp16 下 1.6 GB

转写 10 分钟音频所需的时间。

设备时间实时
M3 Ultra0.8s762x
iPhone 18 Pro1.2s492x
M4 Max1.3s453x
iPhone 17 Pro1.8s334x
iPhone 15 Pro2.1s283x
M1 Mac mini2.4s251x

Open ASR Leaderboard 数据集上的词错误率,Voz 使用自己的文本归一化器,Whisper 的数据取自排行榜

数据集VozWhisper large-v3-turbo
LibriSpeech test-clean2.19%2.13%
LibriSpeech test-other3.86%3.70%
GigaSpeech9.70%8.47%
SPGISpeech3.86%2.79%
Earnings-2212.97%11.07%
AMI11.84%13.87%
Average7.40%7.00%

看与你音频相符的那一行。干净、近距离拾音的语音约落在 2% 到 4%(LibriSpeech、SPGISpeech),播客和网络视频约 10%(GigaSpeech),会议室和电话通话 12% 到 13%(AMI、Earnings-22),这正是模型胜过 Whisper 的地方。各语言各 10 分钟朗读语音的数据在模型卡上,从意大利语的 3.31% 到希腊语的 39.46%。iPhone 计时和 whisper.cpp 对比是我们自己的测试,尚未收录到模型卡里。

应用场景

2 秒转写 10 分钟音频。.

转写任意音频或视频文件

录一段播客或视频,在导出完成之前完整转写就已就绪。可搜索的文本,每个词都带时间戳,支持 25 种语言中的任意一种,全程不离开设备。

清完积压的存量

一批访谈、讲座或会议录音,在 Mac 上一遍就变得可搜索。一百小时 20 分钟过完,什么都不上传。

在词上剪辑

每个词都带起点和终点,因此在转写文本里选中的一段范围,就是视频剪辑器里的一刀。把它和 Clips 搭配做短视频,再用 Title 给它们命名。

灵感

用 Voz 构建的点子。把提示词复制到你的编码 agent 里就能开始。

Voz

Build a podcast player that transcribes every episode and makes it searchable, on the device.

复制提示词
Voz

Build a meeting recorder that hands you a timestamped transcript before you leave the room.

复制提示词
Voz

Turn voice memos into searchable, editable text notes, offline.

复制提示词
Voz

Generate SRT subtitles for any video file on the Mac.

复制提示词
Voz

Add captions to a video player that run on the device.

复制提示词
Voz

Build an interview app for journalists where sources never leave the phone.

复制提示词
VozClipsTitle

Build a full podcast studio: transcribe, clip, and title on the device.

复制提示词
ClearVozClips

Record, clean, transcribe, and clip: a whole creator pipeline, offline.

复制提示词
查看全部 26 个点子

模型能做什么

  • 词的起点平均落在强制对齐器的 83 毫秒以内,词的终点在 95 毫秒以内,帧分辨率为 80 毫秒。
  • 10 分钟音频在 iPhone 17 Pro 上 2 秒、iPhone 18 Pro 上 1.2 秒、iPhone 15 Pro 上 2.1 秒。在 M3 Ultra 上 0.8 秒,762 倍实时。单条短片段跑 50 到 62 倍,因为每个片段都要为一个完整的 15 秒窗口买单。
  • 在 Open ASR Leaderboard 的六个集合上平均词错误率 7.40%(Whisper large-v3-turbo:7.00%),半小时旁白上 2.83%(Whisper:2.72%),AMI 会议上 11.84%(Whisper:13.87%)。
  • 整个计算图都在神经网络引擎上运行,不回退到 CPU 或 GPU;峰值内存不随录音时长增长。
  • 25 种语言:保加利亚语、克罗地亚语、捷克语、丹麦语、荷兰语、英语、爱沙尼亚语、芬兰语、法语、德语、希腊语、匈牙利语、意大利语、拉脱维亚语、立陶宛语、马耳他语、波兰语、葡萄牙语、罗马尼亚语、俄语、斯洛伐克语、斯洛文尼亚语、西班牙语、瑞典语和乌克兰语。
  • 磁盘上 467 MB,按需下载并缓存;下载后首次加载会有一次 20 秒,等 Core ML 做专门化,之后为 0.2 秒。请在引导流程中下载。
  • 任意采样率的单声道音频;SDK 会重采样并下混。较长的音频会在停顿处切成 15 秒的窗口,再在相邻窗口一致认同的词上拼接。
  • 基于 NVIDIA 的 Parakeet TDT 0.6B v3 构建,以 CC BY 4.0 发布。权重未作改动;转换、压缩和神经网络引擎运行时是我们做的。

示例应用

Clipper

使用 Voz, Clips, Title 构建

Generate short clips from a video podcast or a long recording, fully on device. A macOS app and a command-line tool over the same core.

Voz transcribes with a time on every word, Clips ranks the best spans, and Title writes a title and description for each.

macOS 26 or later, Apple Silicon

使用 Homebrew 安装
brew tap desert-ant-labs/tap
brew install --cask clipper

快速上手

只需几行代码,即可为你的 iOS or macOS or web 应用加上 语音识别。 Voz 文档.

iOS, macOS
安装
// Swift Package Manager
.package(url: "https://github.com/Desert-Ant-Labs/desert-ant-core.git", from: "3.5.0")
// target dependency
.product(name: "Voz", package: "desert-ant-core")
示例 - Swift
import Voz

let voz = try await Voz()
let result = try await voz.transcribe(url)
result.text                     // the transcript
result.words.first?.start       // 80ms resolution
result.realtimeFactor           // seconds of audio per second of wall clock
用提示词构建
Add Voz from Desert Ant Labs to this Swift project (iOS, macOS).

What it does: Voz:在神经网络引擎上运行的设备端语音转文字.

SDK:

Swift (iOS, macOS)
Repo: https://github.com/Desert-Ant-Labs/desert-ant-core#readme
// Swift Package Manager
.package(url: "https://github.com/Desert-Ant-Labs/desert-ant-core.git", from: "3.5.0")
// target dependency
.product(name: "Voz", package: "desert-ant-core")

Reference:
- Model page: https://desertant.com/models/voz/
- Full catalog and other models: https://desertant.com/llms.txt

Add the SDK, then follow its README for the exact API and current version. Do not invent API names or method signatures; confirm them against the README.
Android即将推出
Web, Node.js
安装
npm i @desert-ant-labs/voz onnxruntime-web
示例 - TypeScript
import { Voz } from "@desert-ant-labs/voz";

const voz = await Voz.load();
const result = await voz.transcribe(file);   // File, Blob, ArrayBuffer, or samples
result.text;
result.words[0];        // { text, start, end }
result.realtimeFactor;
用提示词构建
Add Voz from Desert Ant Labs to this JavaScript / TypeScript project (Web, Node.js).

What it does: Voz:在神经网络引擎上运行的设备端语音转文字.

SDK:

JavaScript / TypeScript (Web, Node.js)
Repo: https://github.com/Desert-Ant-Labs/desert-ant-core#readme
npm i @desert-ant-labs/voz onnxruntime-web

Reference:
- Model page: https://desertant.com/models/voz/
- Full catalog and other models: https://desertant.com/llms.txt

Add the SDK, then follow its README for the exact API and current version. Do not invent API names or method signatures; confirm them against the README.

规格

速度
10 分钟音频在 Mac 和 iPhone 上 0.8 到 2.4 秒,251 到 762 倍实时;各设备用时见性能
准确度
在 Open ASR Leaderboard 的六个集合上 WER 7.40%;长文本 2.83%;词起点 83 毫秒、终点 95 毫秒平均误差
设备端体积
467 MB 编译后 Core ML,按需下载
语言
25 种欧洲语言;朗读语音上准确度从 3.31%(意大利语)到 39.46%(希腊语)
模型
NVIDIA Parakeet TDT 0.6B v3(CC BY 4.0),由 Desert Ant Labs 转换为 Core ML 并压缩:对数梅尔前端、conformer 编码器、transducer 解码器,全部在神经网络引擎上
平台
iOS、iPadOS、macOS、tvOS、visionOS(Core ML);浏览器和 Node(WebAssembly、ONNX Runtime)

神经网络引擎的速度仅限 Apple:在 iOS、macOS、tvOS 和 visionOS 上,运行时驱动 Core ML,把整张计算图留在神经网络引擎上。浏览器和 Node 用 WebAssembly 和 ONNX Runtime 运行同一个模型,比 Core ML 慢,而且没有原生 Android 构建。Voz 不知道自己听到的是它 25 种语言中的哪一种,遇到它不覆盖的语言,会自信地返回一堆胡话而不是报错,所以请把它和 Ear 搭配使用。准确度因语言差异很大,词的结尾是时间戳里更难的那一半,467 MB 在引导流程中也是一次实打实的下载。

常见问题

Voz 是什么?

设备端语音转文字:在 iPhone 上 2 秒转写 10 分钟音频,比 Whisper 快 14 倍,逐词时间戳精准。

Voz 在设备上运行吗?

是的。Voz 在设备上运行,不调用任何服务器,因此数据始终留在用户手中。

Voz 支持哪些平台?

Voz 以面向 Swift, JavaScript / TypeScript 的原生设备端 SDK 形式提供。

Voz 的价格是多少?

每个模型免费支持最多 10 万台月活跃设备。推理次数不限。 如需定制授权,请联系我们。

Voz 的准确度和速度如何?

10 分钟旁白在 iPhone 17 Pro 上 2 秒、iPhone 18 Pro 上 1.2 秒、iPhone 15 Pro 上 2.1 秒。在 M3 Ultra 上,处理同一段 10 分钟播客比 whisper.cpp large-v3-turbo 快 14 倍,在六个公开英语集合上词错误率为 7.40%,Whisper 为 7.00%。

Voz 能在 Apple 平台之外运行吗?

可以,在浏览器和 Node 中都能运行。JavaScript 包用 WebAssembly 和 ONNX Runtime 运行同一个模型,因此网页应用或服务器无需 Apple 硬件也能转写,包括在 Windows、Linux 或 Android 浏览器上。在 Apple 芯片上,整个模型都跑在神经网络引擎上,那里速度最快。原生 Android 构建目前尚未就绪。

资源