--- library_name: coreai license: apache-2.0 language: - en - zh pipeline_tag: text-to-speech tags: - coreai - text-to-speech - tts - core-ai - on-device - ios - voxcpm base_model: openbmb/VoxCPM2 base_model_relation: quantized --- Core AI is Apple's on-device ML runtime in iOS 27 / macOS 27 and the successor to Core ML: PyTorch models are exported with Apple's `coreai-torch` (LLMs: `coreai.llm.export`) into `.aimodel` bundles that run on the GPU or the Neural Engine, e.g. Qwen3-8B 4-bit decodes at 94 tok/s on an M4 Max GPU, MLX 90 under the same protocol ([apple-silicon-llm-bench](https://github.com/john-rocky/apple-silicon-llm-bench), macOS 27 beta 26A5353q, 2026-06-11). This model has no row on [DeviceMark](https://devicemark.github.io/), the on-device LLM leaderboard. # VoxCPM2 2B — Core AI (on-device, 48 kHz) [OpenBMB **VoxCPM2** (2B)](https://huggingface.co/openbmb/VoxCPM2) converted to **Apple Core AI**, running fully **on-device** on iPhone and Mac — no network. The 2B, 48 kHz successor to [VoxCPM-0.5B-CoreAI](https://huggingface.co/mlboydaisuke/VoxCPM-0.5B-CoreAI). A tokenizer-free diffusion TTS: a **MiniCPM4 28-layer** text-semantic LM + an **8-layer residual** acoustic LM drive a **12-layer LocDiT** flow-matching diffusion head, decoded by a **48 kHz AudioVAE**. Five Core AI bundles + a few host-side projections. ## Use it **New to Core AI? [Start with CoreAIKit 0.7.3](https://github.com/john-rocky/coreai-kit#readme).** Follow its requirements and first-run steps for `qwen3-0.6b`, then open the same release's [ChatDemo](https://github.com/john-rocky/coreai-kit/tree/0.7.3/Examples/ChatDemo). The README records the tested OS/SDK and download size; model and device coverage is stated per example. ⚡ **One line** — run the kit's task op on this model (`import CoreAIOps`; no session, no model plumbing, downloads on first use): ```swift let audio = try await CoreAI.speak(text, options: .model("voxcpm2-2b")) ``` Every op, one shape — [Cookbook](https://github.com/john-rocky/coreai-kit/blob/0.7.3/docs/COOKBOOK.md). ▶️ **Run it (source)** — the [Speak runner](https://github.com/john-rocky/coreai-kit/tree/0.7.3/Examples/Speak) (GUI + CLI, one app for every text-to-speech model in the catalog): ```bash git clone --branch 0.7.3 --depth 1 https://github.com/john-rocky/coreai-kit export DEVELOPER_DIR=/Applications/Xcode-27.0.0-RC.app/Contents/Developer open -a /Applications/Xcode-27.0.0-RC.app coreai-kit/Examples/Speak/Speak.xcodeproj # → Run, then pick "VoxCPM2 2B" in the model picker # agents / headless (macOS): cd coreai-kit/Examples/Speak swift run -c release speak-cli --model voxcpm2-2b --text "Hello from Core AI." --output hello.wav ``` Use Xcode build **27A266a** from the release's `.xcode-pin`; adjust the app path if your installation is named differently. 💻 **Build with it** — complete; the glue is kit API, copy-paste runs: ```swift import CoreAIKit let speaker = try await KitSpeaker(catalog: "voxcpm2-2b") let audio = try await speaker.synthesize(text) // audio.samples: 48 kHz mono PCM in [-1, 1] — play it or write a WAV ``` The take-home is [`Examples/Speak/Sources/QuickStart.swift`](https://github.com/john-rocky/coreai-kit/blob/0.7.3/Examples/Speak/Sources/QuickStart.swift) — this exact code as one typed function, no UI; the CLI is an argument shell over it, and the GUI drives the same `KitSpeaker(catalog:)` and plays the samples. Live playback? `synthesizeStreaming(_:onChunk:)` hands you ~0.5 s chunks as they decode, so audio starts before the whole clip exists. The WAV container is your app's territory (the runner ships a 20-line writer). **Integration checklist** - SPM: `https://github.com/john-rocky/coreai-kit` (exact **0.7.3**) → product **CoreAIKit** - Info.plist: none needed - Entitlements: none needed - First run downloads the model — ~4,721 MB (Mac) / ~4,727 MB (iPhone) — then it loads from the local cache (Application Support; progress via the `downloadProgress` callback) - Measure in Release — Debug is ~3× slower on per-token host work ## What's inside | dir | contents | |---|---| | `macos/` | JIT `.aimodel` bundles (Mac): int8 base/res decode + prefill, fp16 feat_decoder / feat_encoder / vocoder | | `ios/` | the same JIT bundles as `macos/`, laid out the same way (`/.aimodel`); every iPhone generation specializes them on its first load | | `ios-h18p/` | AOT `.aimodelc` bundles (iOS `h18p`, GPU), iPhone 17 Pro only: same five + the two int8 prefill bundles; moved from `ios/` in revision `192cc99f` (2026-09-26) | | `voxcpm2_host_glue/` | embed table + projections / FSQ-512 / stop-head / fusion (`.bin` + manifest) | | `tokenizer/` | the VoxCPM2 tokenizer (Llama fast) | The backbone LMs are **weight-only int8** (the size driver); the diffusion + VAE stay **fp16** (the continuous-feedback path is quant-sensitive — same split mlx-community uses). ## On-device numbers (iPhone 17 Pro, h18p AOT bundles, int8 + prefill + streaming) - **RTF 1.19**, **first-audio 0.65 s**, 48 kHz, ~4.9 GB resident (increased-memory entitlement). - Streaming starts after the first ~0.65 s; the 2B is ~4× the 0.5B, so RTF sits just above realtime. ## iPhone 18 Pro, the JIT bundles in `ios/` (load only) Measured 2026-09-26 on an iPhone 18 Pro (iPhone19,2, iOS 27.0 build 24A437, h19p) with the zoo's DecideGate app in its load-only mode, without the increased-memory entitlement. Each first load was the first after a fresh install of the app. The call is one run on all-zero inputs; a graph with a KV state was loaded only. Each first load wrote a specialization of about the bundle's size into the app container, and the load after a relaunch reused it. One measurement per graph ([knowledge/jit-distribution.md](https://github.com/john-rocky/coreai-model-zoo/blob/main/knowledge/jit-distribution.md)). The same phone refuses an h18p bundle with `incompatibleCompiledAssetArchitecture`. | JIT bundle in `ios/` | MB | first load | first call | load after relaunch | |---|---:|---:|---:|---:| | `voxcpm2_vocoder_fp16_t8.aimodel` | 92 | 0.16 s | 685 ms | 0.06 s | | `voxcpm2_res_int8_decode_cl512.aimodel` | 378 | 0.79 s | — (KV state) | 0.31 s | | `voxcpm2_res_int8_prefill_t32.aimodel` | 378 | 0.73 s | — (KV state) | 0.26 s | | `voxcpm2_feat_encoder_fp16.aimodel` | 420 | 1.17 s | 654 ms | 0.30 s | | `voxcpm2_feat_decoder_fp16.aimodel` | 426 | 5.64 s | 1,739 ms | 0.36 s | | `voxcpm2_base_int8_prefill_t32.aimodel` | 1,324 | 3.49 s | — (KV state) | 1.25 s | | `voxcpm2_base_int8_decode_cl512.aimodel` | 1,324 | 3.20 s | — (KV state) | 0.97 s | ## Use it Runs through **[coreai-kit](https://github.com/john-rocky/coreai-kit)** `VoxCPM2TTS`, wired into the **[coreai-model-zoo](https://github.com/john-rocky/coreai-model-zoo)** `coreai-audio` app ("Voice 2B" tab). Conversion + gates + export scripts: `coreai-model-zoo/conversion/voxcpm/` (`*_v2.py`). ```swift let tts = try await VoxCPM2TTS(paths: .standard(artifactsRoot: root, lm: .int8)) let wav = try await tts.synthesize("On device speech synthesis, running entirely on your iPhone.") // 48 kHz Float PCM ``` ## Verification Reimplemented in exportable Core AI overlays and gated end-to-end against the official model: backbone / feat_decoder / feat_encoder **cos 1.0**, full chain **magspec 0.996**; every exported bundle engine-gated **cos ≥ 0.9999**. ## License Apache-2.0 (commercial OK), inherited from [openbmb/VoxCPM2](https://huggingface.co/openbmb/VoxCPM2). Not affiliated with OpenBMB or Apple. Community port. --- **More models in this format:** [Core AI Model Zoo](https://huggingface.co/collections/mlboydaisuke/core-ai-model-zoo-6a7ff330f753e8dcae04671a) — 75 models, each with the recipe that produced it. **Want a different model on-device?** [Open a request](https://github.com/john-rocky/on-device-requests) — free, open weights only; the export and its measured numbers get published publicly.