Skip to content

Repository files navigation

Pocket TTS Android Engine

An unofficial, fully on-device Android text-to-speech engine for Kyutai Pocket TTS. It wraps the PocketTTS.cpp ONNX runtime in JNI and exposes it through Android's TextToSpeechService, so applications such as Layla can use Pocket TTS as a native system voice.

Pocket TTS Android Engine artwork

Status: experimental community software. The current build targets ARM64 Android devices and has primarily been tested on Android 16.

Features

  • Offline, on-device synthesis after installation and model import
  • Android system TTS integration
  • Streaming PCM audio through the Android TTS callback
  • Importable FP32 or INT8 ONNX language packs
  • Multiple language packs and multiple voices per language
  • Voice cloning from a WAV reference selected in the app
  • Per-model temperature, LSD steps, and CPU-thread settings
  • Configurable silence after sentence boundaries (250 ms by default)
  • Configurable maximum text tokens per generated segment (50 by default)
  • Token-based generation length guard matching the upstream Python runtime
  • Safe removal of installed language packs after confirmation
  • Automatic model selection by requested language, with English fallback
  • Stable per-voice identifiers for TTS clients such as Layla
  • In-app project information and a direct link to this repository

The Android application ID is org.pockettts.android.engine.

Official APKs published by this repository use the following release-certificate SHA-256 fingerprint: 5B:5C:DF:D1:5C:B0:4C:3D:77:7A:DD:21:E2:91:5F:55:AE:2B:7E:70:49:3D:8F:ED:07:9F:1A:C6:96:09:5B:78. Locally built debug APKs intentionally have a different certificate.

Recommended starting values from an Android ARM64 benchmark (108 syntheses, 78 additionally checked with speech recognition):

Model pack Temperature LSD steps Threads Sentence pause Segment size
German FP32 0.5 1 4 250 ms 50 tokens
German 24L FP32 0.3 1 3 250 ms 50 tokens

These are defaults for newly imported release packs, not hard limits. Existing per-pack settings are preserved when the app is upgraded.

The 24-layer packs are experimental. They require considerably more memory and can show variable latency on Android even when the same text and settings are used repeatedly. Test them on the target device before relying on them in an accessibility or interactive workflow.

Install and use

  1. Install the APK from the repository's Releases page.
  2. Download or build a compatible model-pack ZIP.
  3. Open Pocket TTS Android Engine and import the model pack.
  4. Select the official bundled voice, or import a WAV reference that you have permission to use. A clear 5-15 second mono sample usually works well.
  5. Save the selected voice and generation parameters.
  6. Open Android Settings > Text-to-speech output > Preferred engine and select Pocket TTS Android Engine.
  7. Restart the client application if it caches Android voices.

Only import model packs from trusted sources. ONNX graphs are processed by a native inference runtime; archive validation and size limits reduce accidental damage but cannot make a malicious model safe.

For Layla-specific details, see docs/LAYLA.md.

Voice-cloning sample quality

The reference recording is part of the synthesis input, not merely a speaker identifier. Pocket TTS explicitly notes that characteristics and defects in the sample are reproduced, so prompt quality can have a larger audible effect than small changes to the generation parameters. Use a clean, dry recording with natural speech and consistent volume; avoid background noise, music, echo, reverberation, clipping, aggressive noise reduction, and long leading or trailing silence. Testing several 5-15 second excerpts is worthwhile.

In one local German comparison, a clean excerpt from the Thorsten-Voice project produced fewer artifacts than the bundled Jürgen reference. This is an observation about the tested prompt recordings, not a universal ranking of the speakers or models. Thorsten-Voice also provides German models for Piper TTS; Piper itself is not used by this Android engine. See the Pocket TTS voice-cloning guidance and always verify the source and license of any recording before importing or redistributing it.

Build the Android app

The repository intentionally does not commit generated ONNX Runtime binaries, model weights, reference voices, APKs, or local SDK paths.

Requirements:

  • JDK 17
  • Android SDK 35
  • Android NDK 27.2.12479018
  • CMake 3.22.1
  • Git, curl, unzip, and tar
  • An ARM64 Android device for runtime testing
git clone <repository-url>
cd PocketTTS-Android-Engine
./scripts/prepare_android_native_deps.sh
./gradlew :app:assembleDebug

The debug APK is written to app/build/outputs/apk/debug/app-debug.apk. See docs/BUILDING.md for Android Studio, command-line, and device-testing and release-signing instructions.

Build a model pack

Model conversion requires Python 3.12, uv, FFmpeg, and access to Kyutai's gated model repository. Review and accept the upstream model conditions before running the exporter.

huggingface-cli login
./scripts/export_model_pack.sh german de-DE "German (FP32)" fp32

The resulting ZIP is written under build/model-packs/. Release packs may bundle the corresponding official Pocket TTS default voice under CC BY 4.0; see VOICE_ATTRIBUTION.md. Custom builds may omit voices or include only recordings the distributor is authorized to redistribute.

Supported upstream configurations include english, german, italian, spanish, portuguese, and the larger *_24l variants exposed by the pinned Pocket TTS source snapshot. Larger models generally improve quality but consume more storage, memory, and generation time.

See docs/MODEL_PACKS.md for the pack format, manual packaging, INT8 notes, and adding voices.

Repository layout

app/                         Android UI, TTS service, JNI bridge, and resources
scripts/                     Native dependency and model-pack build tools
vendor/PocketTTS.cpp/        Patched C++ runtime and ONNX exporter
vendor/pocket-tts/           Pinned upstream Python source snapshot
docs/                        Build, model, and Layla integration guides
MODEL_LICENSE.md             Model attribution and use conditions
VOICE_ATTRIBUTION.md         Official bundled-voice sources and licenses
THIRD_PARTY_NOTICES.md       Dependency provenance and licenses

Privacy and responsible use

Inference is local. Imported model packs and voice samples are stored in the app's private Android data directory. Do not use a person's voice without their explicit and lawful consent, and do not use synthesized speech to deceive or harm people. Review MODEL_LICENSE.md before distributing converted weights.

Acknowledgements

Development of this project was assisted by OpenAI Codex. The project owner reviewed the implementation, performed on-device testing, and is responsible for the published releases. This project is not affiliated with or endorsed by OpenAI.

Licensing

Original code in this repository is released under the MIT License. Third-party code retains its original license; see THIRD_PARTY_NOTICES.md. Pocket TTS model weights and converted ONNX derivatives are licensed separately; see MODEL_LICENSE.md.

Contributions are welcome. See CONTRIBUTING.md.

About

Unofficial offline Pocket TTS engine for Android system TTS, with importable multilingual ONNX model packs and voice cloning.

Topics

Resources

Contributing

Security policy

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages