Extract phoneme-level timestamps from speeh audio.
-
Updated
Jun 7, 2026 - Python
Extract phoneme-level timestamps from speeh audio.
pytorch model for contexless-phoneme prediction from speech audio
A multilingual phoneme recognizer capable of generalizing zero-shot to unseen phoneme inventories.
Speech Assessment API in FastAPI with HuggingFace 🤗
End-to-end IPA-based phoneme recognition pipeline using Wav2Vec2, featuring preprocessing, vocabulary generation, model training, and evaluation.
Self-hosted Docker speech API for OpenAI-compatible transcription and text-to-speech, streaming ASR, voice cloning, phoneme recognition, and MCP.
My bachelor thesis on Phoneme recognition and alignment on the TIMIT dataset
Official Python SDK for the Vocametrix voice analysis API — AVQI, DSI, jitter/shimmer, pronunciation assessment, speech-to-text, prosody similarity, and 40+ more clinical and acoustic measures.
Natural Language Processing Projects
音素级发音评测引擎。本地运行,毫秒响应,看见每一个声音。核心能力是将一段语音拆解到音素级别,告诉你哪个音读对了、哪个读错了、错在哪里、怎么改。
Record yourself speaking English, see which sounds you got wrong, and get coached on how to fix them.
AI-powered English pronunciation coaching: phoneme- and pitch-level speech analysis with interactive visual feedback. Vue 3 frontend, serverless AWS backend, and Wav2Vec2/SPICE ML model serving (PyTorch/TensorFlow, SageMaker/TorchServe).
Speech and speaker recognition — MFCC feature extraction, HMM alignment and concatenation, and phoneme recognition. (KTH DT2119)
2-People-Spanish-Average-Tone-Speech-Synthesis-Corpus
Context-aware neural decoding for speech BCI: extending DCoND with skip-diphone auxiliary supervision and temporal smoothness regularization on the Brain-to-Text '24 Benchmark.
To associate your repository with the phoneme-recognition topic, visit your repo's landing page and select "manage topics."