Skip to content
View fauxneticien's full-sized avatar

Block or report fauxneticien

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
fauxneticien/README.md

Nay San

Staff Engineer at rime, working on data and modelling for conversational voice AI. Previously a PhD in Linguistics at Stanford (advised by Dan Jurafsky), on improving access to untranscribed speech corpora with AI.

More at resunay.com · Google Scholar


Speech: search and low-resource ASR

  • qbe-std_feats_eval — evaluation of feature-extraction methods for query-by-example spoken term detection in low-resource languages
  • bnf_cnn_qbe-std — query-by-example spoken term detection using bottleneck features and a CNN
  • u2u-asr — "user-to-user" ASR: a proof-of-concept workflow taking users from their own data to a fine-tuned model they can run locally in the browser (a play on "end-to-end ASR")
  • active_learning-w2v2_asr — active learning for fine-tuning wav2vec 2.0 ASR
  • asr-dataset-prep — scripts for preparing datasets for automatic speech recognition

Cross-lingual and self-supervised speech models

Phonetics and phonology

  • phonpack — an R package of fun(ctions) for doing phonetics
  • kphon — helper functions for the Kaytetye Phonological project (KPHON)
  • kaytetye-medial-vowels — processing scripts and datasets for a study of medial vowels in Kaytetye
  • akwelye — text-setting in akwelye (Kaytetye song)
  • census-languages — analysis of ABS Census data on Australian Indigenous languages

Lexicography and dictionaries

  • lexloop — iterative correction tool for data in a domain-specific language: edit a file, re-run, see validation and parsed views in the browser
  • lexicon-grammars — a collection of grammars for parsing backslash-coded lexicons
  • LexDev — a toolkit for generating live feedback on lexicographical data
  • kdict — data-processing functions for the Kaytetye Dictionary Transcriptions project
  • anamR — helper functions to read/write/process data from the Kaytetye database (KDB)

Pinned Loading

  1. CoEDL/vad-sli-asr CoEDL/vad-sli-asr Public

    A pipeline to isolate and transcribe one language in mixed-language speech

    Python 20 3

  2. qbe-std_feats_eval qbe-std_feats_eval Public

    Evaluation of feature extraction methods for query-by-example spoken term detection with low resource languages

    Perl 12 2

  3. CoEDL/vyov CoEDL/vyov Public

    Visualise your own vowels: A short introduction to Praat for complete beginners

    HTML 2

  4. CoEDL/tidylex CoEDL/tidylex Public

    Tidy lexicographical data in backslash-coded formats

    JavaScript 5

  5. CoEDL/yinarlingi CoEDL/yinarlingi Public

    R package for testing Warlpiri dictionary data structures

    R 1