Skip to content

xerktech/Tenir

Repository files navigation

Tenir

Tenir — /teh-NEER/ — from the Latin tenēre, "to hold, to keep": it holds onto everything worth remembering.

A self-hosted speech-to-text recorder for the Even G2 glasses, a browser, and Android: live captions from a self-hosted STT model, with every session recorded and stored — transcript + full audio — and browsable afterwards.

That's the whole product, on purpose. It is the bare-minimum core the previous, much larger feature set (speaker identity, RAG cues, translation, summaries, chat, …) was stripped back to; features return one at a time, slowly.

What it does

  • Streaming transcription — phone, browser or glasses mic → real-time captions from self-hosted STT: a cache-aware streaming model (NVIDIA Nemotron) drives the low-latency live partials while an offline model (NVIDIA Parakeet) decodes each finished turn for the accurate stored transcript (the hybrid path — see docs/stt-model-gpu-benchmark.md)
  • Recorded, stored sessions — every session is persisted as a conversation: transcript segments in Postgres, full audio retained on disk; browse, search, replay, export and delete from the UI
  • One app container — the api serves the WebSocket capture endpoint, the REST surface, and the built web UI on a single origin (:8080)
  • Multi-user auth — login issues a bearer token; every session and conversation is scoped to the user's household
  • Clients — the web UI, an Even G2 glasses app (live captions on the lens, with dedicated Session + History pages on the phone, either surface starting and stopping a session), and an Android app (phone-mic capture + history)

Quick start

# Backend (single-host stack — app + Postgres + LiteLLM + Parakeet STT)
cp .env.example .env           # set API_AUTH_SECRET + bootstrap admin
docker compose up --build      # app (api + web UI) on :8080
curl localhost:8080/health

# Frontends (install once from repo root)
npm install

# Web UI in dev (the built UI is already served by the app container)
VITE_API_HTTP=http://localhost:8080 npm run dev --workspace @tenir/web

# Even G2 glasses app
VITE_API_WS=ws://localhost:8080/ws npm run dev --workspace tenir-even   # :5173
npx @evenrealities/evenhub-simulator -g http://localhost:5173

# Android app
npm run typecheck --workspace tenir-mobile
npm run test --workspace tenir-mobile

Deployment

One compose file at the repo root. The GPU host needs the NVIDIA Container Toolkit for the STT servers; everything else is CPU-only.

Service Port Role
app 8080 ONE container: FastAPI api (/health, /ws, auth + history REST) and the built web UI, served same-origin — no CORS, no separate web container. Built from api/Dockerfile (repo-root context)
postgres-tenir 5432 plain Postgres — transcripts, users (schema.sql applied on first boot)
litellm 4000 LiteLLM gateway — the OpenAI-compatible front door the api uses for STT finals + cues; master-key auth, routing in litellm/config.yaml
parakeet 9401 OpenAI-compatible NVIDIA Parakeet STT — offline model for accurate finals (multilingual + auto language detection), built from parakeet-stt/Dockerfile
nemotron 9402 NVIDIA Nemotron cache-aware streaming STT — WebSocket /v1/audio/stream for low-latency partials, built from nemotron-stt/Dockerfile (hybrid backend only)

Retained audio lives on a bind mount (API_AUDIO_DIR, the "disk" audio backend). Smoke check once up: curl localhost:8080/health, then open http://localhost:8080.

Hybrid STT (partials vs finals)

Live captions come from two models (API_STT_BACKEND=hybrid, the default):

  • Partials stream from nemotron over a persistent WebSocket (API_STT_STREAM_ENDPOINT). It's cache-aware, so each chunk costs ~one chunk of compute regardless of turn length — the live band stays responsive.
  • Finals decode the whole finished turn on parakeet (API_STT_ENDPOINT, else the gateway) for the accurate stored transcript.

If nemotron is unreachable the api degrades to Parakeet-only re-decode partials rather than losing the live band. Set API_STT_BACKEND=voxtral for the single-model (Parakeet-only) path — then the nemotron service isn't needed. Rationale and the GPU benchmark behind the split: docs/stt-model-gpu-benchmark.md.

The LiteLLM gateway

Finals and cues reach their model through a single LiteLLM gateway: one base URL + one key (API_LITELLM_ENDPOINT, API_LITELLM_API_KEY) instead of a per-model endpoint. Routing lives in litellm/config.yaml: the alias the api sends (API_STT_MODEL, default parakeet) fans out to the real model the parakeet server serves. The Nemotron partials stream bypasses the gateway (WebSocket, direct). To split hosts, run the gateway + model servers on the GPU box and point the endpoints at it — no code changes, just env.

Configuration

Secrets come from .env (see .env.example); non-secret config is baked into compose. Key envs on the app container:

API_AUTH_SECRET        bearer-token signing secret (boot refuses the default)
API_AUTH_ADMIN_*       bootstrap admin (username / password / household)
API_LITELLM_ENDPOINT   OpenAI-compatible base URL for STT finals + cues (…/v1)
API_LITELLM_API_KEY    gateway key
API_STT_BACKEND        hybrid (prod) | voxtral (Parakeet-only) | stub (dev/CI)
API_STT_STREAM_ENDPOINT  WebSocket root of the Nemotron partials server (hybrid)
API_PERSISTENCE_BACKEND  postgres | memory | off
API_AUDIO_BACKEND      disk | memory | off      (+ API_AUDIO_DIR)

See docs/contributing.md for repo layout, testing and the contract workflow.

About

No description, website, or topics provided.

Resources

Contributing

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages