Skip to content

Latest commit

ย 

History

8 Commits

Folders and files

NameName
Last commit message
Last commit date
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 

Repository files navigation

๐ŸŽ™๏ธ EchoForge

AI-Powered Radio Drama Production Studio

Turn a story idea into a fully-produced, voice-acted radio drama โ€” automatically.

Version Platforms License: MIT TTS LLM

English ยท ็ฎ€ไฝ“ไธญๆ–‡ ยท ๆ—ฅๆœฌ่ชž


๐Ÿ“ธ Preview

Settings & Providers Music & SFX Library
Settings Library
Agent Workbench Live Progress
Agent Progress

โœจ What is EchoForge?

EchoForge is a local-first desktop studio that produces complete radio dramas (audio dramas) from a script or even just a story idea:

Story idea โ”€โ”€โ–ถ LLM writes the script โ”€โ”€โ–ถ TTS voices every character
                    โ”‚                            โ”‚
                    โ”œโ”€โ”€ ๐ŸŽต Adaptive BGM (intro โž ducking under speech โž outro)
                    โ”œโ”€โ”€ ๐Ÿ”Š Sound effects punctuating the dialogue
                    โ”œโ”€โ”€ ๐ŸŽต Instrumental interludes with timed duration
                    โ””โ”€โ”€ ๐ŸŽฌ Ending credits theme with fade-out
                                 โ–ผ
                    Final MP3 / WAV export

It requires no audio-editing skills โ€” write (or let the AI write) a script in a simple text format, and EchoForge handles casting, performance, music and mixing.

๐ŸŒŸ Key Features

๐ŸŽญ Voice & Casting

  • Three TTS providers โ€” Fish Audio (cloned voices, 4 models) ยท MiniMax (Coding Plan sk-cp keys work! 13 built-in voices) ยท Microsoft Edge TTS (free, no key, 63 zh/en/ja voices)
  • Voice transformation (DSP) โ€” make ONE voice play MANY characters: 12 presets (Boy / Elder / Elf / Robotโ€ฆ) + manual pitch (ยฑ12 semitones, WSOLA time-preserving), formant shift, brightness, breathiness, roughness
  • SSML pronunciation control โ€” <break time="500ms"/> works across ALL providers (native on Microsoft, auto-converted elsewhere); emotion markers [softly] (S2) / (softly) (S1)
  • Per-character voice binding with volume / speed / dB fine-tuning and character profiles
  • Voice preview per voice-model and per role
  • Content-hash caching โ€” successful segments are never billed twice

๐ŸŽต Music Engine

  • BGM with adaptive ducking: each dialogue block is wrapped in [intro โž speech(BGM ducks to 25%) โž outro]
  • Per-directive fade control: #BGM=Night Piano,3,8 sets 3s fade-in / 8s fade-out
  • Instrumental interludes #MUSIC=track,12 โ€” timed music-only segments (auto-loops if the track is shorter)
  • Ending theme #END=track,20 with a long fade-out
  • 5 built-in royalty-free BGM tracks (โ‰ฅ2 min each, mostly piano) + batch upload your own

๐Ÿ”Š Sound Effects

  • 20 built-in programmatically-synthesized royalty-free SFX (thunder, doorbell, footsteps, page-flip, birds, campfireโ€ฆ)
  • #SFX=thunder,1.2 โ€” punctuate dialogue with volume control; independent SFX library with batch upload

๐Ÿค– AI Scriptwriting & Agent Mode

  • Works with any OpenAI-compatible API (OpenRouter, DeepSeek, Kimi, Qwen, Doubao, Ollamaโ€ฆ) and Anthropic Claude (auto-detected)
  • One-click model discovery (โšก fetches the full model list)
  • Context-rich prompting โ€” the full project state (characters + profiles + voice bindings + BGM/SFX libraries + emotion-marker style matching the selected TTS model) is injected into every generation
  • Agent Workbench โ€” fully automated series production:
    1. AI writes the prologue + Episode 1
    2. You give rounds of feedback โ†’ AI revises (round 2, 3, โ€ฆ)
    3. Generate next episode with full story continuity
    4. Apply to workbench โ†’ auto voice-binding โ†’ auto generation

๐Ÿ–ฅ๏ธ Studio UX

  • Clean light UI with grouped navigation (Configure / Create / Produce)
  • Live progress: percentage, stage pipeline (Ready โž Per-segment โž Mix โž Export โž Done), and a card showing exactly which line is being processed
  • Segment timeline table โ€” edit text, BGM, fades and durations inline; regenerate single segments
  • RPM limiting + classified retries (429/503/network back-off, friendly Chinese/English errors for 401/402)
  • Offline mock mode โ€” rehearse the full mix with placeholder voices, zero API cost

๐Ÿ’พ Downloads (all platforms)

Platform File How
๐ŸชŸ Windows EchoForge-vX-win-x64.exe double-click
๐Ÿง Linux EchoForge-vX-x86_64.AppImage (or the raw binary) chmod +x && ./EchoForge*
๐ŸŽ macOS EchoForge-macOS-Linux-source-vX.zip bash install-macos.sh โ†’ EchoForge.app
๐Ÿค– Android EchoForge-Android-Termux-vX.zip bash install-android.sh in Termux
๐Ÿ“ฑ Any phone (PWA) โ€” run EchoForge on any computer with --host 0.0.0.0, open http://pc-ip:8080 on the phone, Add to Home Screen

The macOS/Linux/Android packages contain identical source + one-click installers that create .app / launchers automatically. All three ship every feature โ€” TTS providers, voice transform, SSML, i18n, Agent.

๐Ÿš€ Quick Start

Option A โ€” Ready-to-run EXE (recommended)

  1. Download EchoForge-v2.0.0-win-x64.exe from Releases (~55 MB, single file)
  2. Double-click โ€” first launch self-extracts the runtime (3โ€“10 s) and opens your browser at http://127.0.0.1:17860
  3. Configure keys:
    • TTS: free API key from fish.audio (model s2.1-pro-free is free)
    • LLM (optional): any OpenRouter / DeepSeek / local Ollama key for AI scriptwriting
  4. Try Agent Workbench โ†’ type a story idea โ†’ ๐Ÿš€

SmartScreen note: the binary is unsigned โ€” click More info โž Run anyway. Upgrades are automatic: launching a newer EXE cleanly replaces the running old service.

Option B โ€” From source

git clone https://github.com/TechnologyStar/EchoForge.git
cd EchoForge
pip install -r requirements.txt
python app.py        # opens http://127.0.0.1:17860

Windows users can also double-click build.bat (needs Python 3.11โ€“3.13) to produce their own single-file EXE via PyInstaller.

๐Ÿ“– Script Format

The whole drama is plain text:

// ===== Act 1 =====
#BGM=Rainy Night Piano,3,8        โ† BGM with 3s fade-in / 8s fade-out
#SFX=doorbell                      โ† punctuate with a sound effect
ใ€Narratorใ€‘Late at night, the bell of the old bookstore rang.
Store owner: "This lateโ€ฆ a customer?"
#MUSIC=Suspense Pulse,10          โ† 10-second instrumental interlude
Girl: [softly] I'm looking for a book that never grows old.
#END=Moonlight Lullaby,18         โ† ending theme, fades out
Directive Meaning
ใ€Roleใ€‘line or Role: line Dialogue (both forms, mixable)
#BGM=name[,fadein,fadeout] Background music for subsequent lines
#BGM=none Clear BGM
#SFX=name[,volume] Sound effect insert (volume 0โ€“1.5)
#MUSIC=name,seconds Music-only interlude (timed)
#END=name,seconds Ending theme (long fade-out)
// comment Ignored line
[softly] / (softly) Emotion markers โ€” brackets for S2 models, parentheses for S1

Full specification: docs/script-format.md

๐Ÿ”ง Configuration Guides

๐Ÿ“ฑ Android / macOS / Linux

  • Android (Termux): pkg install python flask requests numpy soundfile lameenc, pip install edge-tts, then bash run-android.sh โ€” serves on port 8080, open from any browser on the same Wi-Fi
  • macOS / Linux: pip install -r requirements.txt && python app.py (identical to Windows source mode; --host 0.0.0.0 --port 8080 for LAN)

๐Ÿ› ๏ธ Tech Stack

Layer Choice
Backend Python 3.11 ยท Flask ยท NumPy ยท soundfile ยท lameenc
Frontend Vanilla JS ยท hand-rolled light-theme design system
TTS Fish Audio REST (/v1/tts, model via header)
LLM OpenAI-compatible + Anthropic native (auto-switch)
Packaging Custom C launcher (static zlib self-extractor, launcher.c) ยท mingw-w64 cross-compiled

The single-file EXE layout: [C launcher | zlib payload | u64 length | magic] โ€” the launcher version-checks a running service, performs automatic upgrades (kills the old service), and self-extracts the Python runtime on first run. No PyInstaller dependency for official builds.

๐Ÿ—บ๏ธ Roadmap

  • UI localization (EN / ZH / JA switcher) โœจ v2.0
  • SSML-level pronunciation control โœจ v2.0
  • Multi-episode batch export with ID3 tags โœจ v2.0
  • Waveform preview timeline โœจ v2.0
  • macOS / Linux builds

๐Ÿค Contributing

Issues and PRs are welcome! The audio synthesizers (bgm_synth.py, sfx_synth.py) are pure NumPy โ€” new royalty-free BGM/SFX recipes are especially appreciated.

๐Ÿ“„ License

MIT โ€” the built-in BGM/SFX are programmatically synthesized and royalty-free.


โญ If EchoForge saved you an afternoon of audio editing, consider starring the repo!

Made with ๐ŸŽ™๏ธ and a lot of piano notes.

About

๐ŸŽ™๏ธ AI-Powered Radio Drama Production Studio โ€” Multi-speaker TTS ยท Adaptive BGM Ducking ยท Sound Effects ยท LLM Scriptwriting ยท Fully Automated Agent Mode

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages