This document provides a detailed reference for the internal Python API of mudio. It is useful for developers who want to use mudio as a library in their own scripts or applications.
For most use cases, use the high-level process_file() and process_files() functions. They provide automatic backups, verification, filtering, parallel processing, and comprehensive error handling.
from mudio.processor import process_file
from mudio.operations import write, append, delete
result = process_file(
"song.mp3",
ops=[
write('artist', 'The Beatles'),
write('album', 'Abbey Road'),
append('genre', 'Rock'), # Add to existing genres
delete('comment') # Remove field entirely
],
backup_dir='./backups'
)
if result['passed']:
print(f"✓ Updated {list(result['changed'].keys())}")
else:
print(f"✗ {result['error']}")Mudio provides a layered API with multiple entry points depending on your use case:
process_file- Process a single file with full controlprocess_files- Process multiple files (takes a list of Paths)process_batch- High-level batch API with directory scanning and result summarizationwrite_fields- Convenient shorthand for writing multiple fields at once
process_file(path: str, ops: List[FieldOperationsType], ...) -> ProcessResultType
Processes a single audio file with comprehensive error handling, backup, and verification.
Parameters:
- path (str): File path to process
- ops (List[FieldOperationsType]): List of operation functions to apply (see
mudio.operations) - filters (List[FilterType], optional): Filter conditions to check before processing
- dry_run (bool): If
True, calculates changes but does not write to disk. Default:False - backup_dir (str, optional): Directory to store backups before modification
- delete_backups (bool): If
True, deletes backups after successful operations. Default:False - force (bool): If
True, overwrites existing backups. Default:False - verify (bool): If
True, re-reads file after writing to verify changes. Default:True - read_schema (str, optional): Schema for reading metadata (
'canonical','extended','raw'). Default:None(uses global config)
Returns: ProcessResultType (Dict) with keys:
'passed'(bool): Whether processing succeeded'path'(str): File path that was processed'ext'(str): File extension'original'(Dict): Original field values'planned'(Dict): New field values that were computed'changed'(Dict): Which fields actually changed'wrote'(bool): Whether write operation succeeded'verified'(Dict): Verification results (if enabled)'error'(str or None): Error message if failed'skipped'(bool): If file was skipped (e.g., filter didn't match)'note'(str): Additional notes (e.g., 'no changes', 'dry-run')
Example:
from mudio.processor import process_file
from mudio.operations import write, prefix
# Process with dry-run to preview changes
result = process_file(
"song.mp3",
ops=[
write('album', 'Greatest Hits'),
prefix('title', '[Remastered] ')
],
dry_run=True
)
if result['passed']:
print(f"Planned: {result['planned']}")
print(f"Would change: {list(result['changed'].keys())}")process_files(files: Iterable[Path], ops: List[FieldOperationsType], ...) -> List[ProcessResultType]
Processes multiple files. Automatically chooses between sequential and parallel processing based on file count and configuration.
Parameters: Same as process_file(), plus:
- files (Iterable[Path]): Iterable of file paths to process
- max_workers (int): Number of parallel threads. Default
0= auto-detect. Set to1for sequential processing
Returns: List of ProcessResultType dictionaries, one per file
Example:
from mudio.processor import process_files
from mudio.operations import write, enlist
from pathlib import Path
# Process files with filtering (only files with specific artist)
results = process_files(
Path('music').glob('*.flac'),
ops=[enlist('genre', 'Electronic;Ambient')],
filters=[('artist', 'Brian Eno', False)], # Plain text match
max_workers=4
)
print(f"Updated {sum(r['passed'] for r in results)} files")process_batch(path: Union[str, Path], operations: List[FieldOperationsType], ...) -> Dict[str, Any]
High-level API that combines file collection with batch processing. Automatically scans directories and provides a summary of results.
Parameters:
- path (Union[str, Path]): Directory or file path to process
- operations (List[FieldOperationsType]): List of operation functions to apply
- recursive (bool): If
True, search subdirectories. Default:False - extensions (List[str], optional): File extensions to include (e.g.,
['.mp3', '.flac']). Default: all supported formats - filters (List[FilterType], optional): Filter conditions
- dry_run (bool): If
True, show changes without writing. Default:False - backup_dir (Union[str, Path], optional): Directory for backups
- delete_backups (bool): If
True, deletes backups after successful operations. Default:False - force (bool): If
True, allow potentially destructive operations. Default:False - verbose (bool): If
True, show detailed progress. Default:False - max_workers (Optional[int]): Number of parallel workers. Default:
None(auto-detect) - verify (bool): If
True, verify writes by reading back. Default:True - read_schema (str, optional): Schema for reading metadata (
'canonical','extended','raw'). Default:None(uses global config)
Returns: Dict with keys:
'processed'(int): Total files processed'successful'(int): Files processed successfully'failed'(int): Files that failed'skipped'(int): Files skipped by filters'results'(List[ProcessResultType]): Detailed results for each file
Example:
from mudio.processor import process_batch
from mudio.operations import find_replace
# Batch replace text in all files under a directory
result = process_batch(
'music/albums',
operations=[find_replace('title', r'\s*\(feat\..*?\)', '', regex=True)],
recursive=True,
extensions=['.mp3', '.m4a'],
backup_dir='./backups'
)
print(f"{result['successful']}/{result['processed']} files updated")
for r in result['results']:
if r.get('error'):
print(f"Failed: {r['path']} - {r['error']}")write_fields(path: Union[str, Path], fields: Dict[str, str], **kwargs) -> Dict[str, Any]
Convenience function to set multiple fields at once. Automatically converts field dictionary into write operations and calls process_batch().
Parameters:
- path (Union[str, Path]): Directory or file path to process
- fields (Dict[str, str]): Dictionary mapping field names to values
- kwargs: Additional arguments passed to
process_batch()(e.g.,recursive,dry_run,backup_dir)
Returns: Same as process_batch()
Example:
from mudio.processor import write_fields
# Convenience function for setting canonical fields only
result = write_fields(
'music/album',
fields={
'artist': 'The Beatles',
'album': 'Abbey Road',
'date': '1969'
},
recursive=True
)
print(f"Updated {result['successful']} files")Note
write_fields() accepts any fields (canonical or custom). It internally uses write() operations for all provided fields. For more complex logic (like appending or conditional writing), use process_batch() with explicit operations.
validate_file(path: Path) -> Tuple[bool, str]
Performs comprehensive file validation. Returns (success: bool, message: str) tuple:
- File existence and type
- Read/write permissions
- File size (empty files or files exceeding limits)
- Format validity (supported extension check)
This module defines how field values are transformed. Operations are used with process_file() and process_files().
Operations are factory functions that return transformation functions:
from mudio.operations import find_replace
# Create an operation
op = find_replace('title', 'Demo', 'Final')
# The operation is a function: List[str] -> List[str]
result = op(['Demo Track']) # Returns ['Final Track']All operations return a transformation function: Callable[[List[str]], List[str]]
| Operation | Purpose | Example |
|---|---|---|
write(field, value, delimiter=';', index=None) |
Overwrite all items (or specific item at index) | write('album', 'Hits', index=0) |
append(field, value, delimiter=';', index=None) |
Append to all items (or specific item at index) | append('title', ' (v2)') |
prefix(field, value, index=None) |
Prepend to all items (or specific item at index) | prefix('artist', 'The ') |
find_replace(field, find, replace, regex=False, delimiter=';', index=None) |
Substitute in all items (or specific item) | find_replace('title', 'Demo', 'Final', index=0) |
enlist(field, value, delimiter=';') |
Add NEW item(s) to list | enlist('genre', 'Rock;Pop') |
delist(field, value, delimiter=';') |
Remove item(s) from list | delist('genre', 'Pop') |
clear(field) |
Set field to empty string (['']) |
clear('comment') |
delete(field) |
Remove field entirely | delete('custom_field') |
Combining Operations:
from mudio.operations import write, enlist, find_replace
ops = [
write('album', 'Best Of Collection'),
enlist('genre', 'Rock;Classic'), # Add genres if not present
find_replace('title', r'\\s+', ' ', regex=True) # Normalize whitespace
]All fields are stored as Dict[str, List[str]] throughout mudio, even single-valued fields:
fields = {
'title': ['My Song'], # Single-valued, but still a list
'artist': ['John', 'Jane'], # Multi-valued list
'album': ['Greatest Hits'], # Single-valued, but still a list
'genre': ['Rock', 'Pop'] # Multi-valued list
}This uniform representation simplifies API usage and makes operations consistent across all field types.
While all fields are represented as lists in memory, mudio distinguishes between single-value and multi-value fields when parsing operations and writing to disk:
- Single-Value (Default): All fields unless specified otherwise (including all custom fields).
- Delimiter Ignoring: Operations on these fields (like
write,enlist) treat the input string as a single literal value and do NOT split it by semicolons. This allows titles like"Whatcha;Whatcha Doin'"to be preserved. - Index 0 Persistence: Only the first item in the list (
index 0) is written to the audio file. Any additional items added to the list in memory will be lost upon saving.
- Delimiter Ignoring: Operations on these fields (like
- Multi-Value Fields:
artist,albumartist,composer,genre,performer, andcomment.- Delimiter Splitting: Operations automatically split input strings by the delimiter (default
;) into multiple list items. - Full Persistence: All items in the list are written to the file (supported formats like FLAC, MP4, and ID3v2.4 handle multiple values natively).
- Delimiter Splitting: Operations automatically split input strings by the delimiter (default
Note
Value-Level Deduplication: Duplicate values are automatically removed from all fields (case-insensitive, order-preserved). This ensures clean metadata even after multiple append/merge operations.
Mudio enforces strict whitespace and empty value handling during both reading and writing:
On Read (read_fields):
- All values are stripped of leading/trailing whitespace
- Values that become empty after stripping are dropped
- If all values for a field are dropped (or the field was empty), the field returns
[''](a list containing a single empty string) rather than being omitted
On Write (write_fields):
- All values are stripped of whitespace
- Empty strings are dropped from lists
- A single explicit
[''](e.g. fromclear()) is preserved and written to the file — this is the "smart empty" behavior - Writing
[](empty list) deletes the field entirely
with SimpleMusic.managed('song.mp3') as sm:
fields = sm.read_fields()
fields['title'] # ['My Song'] — normal value
fields['comment'] # [''] — field exists but is emptyAll fields are treated as multi-valued lists. Operations apply to all items by default, and values are automatically deduplicated (case-insensitive, order-preserved).
Mudio applies value-level deduplication to all fields:
- Rule: Case-insensitive comparison, order-preserved
- Scope: All values from all frames are merged and deduped
- Outcome:
['Rock', 'rock', 'Pop']→['Rock', 'Pop']
Additionally, frame-level deduplication removes identical frames on read:
['Values'](Frame A) +['Values'](Frame B) →['Values']['Val A']+['Val B']→['Val A', 'Val B'](distinct frames kept)
The semicolon (;) is the default delimiter for parsing strings into multiple values.
Default delimiter: ; (customizable via delimiter parameter)
Important
Delimiter parsing is only automatic in operations (write, append, enlist, etc.). When using SimpleMusic.write_fields() directly, you must parse strings yourself using SimpleMusic.parse_list_string() or provide pre-parsed lists.
How it works:
# Writing multiple values with delimiter
write('genre', 'Rock;Pop;Jazz') # Creates ['Rock', 'Pop', 'Jazz']
write('artist', 'John Doe; Jane Smith') # Creates ['John Doe', 'Jane Smith']Multiple delimiters (via SimpleMusic.parse_list_string only):
Pass a list of strings to parse_list_string() to split on any of them:
from mudio.core import SimpleMusic
# Split on semicolons, slashes, or commas
SimpleMusic.parse_list_string('Rock;Pop/Jazz,Blues', delimiter=[';', '/', ','])
# → ['Rock', 'Pop', 'Jazz', 'Blues']Note
Operation functions (write, append, enlist, etc.) accept a single delimiter string only. For multiple delimiters, use SimpleMusic.parse_list_string() directly with SimpleMusic.write_fields().
Operations that use delimiters:
write(field, value, delimiter=';')— parse value string into listappend(field, value, delimiter=';')— parse and add new valuesenlist(field, value, delimiter=';')— parse and add if not presentdelist(field, value, delimiter=';')— parse and remove matching valuesfind_replace(field, find, replace, regex, delimiter=';')— parse replacement results
Example:
from mudio.operations import enlist
# Add multiple genres at once
enlist('genre', 'Electronic;Ambient') # Adds both if not present
# With custom delimiter
enlist('artist', 'John | Jane', delimiter='|') # ['John', 'Jane']Delimiter rules:
- Whitespace around values is automatically stripped:
"A; B ; C"→['A', 'B', 'C'] - Empty values are dropped:
"A;;B"→['A', 'B'] - Whitespace-only values are dropped:
"A; ;B"→['A', 'B']
⚠️ Most users should useprocess_file()instead. UseSimpleMusiconly if you need fine-grained control or custom logic beyond standard operations.
from mudio.core import SimpleMusic
# Always use context manager for proper cleanup
with SimpleMusic.managed("song.mp3") as sm:
# Read metadata
fields = sm.read_fields(schema='extended')
# Custom transformation
if 'artist' in fields:
fields['artist'] = [a.upper() for a in fields['artist']]
# Write back
sm.write_fields(fields)SimpleMusic(path: Union[str, Path])
Opens an audio file for metadata access. Raises FormatError or RuntimeError on failure.
read_fields(schema: Optional[str] = None) -> Dict[str, List[str]]
Reads metadata from the file. Returns a dict where keys are field names and values are lists of strings (even for single-valued fields).
Schema Options:
'canonical': Only standard fields (title, artist, album, date, etc.)'extended'(default): Canonical plus custom/extra fields with normalized keys'raw': Returns all fields with format-specific native keys (e.g.,TIT2,TXXX:FIELDfor MP3,©namfor M4A)
Note
Raw mode returns format-native keys unchanged, while extended mode returns normalized keys. For example:
- Extended:
my_custom_field,artist,title(portable across formats) - Raw:
TXXX:MY_CUSTOM_FIELD,TPE1,TIT2(MP3-specific ID3 frames)
Key Normalization (Extended/Canonical modes only):
- Canonical fields always use standard lowercase names (
artist,album,title, etc.) - Custom field read keys are sanitized to lowercase
[a-z0-9_](e.g.,My-Field!→my_field_) - Custom field write keys are sanitized to uppercase
[A-Z0-9_]for format-specific storage - Prevents duplicate fields with different casing
write_fields(fields: Dict[str, List[str]])
Writes metadata to the file. Custom fields are written as format-specific tags. Fields not in the dict are preserved. To delete a field, pass an empty list: [].
Note
write_fields() expects pre-parsed lists. It does NOT automatically parse delimiter-separated strings. Use parse_list_string() to convert strings like "Rock;Pop;Jazz" to ['Rock', 'Pop', 'Jazz'], or use the operations API which handles this automatically.
SimpleMusic.managed(path) (static method, context manager)
Returns a context manager for safe file handling.
SimpleMusic.parse_list_string(s, delimiter=';') (static)
Utility to split delimited strings into lists. Use this when working with SimpleMusic.write_fields() directly:
# Manual parsing required for SimpleMusic
genres = SimpleMusic.parse_list_string('Rock;Pop;Jazz')
sm.write_fields({'genre': genres}) # ['Rock', 'Pop', 'Jazz']Config
Global configuration settings loaded from environment variables:
MUDIO_SCHEMA: Default schema (canonical,extended, orraw). Default:extendedMUDIO_MAX_WORKERS: Default thread count for parallel processingMUDIO_VERBOSE: Default verbosity (0or1)MUDIO_NAMESPACE: Default namespace for MP4/M4A custom fields (default:com.apple.iTunes)
from mudio.utils import Config
print(f"Default schema: {Config.DEFAULT_SCHEMA}")get_file_hash(path: Path) -> str
Calculates SHA-256 hash of a file for verification.
class MudioError(Exception): ...
class ValidationError(MudioError): ... # Invalid arguments or state
class FormatError(MudioError): ... # File format issues
class PermissionError(MudioError): ... # Filesystem permission issues
class VerificationError(MudioError): ... # Post-write verification failedFieldOperationsType = Callable[[List[str]], List[str]] # Transformation function
FieldValuesType = Dict[str, List[str]] # Metadata dictionary
FilterType = Tuple[str, str, bool] # (field, pattern, is_regex)
ProcessResultType = Dict[str, Any] # Result from process_file()