Skip to content

Latest commit

 

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Mertalizer: Music Structure Recognition

A system for detecting musical structure boundaries and assigning section labels (intro/verse/pre-chorus/chorus/bridge/solo/outro/other) using music-native SSL encoders.

🎵 Features

  • Primary Model: MERT v1 95M SSL encoder with a TCN boundary detector and linear label head
  • Optional Variants: Swap in w2v-BERT-2.0 encoder or a transformer boundary head via config overrides
  • Datasets: Harmonix Set, SALAMI, Beatles/Isophonics, SPAM, CCMUSIC
  • Outputs: Trained checkpoints, CLI & REST inference, evaluation reports
  • Web Dashboard: Interactive audio upload and visualization interface
  • Rust Integration: High-performance web server and audio processing

🏗️ Project Structure

mertalizer/
├── Cargo.toml             # Rust project configuration
├── src/                   # Rust web server (main entry point)
│   ├── main.rs           # Axum web server
│   ├── predictor.rs      # Python CLI integration
│   └── history.rs        # Prediction history management
├── ml/                    # Python ML components
│   ├── data/             # Data processing and ingestion
│   │   ├── ontology.py   # Label mapping system
│   │   ├── fetch_datasets.py  # Dataset downloader
│   │   ├── ingestion.py  # Data normalization
│   │   └── preprocessing.py   # Audio preprocessing
│   ├── modeling/         # Model architectures
│   │   ├── system.py     # Two-head architecture
│   │   └── ssm_novelty.py # Classic baseline
│   ├── training/         # Training scripts
│   │   └── train.py      # LightningModule training
│   ├── evaluation/       # Metrics and evaluation
│   │   └── metrics.py    # Boundary & label metrics
│   ├── inference/        # CLI and REST API
│   │   ├── cli.py        # Command-line interface
│   │   └── api.py        # FastAPI REST server
│   └── export/           # ONNX export
│       └── onnx.py       # Model export utilities
├── configs/              # Model configurations
│   ├── mert_95m.yaml    # MERT configuration
│   └── w2v_baseline.yaml # w2v-BERT configuration
├── static/               # Static web assets
├── templates/            # HTML templates
│   └── index.html       # Web dashboard UI
├── scripts/              # Utility scripts
│   ├── demo.sh          # End-to-end capability walkthrough
│   ├── example.sh       # Sample pipeline runner
│   ├── run_server.sh    # Launch web server
│   └── setup.sh         # Environment bootstrapper
├── data/                 # Dataset storage
└── models/               # Trained checkpoints

🚀 Quick Start

Option 1: Complete Example Pipeline

# See all features in action
./demo.sh

# Run the complete example pipeline
./example.sh

Option 2: Manual Setup

# 1. Setup environment
./setup.sh

# 2. Download datasets
python ml/data/fetch_datasets.py --all

# 3. Process data
python ml/data/ingestion.py --all

# 4. Extract embeddings (repeat for each dataset JSONL)
python ml/data/preprocessing.py --annotations data/processed/ccmusic.jsonl --dataset-name ccmusic

# 5. Train model on precomputed embeddings
python ml/training/train.py --config configs/mert_95m.yaml

# 6. Start web servers
./run_server.sh

The embedding step looks up the dataset field in each JSON line and expects artifacts under data/processed/embeddings/<dataset>/<track_id>.npz. Repeat the command above for every dataset JSONL you ingested (e.g. harmonix.jsonl, salami.jsonl) and adjust --dataset-name accordingly.

🔧 Environment Variables

Create a .env file with:

# Hugging Face token for dataset access
HF_TOKEN=your_hf_token_here

# Weights & Biases API key (optional)
WANDB_KEY=your_wandb_key_here

# Model paths
MODEL_PATH=models/checkpoints/final_model.ckpt
PYTHON_API_URL=http://localhost:8000

# Server settings
PORT=3000

📊 Usage

Web Dashboard

  • URL: http://localhost:3000
  • Upload audio files and visualize structure analysis
  • Interactive waveform with segment boundaries
  • Real-time structure charts and statistics

CLI Inference

# Single file
python ml/inference/cli.py audio.wav --model models/checkpoints/final_model.ckpt

# Batch processing
python ml/inference/cli.py audio_dir/ --model models/checkpoints/final_model.ckpt --batch

# Export results
python ml/inference/cli.py audio.wav --model models/checkpoints/final_model.ckpt --output results.json

REST API

# Health check
curl http://localhost:8000/health

# Predict structure
curl -X POST -F "audio_file=@audio.wav" http://localhost:8000/predict

# API documentation
open http://localhost:8000/docs

ONNX Export

# Export trained model to ONNX
python ml/export/onnx.py --checkpoint models/checkpoints/final_model.ckpt --output models/model.onnx

# Test ONNX model
python ml/export/onnx.py --checkpoint models/checkpoints/final_model.ckpt --output models/model.onnx --test

🎯 Model Architecture

Two-Head Architecture

  • Encoder: Frozen/fine-tunable SSL encoder (MERT 95M by default, w2v-BERT optional)
  • Boundary Head: TCN detector (default) with configurable transformer alternative
  • Label Head: Linear classifier predicting canonical segment classes
  • Loss: Binary cross-entropy with optional focal term for boundaries; cross-entropy for labels

Classic Baseline

  • Self-Similarity Matrix: Cosine similarity between embeddings
  • Novelty Detection: Checkerboard kernel convolution
  • Peak Detection: Threshold-based boundary detection
  • Label Assignment: Majority vote or heuristic rules

📈 Evaluation Metrics

  • Boundary Detection: F@0.5s, F@3.0s, Hit Rate
  • Label Classification: Accuracy, Precision, Recall, F1
  • Cross-Dataset: Generalization across different datasets
  • Error Analysis: Confusion matrices and case studies

🔬 Datasets

Dataset Source Annotations Audio
Harmonix Set Hugging Face Structure + Beats Provided
SALAMI GitHub Hierarchical External IDs
Beatles/Isophonics Hugging Face Structure + Chords External
SPAM GitHub Multi-annotator External
CCMUSIC Hugging Face Chinese Pop External

🛠️ Development

Adding New Datasets

  1. Update ml/data/fetch_datasets.py with dataset configuration
  2. Implement parser in ml/data/ingestion.py
  3. Add label mappings in ml/data/ontology.py

Custom Models

  1. Create new model class in ml/modeling/
  2. Update ml/training/train.py to support new architecture
  3. Add configuration in configs/

Web Dashboard Features

  1. Modify templates/index.html
  2. Update Rust web server in src/main.rs
  3. Add new API endpoints in ml/inference/api.py

📝 Output Format

{
  "track_id": "example_track",
  "sr": 22050,
  "duration": 180.5,
  "boundaries": [0.0, 12.5, 25.0, 37.5, 50.0],
  "labels": ["INTRO", "VERSE", "CHORUS", "VERSE", "OUTRO"],
  "segments": [
    {"start": 0.0, "end": 12.5, "label": "INTRO"},
    {"start": 12.5, "end": 25.0, "label": "VERSE"},
    {"start": 25.0, "end": 37.5, "label": "CHORUS"},
    {"start": 37.5, "end": 50.0, "label": "VERSE"},
    {"start": 50.0, "end": 180.5, "label": "OUTRO"}
  ],
  "version": "mert-95m-tcn@2025-01-24"
}

🤝 Contributing

  1. Fork the repository
  2. Create a feature branch
  3. Make your changes
  4. Add tests if applicable
  5. Submit a pull request

📄 License

MIT License - Educational/Research Use

🙏 Acknowledgments

  • MERT team for the SSL encoder
  • Hugging Face for dataset hosting
  • PyTorch Lightning for training framework
  • Axum for Rust web framework

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages