This project is a machine learning approach to analyze StepMania step charts and provide consistent difficulty ratings across all songs. The system extracts meaningful features from .sm, .ssc, and .dwi files and use them to train models for accurate difficulty prediction and rating standardization across multiple historical rating scales.
It was vibe-coded using Claude Sonnet and Haiku 4.5 models, with human guidance at important points along the way. The result is a system to produce consistently scaled ratings across all stepcharts in your collection.
-
Download the calculated ratings from the Releases page.
-
Clone my fork of itgmania and build it from source.
-
Unzip the downloaded
calculated_ratings-YYYYMMDD.zipin the itgmania working copy'sDatafolder.
Now, when you launch itgmania with the built binary, it should pick up
the calculated_ratings.json values and prefer them to the stepchart ones.
If not, double check that songpack folder names match the ones used in the
calculated_ratings.json file.
Want to generate difficulty ratings from your own StepMania chart collection? Here's the complete workflow:
Your directory structure should look like this:
stepmania/
├── stepml/ # This repository
├── Songs/ # Your StepMania song packs
│ ├── DDR 1st Mix/
│ ├── ITG 1/
│ ├── Custom Pack 1/
│ └── ...
└── Save/ # (Optional) For performance enrichment
└── LocalProfiles/
└── 00000000/
└── Stats.xml
Process all your charts and extract features (the slow step, a few minutes):
uv run extract-featuresOptions:
--songs-dir PATH- Path to Songs directory (default:../Songs)--output-dir PATH- Output directory (default:./data/features)--stats-file PATH- Path to Stats.xml for performance enrichment--no-performance- Disable performance data enrichment--verbose- Show progress for every file-j N,--jobs N- Worker processes (default: one per CPU)
Output:
data/features/features.parquet- Feature cache, independent of rating scaledata/features/generation_stats.json- Statistics about the extraction process
Re-run this only when charts or feature code change. Label changes (ground truth overrides, DDRFreak ratings) need only Step 2.
Label the feature cache for one rating scale (takes seconds):
uv run generate-datasetOptions:
--features PATH- Feature cache (default:./data/features/features.parquet)--output-dir PATH- Output directory (default:./data/processed)--normalization-scale {classic_ddr,modern_ddr,itg}- Target rating scale (default:modern_ddr)classic_ddr: 1-10 scale (DDR 1st through Extreme)modern_ddr: 1-20 scale (DDR X onwards) - recommendeditg: 1-12 scale (In The Groove)
Output:
data/processed/dataset.csv- Full dataset in CSV formatdata/processed/dataset.parquet- Full dataset in Parquet format (more efficient)
Example with custom scale:
# Generate dataset normalized to ITG scale
uv run generate-dataset --normalization-scale itgNote: The normalization scale affects the target rating values in the dataset. If you change the scale, you'll need to retrain your models since the target value range changes (Classic DDR: 1-10, Modern DDR: 1-20, ITG: 1-12).
Train regression models on the extracted features:
uv run train-modelsWhat this does:
- Trains Linear Regression and Random Forest models (the Random Forest produces the ratings)
- Evaluates them on a held-out 20% split
Output:
data/models/{linear_regression,random_forest}.pkl- Trained models (pickled), with scalers and metadatadata/models/random_forest_feature_importance.csv- Impurity-based feature importancedata/models/training_summary.csv- Held-out MAE, RMSE, R² and Spearman correlation
Diagnostics: cross-validation and permutation importance are slower and not needed to produce ratings, so they are a separate step, run on demand:
uv run diagnose-models --dataset data/modern_ddr/processed/dataset.parquet \
--model-dir data/modern_ddr/modelsThis writes *_permutation_importance.csv and diagnostics_summary.csv to
the model directory.
Use the trained model to predict ratings for all charts:
uv run generate-ratingsWhat this does:
- Loads the trained model from
data/models/best_model.pkl - Predicts ratings for all charts in your dataset
- Generates
calculated_ratings.jsonin StepMania-compatible format
Output:
data/output/calculated_ratings.json- Ratings in JSON format to use with itgmaniadata/output/ratings_with_predictions.csv- Full dataset with predicted ratings
Format of calculated_ratings.json:
{
"Songs/Pack Name/Song Name/chart.ssc": {
"single_beginner": 3.2,
"single_easy": 5.8,
"single_medium": 8.4,
"single_hard": 11.2,
"single_challenge": 14.7
}
}Copy calculated_ratings.json to your StepMania/itgmania Data folder:
cp data/output/calculated_ratings.json ../Data/Launch itgmania (using the fork that supports this feature) and it will use the calculated ratings instead of the chart-specified ones.
Once you have calculated ratings, you can generate course playlists:
uv run generate-playlistsWhat this does:
- Creates random course playlists based on difficulty ranges
- Generates
.crsfiles indata/courses/ - Organizes courses by difficulty tier (beginner, intermediate, advanced, expert)
Output:
data/courses/*.crs- StepMania course files- Each course contains 4-5 songs within a similar difficulty range
Using the playlists:
# Copy courses to your StepMania Courses directory
cp data/courses/*.crs ../Courses/Launch StepMania and access the courses through the "Course Mode" menu (e.g. "Marathon" mode if using Simply Love theme).
Full list of available commands:
| Command | Purpose |
|---|---|
| uv run generate-baseline | Regenerate regression test baseline |
| uv run extract-features | Extract chart features (feature cache) |
| uv run generate-dataset | Label features to create ML dataset |
| uv run generate-ratings | Generate calculated difficulty ratings |
| uv run generate-playlists | Create StepMania course playlists |
| uv run train-models | Train ML models on dataset |
| uv run diagnose-models | Cross-validate, permutation importance |
| uv run sync-favorites | Sync StepMania favorites lists |
| uv run analyze-performance | Analyze performance-enriched dataset |
For technical details, see documentation in the doc directory.
All copyright disclaimed; see UNLICENSE.
This is a personal project I created to scratch my own itch. If you need help with it, your first line of attack should be feeding the codebase into an LLM and asking it for clarification.
If you are still stuck after doing that, you are welcome to file an issue and I'll try to help. PRs are also welcome, although for big changes you might be better off forking.
0 comments
log in to comment.