Evolutionary neural architecture search for automatic chord estimation: a reproducible, vocabulary-aware study of structured harmonic sequence models
| dc.contributor.advisor | Akilan, Thangarajah | |
| dc.contributor.author | Frost, Russell | |
| dc.contributor.committeemember | Yassine, Abdulsalam | |
| dc.contributor.committeemember | Carastathis, Aris | |
| dc.contributor.committeemember | Zhou, Yushi | |
| dc.date.accessioned | 2026-09-15T14:38:46Z | |
| dc.date.created | 2026 | |
| dc.date.issued | 2026 | |
| dc.description | Thesis is embargoed until September 15 2027. | |
| dc.description.abstract | Automatic chord estimation (ACE) takes a musical recording and produces a time-aligned sequence of labels. These labels, such as C major, A minor 7, or G7/B, provide a compact description of a song’s harmony and support transcription, teaching, accompaniment, harmonic search, and large-scale musical analysis. Automating this process is difficult because the same harmony can be voiced in different ways, chords often share pitches, chord labels are contextual to surrounding harmonic information, detailed chord types may be rare, and human annotators may disagree. This thesis presents a reproducible evolutionary neural architecture search (NAS) framework for large-vocabulary ACE. Starting from a validated convolutional–recurrent baseline, the framework searches for improved convolutional and recurrent architectures while keeping the musical representation, chord vocabulary, training procedure, and decoder fixed. Multiple independent searches are conducted using a validation objective that balances recognition performance with model complexity, after which selected architectures are independently retrained and evaluated. On the McGill Billboard benchmark test set, the proposed model achieved a Weighted Chord- Symbol Recall (WCSR) of 63 49 0 36%, compared with 60 09 0 16% for the baseline, with a mean paired improvement of 3 39 0 42 percentage points across three initialization seeds. After topology freeze, the same model was retrained from scratch on a genre-stratified, song-disjoint split of Chordonomicon, with the corresponding chord progressions rendered as synthetic audio; it achieves 96 50 0 12% WCSR compared with 95 26 0 14% for the baseline, an improvement of 1 23 0 23 percentage points. | |
| dc.identifier.uri | https://knowledgecommons.lakeheadu.ca/handle/2453/5647 | |
| dc.language.iso | en | |
| dc.subject | Musical analysis | |
| dc.subject | Machine-learning | |
| dc.subject | Automation | |
| dc.title | Evolutionary neural architecture search for automatic chord estimation: a reproducible, vocabulary-aware study of structured harmonic sequence models | |
| dc.type | Thesis | |
| etd.degree.discipline | Engineering : Electrical & Computer | |
| etd.degree.grantor | Lakehead University | |
| etd.degree.level | Master | |
| etd.degree.name | Master of Science degree in Electrical and Computer Engineering |
Files
License bundle
1 - 1 of 1
Loading...
- Name:
- license.txt
- Size:
- 2.23 KB
- Format:
- Item-specific license agreed upon to submission
- Description:
