Cluster Analysis–Part 1

—

by

in

This week we talk about how to apply cluster analysis to the Skiptune database. Skiptune is unusually well suited to clustering because we have both the symbolic melodies and extensive metadata. We distinguish clustering the tunes themselves from clustering composers, genres, historical periods, or musical fragments.

The most interesting analysis, given what we’ve already built, would be to cluster tunes using only their musical content, deliberately withholding our labels. Then ask whether the resulting clusters independently recover things like genre, period and composer. That gives you an empirical test of how much stylistic information is actually encoded in melody.

You might why bother with cluster analysis when we already have high-quality labels as part of the metadata. The answer is that while we defend the labels as high-quality, they are not perfect. Furthermore, some may be so highly correlated with each other that one of them is redundant. When we get started on developing an AI model, using redundant labels adds to the cost of computing (and time spend) while not adding to the quality of the model. Identifying them with cluster analysis now will clean up our modeling effort later.

Candidates

We’re going to walk through some cluster analysis ideas that look interesting before deciding which ones are worthwhile and in what order.

Historical Clustering

Because we know from our earlier Heaps Curve analysis that the era associated with a tune matters enormously in terms of vocabulary, we wonder if the tunes would group themselves in that way even if we don’t tell the algorithm anything about dates. We would cluster solely from melodic content and afterward calculate the median composition year of each cluster.

Say clusters naturally order themselves something like:

  • Cluster A median: 1740
  • Cluster B: 1805
  • Cluster C: 1870
  • Cluster D: 1925
  • Cluster E: 1970

That would demonstrate quantitatively that melodic language contains historical information.

More interesting would be identifying transitional composers or tunes. For example, a tune written in 1820 might consistently cluster with music from 1870. That gives us an empirical definition of something being stylistically ahead of its time. Conversely, something written in 1920 that clusters strongly with 1870 material would be quantitatively conservative.

Composer “Neighborhoods”

For composers with sufficient representation, aggregate their melodies into composer-level vectors. It might look something like this dendrogram:
                 ┌─ Composer A
             ┌───┤
             │   └─ Composer B
         ┌───┤
         │   └───── Composer C
─────────┤
         │       ┌─ Composer D
         │   ┌───┤
         │   │   └─ Composer E
         └───┤
             └───── Composer F

The branch lengths quantify melodic distance. We could then ask whether composers historically regarded as stylistically related actually appear close together when judged only from melody.

Genre Purity

Cluster all tunes blindly and calculate the genre composition of each cluster. That lets us calculate something like cluster purity or normalized mutual information.

Consider a group of labels like folk, march, waltz, ragtime, baroque, classical, etc. If musical clustering predicts those human labels well, your labels correspond to genuine melodic distinctions. If it doesn’t, that’s equally interesting. It might show, for example, that genre is substantially determined by instrumentation, harmony, production and cultural context rather than melody itself. Or it might demonstrate a weak relationship suggesting that it is melody combined with harmony, etc., that determines genre purity.

That would fit directly into the larger question we’ve been investigating:  What properties of music reside in melody?

Find “Musical Orphans”

Clustering also identifies outliers. Some tunes will be far from every cluster centroid. Those may be particularly interesting compositions because their melodic structure is unusual relative to the rest of the ~84,000-token corpus. We could generate a ranked list:

RankTuneNearest clusterDistance
1XFolk-3.89
2YPopular-7.84
3ZClassical-2.81

That effectively asks, “What are the strangest melodies in Skiptune?” That’s probably as far as cluster analysis would take us. We’d likely have to listen to each one to figure out what makes them “strange”.

Intervallic Clustering

This is the opposite of the previous cluster. Throw away the duration and consider only the intervals in the melody. We would use our pitch differentials to establish a string of melody without duration values. If our pitches are MIDI numbers:
60, 62, 64, 67, 65
then take successive differences:
+2, +2, +3, -2
Those are directed melodic intervals in semitones, which is the same as our pitch differentials.

The next question is how one turns those differentials into variables suitable for clustering. We could derive features such as mean pitch differential, mean absolute differential, standard deviation, percentage of 0, ±1, ±2, ±3, etc., percentage of steps versus leaps, ascending/descending proportions, or interval n-gram frequencies.

Tuple Clustering

Our use of tuples, each containing a pitch differential and a duration ratio, is a perfect candidate for cluster analysis. We would have to normalize each tune for its length so a 200-note tune isn’t automatically removed from a cluster of a 40-note tune.

We would then calculate the similarity between every pair of tunes. Cosine similarity or Jensen-Shannon distance would be reasonable starting points. We’ll explain those more as we use them. Throw that mix into the cluster analysis absent all other metadata, such as year, composer, genre, etc. After the clustering is complete, assign the labels to each tune and calculate the percentages for each label in each cluster.

We might get a final cluster that looks like this:

  • 72% folk
  • 14% country
  • 8% hymn
  • 6% other

Other clusterings might have completely different labels dominated in percent by, say, the baroque era. That would be strong evidence that our symbolic tuple representation captures meaningful stylistic structure.

Next week we dive in and begin the cluster analysis. We’ll start with our 34 metrics detailed elsewhere on the website, expand it for completeness, and then evaluate the ones we should use for cluster analysis. We don’t know if this will be a useful exercise, but the cost of performing it isn’t high and the payoff is huge if it turns out that our tuples contain the information we think it does to define melodic structure.