Why Metadata Is Everyone's Problem Now
From record store bins to online shopping malls, and from library shelves to AI training data, the same mistake keeps showing up
[TL;DR] A taxonomy might look like a neutral list of categories. It’s not. Building one always means someone decides what content earns its own category, what gets lumped together, and what gets left out altogether — and those decisions aren’t neutral.
My friend Rahel Anne Bailie once told me about a taxonomy project she worked on for a European mobile carrier. If I recall correctly, (most days, that’s questionable 🤨), the job was to make an online music library work for listeners in different countries by organizing it so the people in each market could find the music they wanted.
Bailie, founder of Content Seriously and creator of the Content Integrity Model, was helping a client solve the information and data challenges behind a service intended to compete with Apple iTunes. Her team’s work included categorizing the music library, which seems simple, until you try doing it for music lovers in multiple countries.
How Metadata Enters The Equation
Assigning a song to a genre category creates metadata; information that tells a system what something is and where it belongs. Bailie’s team started with a taxonomy shaped by a classification scheme commonly used in the commercial music industry. That taxonomy wasn’t localized for describing music in the new markets Bailie’s client planned to offer the streaming music service to.
👉🏾 Localization is the process of adapting content to be meaningful for a particular culture, locale, or market.
Source: The Language of Localization (XML Press)
To be clear, the categories the music industry used were well-suited for the audiences for which they were built. The problem materialized as Bailie’s team recognized the incongruities in the way music is categorized in various target markets. The solution to the classification problem they faced involved squaring the pre-existing music taxonomy with how each new target market already organized its music. A big challenge given that the original categories hadn’t been designed with this kind of flexibility in mind.
Somewhere in their work, Rahel’s team noticed that world music wasn’t really a genre at all. It was a container: a catch-all category for anything that didn’t fit into existing ones.
That begs the question: Is sorting music by genre and geography the only way to organize it — or just the way a US- and UK-dominated music industry happened to do it?
Some Countries Organize Music By The Emotion It Evokes
India’s classical music tradition doesn’t organize primarily by genre or geography. It sorts by rasa — the specific emotion a piece of music is built to produce — a system old enough to appear in the Natya Shastra, a text roughly two thousand years old.
Some Countries Organize Music By Color
Traditional Chinese music theory goes further: the five-note scale is mapped directly onto five colors, each note tied to a color, a direction, a season.
In Brazil, the sacred music of Candomblé is organized by orixá — which deity a rhythm calls. Each orixá carries its own color and character, so rhythm, color, and spirit are treated as a single category, not three separate ones.
None of these systems sort music by where it’s from (geography-based music categorization). “World music” doesn’t just group unfamiliar sounds into one bin — it assumes geography-based categorization is the correct way to organize music, an assumption some of the world’s major musical traditions don’t share. That assumption has its own history; and its own paper trail.
It’s not just categories that get stuck at the border. Names can, too. Some of the biggest Western pop stars carry an entirely different identity once they cross into other markets. Britney Spears isn’t categorizes as as “Britney Spears” in China. Chinese fans know her as 小甜甜 (Xiǎo Tiántián), “Little Sweetie,” a nickname with no relationship to a transliteration of her English name. A platform that only indexes her under “Britney Spears” is invisible to anyone searching for 小甜甜.
This Has Been In Heavy Rotation Since 1987💿
The phrase “world music” had existed since the early 1960s, coined by ethnomusicologist Robert E. Brown at Wesleyan University to describe an academic program, not a music marketing category. In 1987, a group of UK record labels met to market a new umbrella term to record stores, so retail shops would have a single bin to store artists that didn’t fit the Western pop and rock mainstream.
That meeting is where the commercial version of world music comes from. It was a marketing category built to categorize music in ways that supported how record stores that sold music actually worked.
Same As It Ever Was, Same As It Ever Was
I was 17-years-old in 1981 when I discovered the song, “Once in a Lifetime,” which turned me on to the new wave band, Talking Heads. Decades later, I was intrigued to find that David Byrne — the Scottish-born American singer and musician best known as the band’s lead singer and guitarist, had something to say about this toss-it-all-in-a-single-bin categorization approach.
In a 1999 New York Times essay, Byrne wrote that:
“….use of the term world music was a way of dismissing artists or their music as irrelevant. It’s a label for anything at all that is not sung in English or anything that doesn’t fit into the Anglo-Western pop universe.”
It’s clear to me that Rahel’s team didn’t stumble onto a brand-new music categorization problem. Instead, they bumped up against an issue that had been openly discussed by people like Byrne. The lesson here is that taxonomies can (and clearly do) get handed down to others without their history attached. That’s why it’s not surprising that the people applying taxonomy data typically don’t know what assumptions are baked in to the category decision-making process, nor why those choices were made in the fist place.
The Recording Academy has recently revised its award categories in ways that reflect some of the same criticism Byrne highlighted in his commentary. In 2020, it renamed Best World Music Album to Best Global Music Album, explaining that the old name carried connotations of colonialism, folk, and “non-American” music. It also added Best Global Music Performance for the 2022 awards. In June 2026, the Academy announced Best Asian Pop Music Performance, a new category that includes K-pop, J-pop, and C-pop. The broad global categories haven’t disappeared, but the Academy has also begun creating more specific categories alongside them.
The Same Pattern Shows Up In Libraries
This isn’t just a quirk of the music industry: the same categorization pattern appears in a scheme many of us rely on: the Dewey Decimal Classification system. It’s the world’s most widely-applied numerical library classification system, in use in more than 135 countries.
In this system, the category “religion” sits in the 200s. The Bible and Christianity receive most of the range from 220 through 289 (meaning most of the Dewey categorization numbers in the religion section are dedicated to identifying books associated with Christianity). Religious traditions aside from Christianity (e.g., Islam, Judaism, Hinduism, Buddhism, Sikhism, and a variety of Indigenous traditions) are concentrated within 290 through 299, often stacked on top of each other in the same three-digit number.
A 2024 study — using a dataset of more than three million books from a consortium of Ohio academic libraries — found that Dewey’s religion categories show measurably more Western skew than its history or literature categories do, and the Dewey system itself is measurably more biased than the Library of Congress system many academic libraries use instead.
Whatever their intentions, both categorization systems reflected the practical and cultural assumptions and biases of the people who created them. Melvil Dewey built a system that made sense from where he stood, in the late 1800s, and a handful of UK record label executives promoted a system for categorizing music that didn’t fit anywhere else in 1987.
Those who inherited Dewey’s system kept using it without questioning whose perspective it reflected. That’s the reason I wanted a second example here, beyond music. I’m assuming most of you reading this haven’t built a music genre taxonomy, but all of you have organized information for someone other than yourself.
This Problem Shows Up Everywhere
The least surprising version of this problem is ubiquitous and shows up in e-commerce, which is why it’s worth mentioning. Retailers operating across multiple countries run into mismatched categorization constantly. A product that fits cleanly into one country’s taxonomy may be awkwardly placed, duplicated, or missing in another’s. It’s such a routine problem that entire product-information-management vendors exist just to manage it.
Zalando, the pan-European fashion retailer, offers an adjacent example. When the company established a Nordic fulfillment center, it said that being closer to customers would help it provide a more locally relevant assortment, including a larger selection of Nordic brands. The example involves merchandising and logistics as well as information architecture, but the underlying requirement is familiar: a system designed for multiple markets has to account for what people in each market expect to find.
Zalando addressed a related problem by adapting its assortment to regional expectations. It’s also the exception to the rule. Most organizations don’t get there until after the mismatch has already cost them customers, damages market share, or created some other sort of expensive business pain.
What Music, Libraries, And Online Shopping Have In Common
Put the three examples side-by-side and the pattern stops being about music, or libraries, or shopping. It starts being about decisions. A taxonomy looks like a neutral description of what exists. It’s not. Not even remotely. Building one means making decisions — and decisions aren’t neutral.
Someone has to decide what earns its own category, what gets folded into something else, and what gets left off the list entirely.
Dewey already made those decisions when he built out the 200s. Rahel's carrier ran into the same kind of decisions, made by default instead of on purpose. The taxonomy they inherited hadn’t been designed with any other market's categories in mind, and that gap only surfaced once the team started deciding which categories the platform would actually use.
Dewey classified the world’s knowledge based on what he knew and believed back in the late 1800s, and that understanding is still baked into the numbers today. Rahel’s team inherited categories created elsewhere. In Dewey’s case, the creator’s own perspective shaped the system. In Rahel’s project, another organization’s commercial categories arrived already embedded in the work.
You don’t notice a taxonomy’s blind spots until you try to make it work for people it wasn’t built for.
That’s exactly the position Rahel’s team found itself in. Their job was to solve the categorization side of making the platform work across markets, which meant mapping, adapting, or adding categories for each target market — categories the original taxonomy never had. Zalando addressed a related problem by adapting its assortment to regional expectations before customers ever ran into the mismatch.
Every organization that documents, categorizes, or governs information has made these same kinds of decisions, whether or not anyone meant to do so. Support ticket types. Customer segments. Product categories. Content types in your component content management system (CCMS).
Each of those schemas is built around an assumed “normal” case. Nobody likely wrote their assumption down, because the people who built the system already agreed on what normal meant.
Every example in this post so far has focused on a system that organizes information for people to find. The next one organizes information before a person ever sees it.
A Newer Version Of The Same Argument
There’s a newer version of this happening right now, and it’s the one with the highest stakes for anyone who documents things for a living. Increasingly, readers don’t read our docs directly. They ask an AI-powered answer engine a question, and the model that powers it answers using whatever knowledge it has access to.
Sometimes that means retrieving answers from our docs when a question is asked. In that case, labeling, structure, terminology, and other discoverability signals affect whether the system identifies our content as relevant.
Other times, the model answers from patterns learned during training. There, our product info and docs influence an answer only if it was included and retained in the training process in a form the model could learn from.
Either way, it’s the same argument. Metadata helps determine not just how your content is organized, but whether it’s findable, chosen, and understood at all — by a person or by an A-powered tool.
It’s important to note that the corpus of information that models learn from isn’t neutral either. Two influential earlier models illustrate the imbalance: GPT-3‘s training data was roughly 93 percent English, while LLaMA-2‘s was nearly 90 percent. CommonCrawl, one of the largest sources feeding models like these, is roughly 45 percent English — with every other language, including widely spoken ones, trailing far behind.
A 2024 study tested ChatGPT and Mixtral on music specifically. Researchers asked the models to identify leading musical contributors and evaluate the musical cultures of different countries. Both experiments revealed a strong preference for Western music cultures — the same kind of ethnocentrism Byrne criticized in 1999, now appearing in model outputs rather than record-store bins.
An AI training corpus raises a similar problem. It’s shaped by decisions about what to include, not a neutral sample of what exists. The underlying issue is close to what Rahel’s carrier ran into: a system built from one set of assumptions doesn’t necessarily work across every market — or every language — it intends to serve. Now it’s happening to anyone whose documentation gets read by a model instead of a person.
That’s what makes this everybody’s problem: the same decision keeps showing up wherever someone has to sort things into categories, which turns out to be nearly everywhere — a record label, a library, a retailer, and now anyone whose documentation gets read by a model instead of a person.
Rahel’s carrier found out because a new market made the gap impossible to ignore. Zalando addressed a related problem by adapting its assortment to regional expectations. Either way, someone eventually has to ask whether the taxonomy actually works for everyone it needs to serve. Organizations may never ask that question unless someone makes it part of their governance process.
The people closest to the taxonomy — tech writers, content strategists, and CCMS owners — are well positioned to notice those gaps before a customer, a market, or a model exposes them. Not every organization gives someone explicit responsibility for doing so. 🤠









