Work in progress
This page is being shared to explain the method. The library is still growing and the classification is re-run as it changes, so the exact counts here will keep evolving — treat them as indicative, not final.

Parochial to Cosmopolitan: Debiasing AI Clustering

One library of 1,232 books, classified three times. Each pass corrected a defect in the previous one.
This page covers classification accuracy only — what each book was filed as. Cost (cloud vs. local models, tokens) is measured separately. Terms: the lattice is the classification structure — a set of facets (Discipline, Form, Place, Time, and others), each branching into fields and then into leaves. A leaf is the most specific category a book can be assigned, such as "Cognitive psychology."

The bias in this library

Left to cluster the library on its own, the model grouped non-Western books by the author's identity rather than the book's content. Two cases caught our attention. In each, what the classifier saw is on the left; what the books are actually about is on the right.

What the classifier saw
Siddhartha Mukherjee
Siddhartha Mukherjee
Shelved under: India
What the books are about
GeneticsThe Gene
CancerThe Emperor of All Maladies
Cell biologyThe Song of the Cell
What the classifier saw
Taiichi Ohno
Taiichi Ohno
Shelved under: Japan
What the books are about
Lean manufacturingToyota Production System
Factory operationsTaiichi Ohno's Workplace Management
Filed by who wrote it, not what it says.
Portraits: Siddhartha Mukherjee — Moody College of Communication, CC BY-SA 2.0; Taiichi Ohno — Toyota Motor Corp., public domain. Both via Wikimedia Commons.

In both, the book's argument was discarded in favour of who wrote it and where they were from. The same pull collapsed the whole Hindu tradition into a single group while Western material was divided finely.

Those two were failures of the first pass. Here is the flip side — a book the corrected lattice places well, which also previews the fix the rest of this page explains.

How Champions Think
One book, corrected
How Champions Think — a book on the psychology of elite performance
Before the fix: — no discipline —
After the fix: Sport & performance psychologyJudgment & decision-making
A different kind of case: not identity bias but a missing branch. With no Psychology in the lattice, the book could not be classified at all. Once the branch was added, it found its shelf — Sport & performance psychology.

The bias in every library

We then checked whether an existing standard system would fix this. It would not — the same bias is built into the tools. In Dewey Decimal, the world's dominant scheme for over a century, Christianity occupies the range 200–289 (about ninety numbers); every Indic religion shares the single number 294. Library of Congress encodes the same imbalance.

Dewey Decimal · the 200s · shelf space vs. followers

How many numbers each religion gets — and how many people it has

A “box” is one three-digit Dewey number. Christianity holds the whole 200–289 range; every other religion on earth is given a single number — or, like Shinto, only a decimal — no matter how many people follow it.

Religion
Numbers in Dewey (one box = one number)
Followers
ChristianityDewey 200–289
90
2.3 billion followers
IslamDewey 297Qur’an 297.1 · theology 297.2 · law 297.5 · Sufism 297.4 · Sunni & Shia 297.8 — all decimals of one number
1
2.0 billion followers
Religions of Indic originDewey 294Hinduism, Buddhism, Jainism & Sikhism share it
1
~1.5 billion followers
ShintoDewey 299.56no whole number — a decimal inside 299, “not provided for elsewhere”
0
~100 million* followers
JudaismDewey 296
1
~15 million followers
ZoroastrianismDewey 295
1
~200,000 followers

Of 94 Dewey numbers across these traditions, 90 go to Christianity. Islam and the Indic religions each get one; Shinto gets no whole number at all. The shelf space a religion receives tracks the cataloguer's world — not the world's people.

Boxes are three-digit numbers in DDC 23: Christianity 200–289; Islam 297; religions of Indic origin 294; Judaism 296; Zoroastrianism 295; Shinto has no whole number — only the decimal 299.56, inside 299 (“religions not provided for elsewhere”). Islam's own subfields likewise sit in decimals of 297. Followers: Pew Research Center global estimates, 2020 (rounded; Indic combines Hinduism, Buddhism, Jainism, Sikhism). *Shinto is hard to count — tens of millions practise in Japan, often alongside Buddhism.

The deeper form: the West is the unmarked default — "life sciences," "modern law," "medicine" carry no cultural tag and read as neutral — while every other tradition is marked as particular: "Traditional Chinese Medicine," "Hindu law," "Confucian education." The West disappears into "universal," and everyone else becomes a labelled exception to it.

The deeper asymmetry · who has to be labelled
The West — no cultural tag
Every other tradition — marked
Medicine
Traditional Chinese Medicine
Modern law
Hindu law
Education
Confucian education
Same field, same rigour. The West's version carries no label and reads as the neutral default; every other tradition is marked as a particular exception to it.

The model was re-instantiating a bias embedded in library classification for a century.

Why we did not correct it by hand

Two quick fixes were available. We rejected both, because each substitutes one bias for another.

Personal priming. The curator (Bharat) could hand-correct each misfiling from his own knowledge of the books. But a personal scheme encodes the curator's own experience and blind spots. Every classification reflects its author's worldview — Dewey encodes Christianity; a Marxist library foregrounds Marxism; an individual foregrounds the subjects he knows best. Correcting by hand would replace the model's bias with the curator's, not remove the bias. So we deliberately kept personal priming out of the method.

A single external standard. Adopting Dewey, Library of Congress or BISAC wholesale imports that system's own century-old asymmetry. No off-the-shelf scheme is neutral.

The method

Cosmopolitan by construction.

The approach rests on one finding from surveying how the world classifies: every culture that classified itself achieved parity — fine, even resolution — either by faceting (India's Colon classification, Belgium's UDC, Britain's Bliss) or by deep local enumeration (Japan's NDC). The coarseness appears only when one tradition classifies another.

Painted portrait of S. R. Ranganathan
Inspiration — S. R. Ranganathan (1892–1972). The Indian mathematician and librarian who invented faceted classification — the Colon Classification, 1933. His idea, to describe a book by a few fundamental facets rather than file it in one place, is the backbone of the lattice. Decades later, Italy's Nuovo Soggettario settled on almost the same facets for a different language and culture — a sign the axes are close to universal.
Painting: Bhaskar Shikha Saikia, CC BY-SA 4.0 · Wikimedia Commons (cropped)

So the method is a parity-audited mosaic: source each tradition's categories from its own native structure, mark the West as one tradition among many rather than the default, and run a parity audit — a measurable check that non-Western material is divided as finely as Western material, sub-dividing until it is. The aim is not to flatten distinctions but the opposite: to give every tradition the same fineness of clustering, none left coarse. The bias is removed by construction, not by a human hand applied after the fact.

What follows is how that played out in practice, in three stages.

1

Stage 1 — grouped by identity, not content

Initial clustering engine
The initial engine sorted the library into 321 groups. As shown above, it filed non-Western books under identity labels — the author's origin or a broad theme — rather than by content. The largest of those identity groups are listed below.
321
groups
~60
books filed as "Indian"
The identity labels (books grouped by author origin, with cluster sizes)
22 Indian-Spirituality-Classic
22 Indian-Spirituality-Modern
19 Buddhist-Mindfulness-Classic
19 Buddhist-Mindfulness-Modern
19 Japanese-Aesthetics-Classic
Problem: books were grouped by author identity, not content. Non-Western cultures were compressed into a few large buckets; Western material was divided finely by discipline.
↓   Diagnosed the bias. Built a balanced lattice that classifies by content.   ↓
2

Stage 2 — bias gone, gaps remained

Content-based classification, before the final revision
The lattice classifies each book by content, across balanced facets, so non-Western material is divided as finely as Western material. This removed the identity bias. But the lattice was incomplete: it had no Psychology branch and no Design branch, so books in those fields had no valid category.
326
clusters
0
Psychology / Design clusters
139
books with no discipline
The remaining defect was coverage, not bias. Several modern fields were missing from the lattice, so their books were left unclassified.
↓   Added the Psychology and Design branches and a country tier.   ↓
3

Stage 3 — the gaps filled

After the final revision
The missing branches were added to the lattice. 231 books moved to a Psychology or Design category, and those fields now form clusters that did not exist before. A country tier was added, so a book about Japan can be filed under "Japan" rather than only a sub-region.
361
clusters
20
Psychology clusters
9
Design clusters
231
books rehomed
Psychology — clusters that did not exist a day earlier
53 Judgment & decision-making
30 Positive psychology & well-being
29 Social psychology
27 Cognitive psychology
17 Sport & performance psychology
13 Personality psychology
Design — likewise new
18 Design theory & criticism
11 Fashion & apparel design
9 Design thinking & methods
9 Design history
8 Interior & spatial design
7 Industrial & product design
The lattice is now unbiased and complete: books are classified by content, and every field has a category.
Next: the same clustering task, run through seven AI models  →  How seven AI models cluster one library
Antilibrary · classification accuracy. Computed from three preserved states: bk/clusters.json (stage 1), the 2026-07-11 combined baseline (stage 2), and its v2 (stage 3). All classification on this page was done by one model (Claude). A separate study runs the same task through seven models; see How seven AI models cluster one library. Built 2026-07-11.