Left to cluster the library on its own, the model grouped non-Western books by the author's identity rather than the book's content. Two cases caught our attention. In each, what the classifier saw is on the left; what the books are actually about is on the right.





In both, the book's argument was discarded in favour of who wrote it and where they were from. The same pull collapsed the whole Hindu tradition into a single group while Western material was divided finely.
Those two were failures of the first pass. Here is the flip side — a book the corrected lattice places well, which also previews the fix the rest of this page explains.
We then checked whether an existing standard system would fix this. It would not — the same bias is built into the tools. In Dewey Decimal, the world's dominant scheme for over a century, Christianity occupies the range 200–289 (about ninety numbers); every Indic religion shares the single number 294. Library of Congress encodes the same imbalance.
Dewey Decimal · the 200s · shelf space vs. followers
A “box” is one three-digit Dewey number. Christianity holds the whole 200–289 range; every other religion on earth is given a single number — or, like Shinto, only a decimal — no matter how many people follow it.
Of 94 Dewey numbers across these traditions, 90 go to Christianity. Islam and the Indic religions each get one; Shinto gets no whole number at all. The shelf space a religion receives tracks the cataloguer's world — not the world's people.
Boxes are three-digit numbers in DDC 23: Christianity 200–289; Islam 297; religions of Indic origin 294; Judaism 296; Zoroastrianism 295; Shinto has no whole number — only the decimal 299.56, inside 299 (“religions not provided for elsewhere”). Islam's own subfields likewise sit in decimals of 297. Followers: Pew Research Center global estimates, 2020 (rounded; Indic combines Hinduism, Buddhism, Jainism, Sikhism). *Shinto is hard to count — tens of millions practise in Japan, often alongside Buddhism.
The deeper form: the West is the unmarked default — "life sciences," "modern law," "medicine" carry no cultural tag and read as neutral — while every other tradition is marked as particular: "Traditional Chinese Medicine," "Hindu law," "Confucian education." The West disappears into "universal," and everyone else becomes a labelled exception to it.
The model was re-instantiating a bias embedded in library classification for a century.
Two quick fixes were available. We rejected both, because each substitutes one bias for another.
Personal priming. The curator (Bharat) could hand-correct each misfiling from his own knowledge of the books. But a personal scheme encodes the curator's own experience and blind spots. Every classification reflects its author's worldview — Dewey encodes Christianity; a Marxist library foregrounds Marxism; an individual foregrounds the subjects he knows best. Correcting by hand would replace the model's bias with the curator's, not remove the bias. So we deliberately kept personal priming out of the method.
A single external standard. Adopting Dewey, Library of Congress or BISAC wholesale imports that system's own century-old asymmetry. No off-the-shelf scheme is neutral.
Cosmopolitan by construction.
The approach rests on one finding from surveying how the world classifies: every culture that classified itself achieved parity — fine, even resolution — either by faceting (India's Colon classification, Belgium's UDC, Britain's Bliss) or by deep local enumeration (Japan's NDC). The coarseness appears only when one tradition classifies another.
So the method is a parity-audited mosaic: source each tradition's categories from its own native structure, mark the West as one tradition among many rather than the default, and run a parity audit — a measurable check that non-Western material is divided as finely as Western material, sub-dividing until it is. The aim is not to flatten distinctions but the opposite: to give every tradition the same fineness of clustering, none left coarse. The bias is removed by construction, not by a human hand applied after the fact.
What follows is how that played out in practice, in three stages.
bk/clusters.json (stage 1), the 2026-07-11 combined baseline (stage 2), and its v2 (stage 3).
All classification on this page was done by one model (Claude). A separate study runs the same task through seven
models; see How seven AI models cluster one library.
Built 2026-07-11.