Six probes went out to look at one question — how should an engine choose what a personal antilibrary adds and subtracts — each from a different lens. What each probe found is filed below, unedited and unranked.
Nothing here is organized, merged, or ordered by importance. Each block is one probe's own report in its own words. They are listed by lens, not by rank: the first four are the lenses you named (artistic, scientific, mathematical, fractal); the last two were added to the sweep (economic / via-negativa, ecological / living-systems). A menu of lenses not yet probed sits at the end.
What this lens uniquely contributes. The engineer optimizes objects; the curator composes relationships and intervals. To an artist, a collection is not a set of items but a single hung composition — and the most charged part is the empty space between the works. This lens gives Khoj three things the graph-theory lenses miss: it can value the gap as much as the node, it can recommend removal as a creative act rather than a failure, and it treats surprise as something you engineer through constraint, not something you wait for.
Sources: Warburg, Law of the Good Neighbour (Cornell) · Iconology of the Interval · AAMD Deaccessioning Policy · Yohaku-no-bi & the art of subtraction · Shakkei / borrowed scenery
What this lens uniquely contributes. RecSys research long ago proved that optimizing for predicted relevance alone produces a degenerate, self-narrowing list — the "beyond-accuracy" literature exists precisely to formalize the objectives an antilibrary actually cares about (novelty, diversity, serendipity, coverage). It hands Khoj exact, computable objective functions rather than intuitions. The cognitive-science half (active learning, Bayesian surprise, information foraging, learning-progress curiosity) supplies the missing epistemic criterion: not "what will you like?" but "what will teach you the most, and when should you leave a topic?" Together they let Khoj rank ADD/SUBTRACT by measurable expected-learning, not taste.
D_KL(posterior ‖ prior). It beat 11 other metrics at predicting human gaze. Mechanism: score a candidate book by how much its idea-units would shift your cluster/thread distribution if ingested — high expected posterior-shift = high acquisition priority. The rigorous form of "maximize expected learning."−log₂ p(item); intra-list diversity penalizes near-duplicates. Mechanism: demote books whose idea-units are near-duplicates of higher-signal books already held (low marginal self-information, high pairwise similarity) — the book adds no novelty to its cluster.R_unexp = R \ PM_u — but the item must still be relevant. Mechanism: build a trivial "obvious next book" model (same author/cluster); surface only candidates outside it that still connect to a live thread. Surprise with a reason to exist.Key sources: Vargas & Castells 2011, Steck 2018, Itti & Baldi 2009, Pirolli & Card 1999, Adamopoulos & Tuzhilin 2014, Schmidhuber compression progress, popularity-bias survey.
What this lens uniquely contributes. It reframes "what to buy / what to cull" as subset selection in embedding space with provable guarantees. Where the taste- and narrative-lenses argue, this lens computes: it turns the ~5,700 idea-units + embeddings + tag graph + 321 clusters into objective functions whose optima are the ADD and SUBTRACT lists, and it supplies the one operator the others lack — a tunable dial between coverage (relevance) and dispersion (surprise). The deep move: ADD is submodular coverage/gap-filling, SUBTRACT is submodular pruning/coreset, and serendipity is a stochastic perturbation of the same objective — one math, three signs.
The question is "which k books maximally extend what the antilibrary knows about." Candidate books are embedded (title/description/back-cover → same space as idea-units) and their tag-facets projected onto the tag universe.
f(S) = weighted tag-facets/concept-regions newly covered by adding book-set S, subject to budget k. Greedy — repeatedly add the book with largest marginal gain f(S∪{b})−f(S) — gives the tight (1−1/e) ≈ 0.63 guarantee. Cost: O(k·N·cost(f)); lazy/CELF evaluation makes it near-linear in practice. This is the spine.f: facility-location Σ_i max_{b∈S} sim(i,b) rewards a book that becomes the nearest representative of currently-underserved idea-units; max-k-cover rewards raw tag-universe hit-count. Both inherit the (1−1/e) bound.arg max λ·rel(b) − (1−λ)·max_{j∈S} sim(b,j). O(kN), λ is the relevance-vs-novelty knob surfaced to the user.The question inverts: "which owned books are redundant representatives we can drop with least coverage loss." Protect the informative frontier, discard the interior duplicates.
f (another book covers its idea-units) → safe to cull. Marginal-coverage-loss ranking, same greedy machinery.det(L_S); the kernel L encodes quality×diversity, so draws are provably anti-redundant and stochastic — the principled alternative to greedy MMR. Exact sampling O(n³) eigendecomp then cheap draws; k-DPP fixes the count.Sources: Kulesza & Taskar, DPPs for ML · Nemhauser–Wolsey–Fisher greedy (1−1/e) · Persistent homology / TDA
What this lens uniquely contributes. Every other lens ranks items; this one measures the shape of the coverage itself and treats gaps as objects with a size, a location, and a scale. A library's grasp of concept-space is a point set; you can box-count that set at many resolutions to ask not just "what's missing" but "at which zoom level is the collection thinnest, and is its texture smooth or clumpy?" The payoff is a single scale-resolved gap map that makes ADD, SUBTRACT, and SERENDIPITY fall out of one measurement instead of three heuristics.
N(ε) ≈ C·ε^(−d). Khoj mechanism: at each scale, the empty boxes adjacent to occupied ones are ranked ADD candidates; the scale where the occupied/total ratio drops fastest is where the collection is most porous, so weight acquisitions there. A missing domain and a missing sub-subject are the same operation at different ε.Book → Book + [adjacent sub-subject seed]. Khoj mechanism: each turn, apply one production to the current thinnest branch — organic, rule-driven growth that reaches new territory without a human picking every node.Sources: Minkowski–Bouligand (box-counting) dimension; Scale-free networks & preferential attachment.
What this lens uniquely contributes. This is the antilibrary's own bloodline: Eco→Taleb, the unread book as the highest-value asset. Every other lens treats a book as content to be matched; this lens treats each acquisition as a financial position — an option with an asymmetric payoff — and the whole library as a portfolio of bets under irreversible uncertainty. Its deepest gift is the SUBTRACT limb: it is the only lens with a rigorous, named account of why letting go feels wrong (loss aversion, endowment, sunk cost) and how to engineer around it. It reframes the scarce resource not as shelf space but as attention — the one budget that cannot be expanded.
Sources: Antifragile / barbell, Klotz, Subtract, Lindy effect, Endowment effect / loss aversion (Kahneman-Knetsch-Thaler), Dixit-Pindyck, Investment under Uncertainty, Markowitz efficient frontier, Secretary problem / 37% rule.
What this lens uniquely contributes. Every other framing treats the antilibrary as an information store to optimize. Ecology treats it as a living community with a carrying capacity, a successional stage, and a health that depends on diversity, not size. That reframes both lists: ADD is colonization of empty niches and keystone introduction; SUBTRACT is weeding and thinning for the health of the whole, not tidying. Crucially, ecology supplies the only field-tested, professionally-codified SUBTRACT methodology in existence — the librarians' CREW/MUSTIE protocol — plus a far richer diagnostic toolkit (evenness, richness, rarity) than plain Shannon entropy.
The library-science backbone. CREW (Continuous Review, Evaluation & Weeding; Texas State Library, 1976) uses the MUSTIE test. Map each criterion to an antilibrary signal:
Two ecological overlays on top of MUSTIE:
Sources: CREW/MUSTIE — Texas State Library CREW Manual; Hutchinson niche & competitive exclusion — Wikipedia "Vacant niche" / "Ecological niche"; keystone species — Paine 1966, Nature Scitable & American Naturalist; r/K — MacArthur & Wilson via Wikipedia "R/K selection theory"; indices — Bio LibreTexts "Diversity Indices," Pielou's J′.
These have not been run — no findings exist for them yet. They are candidates, listed so you can pick which to send a probe into. Say the word on any and I'll dispatch it the same way as the six above.
Neti-neti ("not this, not that") as the original via-negativa; aparigraha (non-possession) as a subtract ethic; Manthan (the churning of the ocean) as the project's own metaphor; darshan and guru–shishya as discovery. On-brand with the Sanskrit naming.
How canons and lineages actually form; Harold Bloom's "anxiety of influence"; memetics and the sociology of the citation; what the last thinker in a school owes the first — acquisition as joining a live intellectual descent.
Forgetting curves, spaced repetition, memory consolidation, the generation effect. Bears directly on the hardest subtract question: what to reread, what to demote once internalized, what a personal library owes to a finite memory.
The spec's own proactive-gap question made into a lens: "what is the book that would most disturb the existing structure?" Steelman the opposite library; dialectic; buy your strongest disagreement, not your next agreement.
What other people's libraries reveal about the gaps in yours; collaborative filtering done ethically (not popularity); the library as gift, inheritance, and shared object rather than a solo optimization.
The collection as memoir — the story the add and subtract lists tell about who the owner is becoming. Bibliomemoir, the shelf as self-portrait; acquisitions chosen for the arc, not only the coverage.
Six probes filed above; six more lenses waiting. None of it has been ranked, merged, or reconciled — that's yours to do when you've looked through.