Antilibrary Engine · Khoj खोज — deep research outputs

Khoj Deep Research Outputs

Six probes went out to look at one question — how should an engine choose what a personal antilibrary adds and subtracts — each from a different lens. What each probe found is filed below, unedited and unranked.

Nothing here is organized, merged, or ordered by importance. Each block is one probe's own report in its own words. They are listed by lens, not by rank: the first four are the lenses you named (artistic, scientific, mathematical, fractal); the last two were added to the sweep (economic / via-negativa, ecological / living-systems). A menu of lenses not yet probed sits at the end.

Probe 01 · Artistic

The curatorial / artistic eye

The lens of the serious artist and curator: composition, the interval, deaccession, and engineered surprise.

What this lens uniquely contributes. The engineer optimizes objects; the curator composes relationships and intervals. To an artist, a collection is not a set of items but a single hung composition — and the most charged part is the empty space between the works. This lens gives Khoj three things the graph-theory lenses miss: it can value the gap as much as the node, it can recommend removal as a creative act rather than a failure, and it treats surprise as something you engineer through constraint, not something you wait for.

Add — what the curatorial eye acquires
  • Warburg's "Law of the Good Neighbour." Warburg shelved by elective affinity, not classification — his creed was "the book one knew was rarely the book one needed." Mechanism: for each cluster, Khoj proposes the book that would sit best beside it — highest affinity to the shelf, but from a different existing class. Recommend by desired neighbour, not by category match.
  • Shakkei (borrowed scenery). A Japanese garden frames a distant mountain it doesn't own, making the view part of the composition. Mechanism: a "borrowed view" tier — surface canonical works just outside the collection's territory that a thread already points at, so the antilibrary annexes an adjacent field through one framing acquisition rather than filling it in.
  • The anchoring counterpoint / completing the composition. Curators buy the one piece that resolves a wall's imbalance. Mechanism: find the cluster whose internal tension is unresolved (one strong voice with no interlocutor) and recommend its antithesis — the book that makes the argument two-sided.
  • Wunderkammer juxtaposition. The cabinet of curiosities derived meaning from a nautilus shell beside an astrolabe. Mechanism: deliberately rank a small quota of high-surprise cross-domain pairings — books whose only link is a shared thread, not a shared cluster.
  • Yohaku-no-bi (the beauty of the deliberately unfilled). Mechanism: not every gap is a buy signal. Tag some diversity-holes "leave blank" — negative space the collection is meant to hold — so Khoj recommends against completing certain axes.
Subtract — how the artist decides what to cut
  • Deaccession criteria (AAMD). Museums remove on defined grounds: redundant duplicate in a series, poor quality / no study value, false attribution — and, critically, never for market price. Mechanism: Khoj's subtract score keys on redundancy-within-cluster and low idea-unit yield — never on resale value or shelf cost.
  • "Kill your darlings" (Arthur Quiller-Couch). The line you love most is often the one bloating the whole. Mechanism: flag over-loved zones — clusters far denser than their thread-connectivity justifies — as prime pruning candidates, precisely because attachment hides the redundancy.
  • The piece that no longer converses. A curator cuts the work that has stopped talking to its neighbours. Mechanism: subtract books whose bridge-betweenness has decayed to zero — nodes that connect to nothing, orphaned by the collection's growth around them.
  • Pruning for the whole, not the specimen. A gardener cuts a healthy branch so the tree reads. Mechanism: rank subtractions by marginal gain to composition (diversity, legibility, thread-clarity), not by the book's own weakness — remove what deadens the ensemble even if it is individually fine.
Serendipity — engineering the happy accident
  • Cage / Duchamp chance operations. Cage used the I Ching to compose; Duchamp let dropped threads (3 Standard Stoppages) become the ruler. Mechanism: a chance operator that periodically picks a random cluster-pair and asks Khoj to justify a bridge book between them — structured randomness as a recommendation source.
  • Exquisite corpse (Surrealist). Each contributor adds blind to the others; the join is the discovery. Mechanism: generate a candidate by chaining three threads without checking coherence first, then surface the improbable book at the seam.
  • Constraint as serendipity generator (Oulipo / Sol LeWitt). LeWitt's instructions and Oulipo's lipograms prove a tight rule forces novelty. Mechanism: run recommendation rounds under a forced constraint ("only pre-1900," "only translated," "only under-represented axis") — the constraint drives Khoj somewhere its default optimizer never would.
  • The generative mis-shelving. The wrong shelf reveals a neighbour you'd never have sought. Mechanism: occasionally mis-file a candidate into a foreign cluster and show what it illuminates there — deliberate error as a discovery probe.
The sharpest single idea

Warburg's iconology of the interval (Ikonologie des Zwischenraums). In the Mnemosyne Atlas, Warburg argued the image "has no intrinsic meaning — it is the encounter between images, their spacing, that produces sense." Meaning lives in the Zwischenraum, the gap between works. This is the deepest gift to Khoj: the owner's real assets are the 108 threads — the intervals — not the 1,250 books. So Khoj should rank both ADD and SUBTRACT by their effect on the intervals: acquire the book that makes a thread suddenly cohere; cut the book that clutters the space between two clusters that want to speak. Optimize the gaps, and the books arrange themselves.

Sources: Warburg, Law of the Good Neighbour (Cornell) · Iconology of the Interval · AAMD Deaccessioning Policy · Yohaku-no-bi & the art of subtraction · Shakkei / borrowed scenery

Probe 02 · Scientific

Recommender science & the cognitive science of discovery

Beyond-accuracy recommender research plus active learning, Bayesian surprise, and information foraging.

What this lens uniquely contributes. RecSys research long ago proved that optimizing for predicted relevance alone produces a degenerate, self-narrowing list — the "beyond-accuracy" literature exists precisely to formalize the objectives an antilibrary actually cares about (novelty, diversity, serendipity, coverage). It hands Khoj exact, computable objective functions rather than intuitions. The cognitive-science half (active learning, Bayesian surprise, information foraging, learning-progress curiosity) supplies the missing epistemic criterion: not "what will you like?" but "what will teach you the most, and when should you leave a topic?" Together they let Khoj rank ADD/SUBTRACT by measurable expected-learning, not taste.

Add — what science says to acquire
  • Active learning / uncertainty sampling (Settles, 2009). Acquire the sample the model is least certain about, because it reduces model uncertainty most. Mechanism: treat each of the 321 clusters as a region where your "model" (thread coherence) is confident or shaky. Acquire books that land in high-entropy zones — clusters that are sparse, internally contradictory, or that a thread reaches toward but can't yet close. The formal version of "buy the book that resolves the open question."
  • Query-by-Committee (Seung / Freund; Settles). Informativeness = disagreement among an ensemble. Mechanism: run several clustering/thread models; books that would be classified differently by different models (a book two threads both claim) are maximally informative bridge acquisitions — buy where your own models disagree.
  • Bayesian surprise (Itti & Baldi, 2005 / 2009). Surprise = KL divergence between prior and posterior belief, D_KL(posterior ‖ prior). It beat 11 other metrics at predicting human gaze. Mechanism: score a candidate book by how much its idea-units would shift your cluster/thread distribution if ingested — high expected posterior-shift = high acquisition priority. The rigorous form of "maximize expected learning."
  • Marginal Value Theorem (Charnov 1976) via Information Foraging (Pirolli & Card 1999). MVT says leave a patch when its within-patch gain rate drops below the environment's average rate. Mechanism: track diminishing returns per cluster ("patch"); when the marginal idea-unit yield from a well-stocked cluster falls below your library-wide average, the patch-leaving signal fires → the ADD list should point to a fresh, under-foraged patch, not more of the saturated one. The "information scent" model also says rank candidates by proximal cues (title, citations) predicting downstream value.
  • Coverage / low-density regions (Herlocker et al. 2004; Adomavicius on aggregate diversity). Acquire into catalog regions with near-zero coverage — the literal empty shelves in your 321-cluster map.
Subtract — what science says to remove
  • Redundancy / low marginal self-information (Vargas & Castells, RecSys 2011). Novelty is formalized as self-information −log₂ p(item); intra-list diversity penalizes near-duplicates. Mechanism: demote books whose idea-units are near-duplicates of higher-signal books already held (low marginal self-information, high pairwise similarity) — the book adds no novelty to its cluster.
  • Popularity-bias overweighting (Klimashevskaia et al. 2024; Anderson, The Long Tail, 2006). Algorithms systematically over-hold head items. Mechanism: flag over-represented canonical/popular titles that crowd the head while contributing little unique coverage — candidates to deaccession or demote so long-tail books surface.
  • Over-exploited patches (MVT again). The same patch-leaving math that drives ADD also drives SUBTRACT: a cluster far past its marginal-value point is over-invested — trim its weakest members.
  • Schmidhuber's boredom criterion (Driven by Compression Progress, 2008). What yields no compression progress — neither random-noise nor already-fully-learned — is worthless to keep foregrounded. Books whose ideas you've fully internalized (learning-progress ≈ 0) drop off the active mirror.
Serendipity — productive surprise, not noise
  • Serendipity = unexpectedness × relevance (Ge et al. 2010; Adamopoulos & Tuzhilin, 2014). The formal guard against random noise: unexpectedness is measured relative to a "primitive prediction model" PM — R_unexp = R \ PM_u — but the item must still be relevant. Mechanism: build a trivial "obvious next book" model (same author/cluster); surface only candidates outside it that still connect to a live thread. Surprise with a reason to exist.
  • Calibrated recommendation (Steck, RecSys 2018). Match the genre/topic distribution of recommendations to the user's own distribution via KL divergence — diversity that respects your actual profile instead of injecting noise. Mechanism: keep the ADD list's cluster-mix calibrated to your library's shape, then perturb deliberately — controlled, not random, novelty.
  • Learning-progress curiosity (Schmidhuber; Oudeyer & Kaplan). Reward the first derivative of compressibility — target the zone of intermediate difficulty where learning rate is maximal. Mechanism: recommend books at the edge of your current threads — connected enough to be learnable, novel enough to move the compression frontier. The antidote to filter bubbles (Nguyen et al. 2014 showed recommenders narrow content diversity): curiosity as directed exploration, not entropy injection.
Sharpest single idea

Frame Khoj as an active-learning acquirer for your own knowledge model: score every candidate by expected Bayesian surprise (posterior-vs-prior KL over your cluster/thread structure), gate it with unexpectedness-relative-to-a-primitive-model × relevance so surprise stays productive, and let Marginal Value Theorem patch-leaving decide when an ADD should jump to a fresh topic and a SUBTRACT should trim a saturated one. One unifying currency — expected shift in your knowledge model — drives ADD, SUBTRACT, and serendipity alike.

Key sources: Vargas & Castells 2011, Steck 2018, Itti & Baldi 2009, Pirolli & Card 1999, Adamopoulos & Tuzhilin 2014, Schmidhuber compression progress, popularity-bias survey.

Probe 03 · Mathematical

The optimization core

Subset selection with provable guarantees: submodular coverage, DPPs, topological gap-finding, optimal stopping.

What this lens uniquely contributes. It reframes "what to buy / what to cull" as subset selection in embedding space with provable guarantees. Where the taste- and narrative-lenses argue, this lens computes: it turns the ~5,700 idea-units + embeddings + tag graph + 321 clusters into objective functions whose optima are the ADD and SUBTRACT lists, and it supplies the one operator the others lack — a tunable dial between coverage (relevance) and dispersion (surprise). The deep move: ADD is submodular coverage/gap-filling, SUBTRACT is submodular pruning/coreset, and serendipity is a stochastic perturbation of the same objective — one math, three signs.

Add — proposing acquisitions

The question is "which k books maximally extend what the antilibrary knows about." Candidate books are embedded (title/description/back-cover → same space as idea-units) and their tag-facets projected onto the tag universe.

  • Submodular maximization, greedy (Nemhauser–Wolsey–Fisher 1978). Maximize a monotone submodular coverage f(S) = weighted tag-facets/concept-regions newly covered by adding book-set S, subject to budget k. Greedy — repeatedly add the book with largest marginal gain f(S∪{b})−f(S) — gives the tight (1−1/e) ≈ 0.63 guarantee. Cost: O(k·N·cost(f)); lazy/CELF evaluation makes it near-linear in practice. This is the spine.
  • Facility-location & max-coverage objectives — two concrete submodular f: facility-location Σ_i max_{b∈S} sim(i,b) rewards a book that becomes the nearest representative of currently-underserved idea-units; max-k-cover rewards raw tag-universe hit-count. Both inherit the (1−1/e) bound.
  • Gap-geometry, for where the shelf is empty. Compute largest-empty-ball in the embedding cloud (via Delaunay/Voronoi — the biggest empty circumsphere sits at a Voronoi vertex; O(n log n) in low-dim, approx in high-dim) and kernel-density low-density regions. Rank candidate books by how well they land in a genuine void rather than a crowded region.
  • Persistent homology (TDA), the sharpest gap-finder. Build a Vietoris–Rips filtration over idea-units; H1 loops / H2 voids with high persistence are real topological holes — a ring of adjacent topics with a hollow center (e.g., ideas that circle a subject no owned book crosses). Rank candidates by their power to fill or close a persistent hole. Cost: worst-case super-cubic; use spectral/landmark (witness-complex) reductions on ~5,700 points.
  • MMR (Carbonell & Goldstein 1998) as the cheap online ranker within the shortlist: arg max λ·rel(b) − (1−λ)·max_{j∈S} sim(b,j). O(kN), λ is the relevance-vs-novelty knob surfaced to the user.
Subtract — proposing removals

The question inverts: "which owned books are redundant representatives we can drop with least coverage loss." Protect the informative frontier, discard the interior duplicates.

  • Submodular pruning / coreset (prototype) selection. Find the subset that preserves facility-location coverage; the complement are removal candidates. A book is a redundant point if removing it barely lowers f (another book covers its idea-units) → safe to cull. Marginal-coverage-loss ranking, same greedy machinery.
  • k-medoids (PAM). Medoids are the irreplaceable representatives to keep; non-medoids whose loss is small are demotion candidates. Distinguishes the two removal types cleanly.
  • Redundancy via pairwise similarity / mutual information — high max-similarity to a retained neighbor = redundant. CUR / column-subset-selection (leverage scores) on the book×idea matrix: low-leverage columns are linearly explained by others → deaccession; high-leverage columns are structurally unique → protect.
  • The outlier-protection rule (antilibrary-specific). k-medoids/LOF separate redundant (near a medoid) from outlier (far from everything). Standard pruning discards both; Khoj must protect outliers — a lone book far from every cluster is the antilibrary's whole point (unknown-unknown). So: cull high-similarity/low-leverage redundant points, shield low-density outliers. This inversion is the lens's signature.
Serendipity — controlled surprise
  • Random-walk-with-restart on the tag/idea graph. From owned clusters, random-walk with restart probability α; books reached at low stationary probability but non-zero reachability = "surprising yet connected." Tunes surprise via α. Cost: sparse power-iteration.
  • Determinantal Point Processes (Kulesza & Taskar 2012). Sample a diverse set with probability ∝ det(L_S); the kernel L encodes quality×diversity, so draws are provably anti-redundant and stochastic — the principled alternative to greedy MMR. Exact sampling O(n³) eigendecomp then cheap draws; k-DPP fixes the count.
  • Entropy-regularized / information-directed selection. Add τ·H(S) to the objective and Boltzmann-sample; τ dials exploration continuously.
  • Bandits, formalized. Treat clusters/facets as arms; Thompson sampling (sample from posterior over each facet's expected value) or UCB (mean + √(2ln t / n)) to allocate the acquisition budget across explore/exploit; Gittins index is the exact optimum for the discounted case.
  • Secretary / optimal stopping (37% rule) for buy-now-vs-wait on a single book seen in a stream: reject the first ~n/e, then take the next best-so-far — near-optimal single-choice policy.
Sharpest single idea

One submodular coverage objective f runs the whole engine, and the sign of its marginal decides everything. ADD = books with the largest positive marginal gain f(S∪b)−f(S) (with persistent-homology holes weighting the gaps); SUBTRACT = owned books with the smallest marginal loss f(S)−f(S∖b), except protected outliers; SERENDIPITY = don't take the argmax — sample proportionally (DPP / entropy-regularized), letting a temperature τ inject exactly as much surprise as the curator wants. Coverage, pruning, and surprise become one function read three ways.

Sources: Kulesza & Taskar, DPPs for ML · Nemhauser–Wolsey–Fisher greedy (1−1/e) · Persistent homology / TDA

Probe 04 · Fractal

Scale-invariance & the shape of coverage

Box-counting the collection at every zoom, preferential attachment, L-system growth, gaps at all scales.

What this lens uniquely contributes. Every other lens ranks items; this one measures the shape of the coverage itself and treats gaps as objects with a size, a location, and a scale. A library's grasp of concept-space is a point set; you can box-count that set at many resolutions to ask not just "what's missing" but "at which zoom level is the collection thinnest, and is its texture smooth or clumpy?" The payoff is a single scale-resolved gap map that makes ADD, SUBTRACT, and SERENDIPITY fall out of one measurement instead of three heuristics.

Add — acquisitions from the empty boxes
  • Box-counting the coverage (Minkowski–Bouligand dimension). Embed the ~5,700 idea-units in the nested tree (domain → subject → sub-subject → book → unit). At each scale ε — coarse cells for domains, fine cells for sub-subjects — count occupied boxes N(ε). The box-counting dimension is the slope d where N(ε) ≈ C·ε^(−d). Khoj mechanism: at each scale, the empty boxes adjacent to occupied ones are ranked ADD candidates; the scale where the occupied/total ratio drops fastest is where the collection is most porous, so weight acquisitions there. A missing domain and a missing sub-subject are the same operation at different ε.
  • Feed the tail against preferential attachment (Barabási–Albert). Collections grow "rich-get-richer": the cluster you already own most of attracts the next book (degree ∝ existing degree, γ≈3 for scale-free graphs). Khoj mechanism: compute the cluster-size distribution; flag Zipf-head clusters as saturated and down-weight their marginal ADD score, while boosting singleton/doubleton clusters — deliberately anti-preferential acquisition that flattens the exponent and thickens the tail.
  • Grow by an L-system rule (Lindenmayer). Treat acquisition as a production grammar: a book with n idea-units in an under-covered sub-subject "expands" per a rule like Book → Book + [adjacent sub-subject seed]. Khoj mechanism: each turn, apply one production to the current thinnest branch — organic, rule-driven growth that reaches new territory without a human picking every node.
Subtract — restoring a balanced dimension
  • Prune monofractal spikes. A cluster whose local box-count density towers over the library's median is a clump — over-dense, low marginal information. Khoj mechanism: rank for SUBTRACT the idea-units in the highest-density boxes whose removal raises the global box-counting dimension (makes coverage more uniformly space-filling). You are smoothing a multifractal spike toward monofractal evenness.
  • Cut redundant self-similar copies. Test self-similarity: is the subject-mix of one shelf a scaled echo of the whole library? Where a sub-shelf is a near-duplicate of the global distribution, it adds no new shape. Khoj mechanism: compute similarity between each shelf's normalized profile and the parent's; the most redundant, most self-similar shelves are deaccession/demote candidates — keep one representative and prune the copies.
  • Devil's-staircase test for dead plateaus. A Cantor-like staircase has long flat runs of zero measure. Long runs of idea-units that never appear in any thread (108 threads) are inert plateaus. Khoj mechanism: SUBTRACT-rank units with zero thread-participation inside already-dense boxes.
Serendipity — scale-jumping as an engine
  • Zoom-level leaps. Standard recommenders move sideways within a subject. Khoj mechanism: deliberately jump ε — from a fine idea-unit straight to a distant coarse domain box that shares a rare motif — surfacing a book that is near in pattern but far in catalogue location. Self-similarity is the bridge: the shape of one shelf points to an analog shelf three domains away.
  • Strange-attractor random restart. A bandit exploits basins it already knows. Khoj mechanism: with small probability, reseed the recommendation walk at a random empty box far from the current orbit (edge-of-chaos restart) rather than the nearest neighbor — escaping the attractor's basin the collector keeps falling into.
  • Forking-path branch (Borges). Log the acquisition tree's un-taken branches — the L-system productions Khoj considered but didn't fire. Khoj mechanism: periodically resurface one abandoned fork ("the garden you didn't walk"), a structurally-motivated wildcard that is neither pure noise nor pure exploitation.
Sharpest single idea

Build one scale-resolved gap map: box-count the collection at every zoom level and render occupancy N(ε)/total as a curve across scales. The scale where that curve sags is the collection's true weak point, and it dictates everything — ADD fills the empty boxes there, SUBTRACT flattens the spikes that distort the curve, and SERENDIPITY leaps across scales along self-similar echoes. Gaps at every zoom, from one measurement.

Sources: Minkowski–Bouligand (box-counting) dimension; Scale-free networks & preferential attachment.

Probe 05 · Economic added lens

Economics & via negativa

The antilibrary's own Eco–Taleb bloodline: options, the barbell, the Lindy effect, and improvement by subtraction.

What this lens uniquely contributes. This is the antilibrary's own bloodline: Eco→Taleb, the unread book as the highest-value asset. Every other lens treats a book as content to be matched; this lens treats each acquisition as a financial position — an option with an asymmetric payoff — and the whole library as a portfolio of bets under irreversible uncertainty. Its deepest gift is the SUBTRACT limb: it is the only lens with a rigorous, named account of why letting go feels wrong (loss aversion, endowment, sunk cost) and how to engineer around it. It reframes the scarce resource not as shelf space but as attention — the one budget that cannot be expanded.

Add — the acquisition strategy
  • Barbell (Taleb). Split the acquisition budget explicitly: ~85–90% into the canonical core (Lindy-proven, foundational to owned clusters/threads) + ~10–15% into a wild-bet sleeve of high-variance, off-thesis books. Khoj rule: two ranked ADD lists, "Core" and "Wild," with a hard budget ratio — never let the middlebrow "safe-but-forgettable" mid-list dominate. The middle is where you die.
  • Buy the option, not the certainty (optionality/convexity). A cheap unread book is a real option on a future self: limited downside (its price + a slot), open-ended upside (it reframes a whole field). Khoj rule: score each candidate by convexity = (potential upside if it hits) ÷ (cost if it doesn't). Rank the wild sleeve by convexity, not by predicted relevance — you are buying payoff shape, not a prediction.
  • Value of waiting (Dixit–Pindyck real options). Acquisition is irreversible (money + attention spent). Under high uncertainty, waiting has option value. Khoj rule: for expensive or high-uncertainty candidates, output "Buy now" vs. "Watchlist" — recommend now only when the option-to-expand (a live thread actively demanding it) outweighs the option-to-wait.
  • Diversify against correlation (Markowitz). A library's risk is set by how much its holdings move together, not their average quality. Khoj rule: compute concentration across the 321 clusters; when a candidate is highly correlated with what's already owned (redundant coverage), penalize it. Reward books on the efficient frontier — new return (coverage) per unit of correlated risk. Diversification is the only free lunch.
  • Fixed budget → optimal stopping (secretary/37% rule). For a capped monthly acquisition budget, Khoj rule: survey ~37% of the candidate window as calibration (buy nothing), then take the next candidate that beats everything seen so far. Stops "buy the first shiny thing" and "endless deliberation" alike.
Subtract — the deaccession logic
  • Via negativa (Taleb). Knowledge grows by removing what's wrong or noise. Khoj rule: SUBTRACT is a first-class output, not an afterthought — every cycle must propose removals, because subtraction improves signal even when addition is zero.
  • Anti-Lindy hype cull. The Lindy filter runs both ways: survival predicts survival for the timeless, so recency + hype is a negative signal for the perishable. Khoj rule: flag books that are recent, trend-driven, and have low thread-attachment as decay candidates — the un-Lindy get culled first.
  • Redundancy given a dominating book. Khoj rule: if you own a strictly better book on the same node (higher centrality, more atomic-units drawn from it), the weaker one is dead weight — surface the pair, recommend demotion.
  • Low-optionality dead weight. Khoj rule: rank SUBTRACT candidates by inverse optionality — never opened, zero atomic-units, no cluster/thread links, low convexity. These are options that expired worthless.
  • Defeating endowment / loss aversion / sunk cost. These biases (Kahneman–Knetsch–Thaler) make us value what we own ~2× and frame removal as loss. Khoj counter-rules: (a) Reframe as opportunity cost — show what the freed attention/slot buys (a specific ADD candidate), converting a loss frame into a gain frame; (b) Ignore sunk cost explicitly — the price paid is gone; score only forward option value, and say so in the UI; (c) Demote, don't delete — a reversible "storage/deaccession" tier lowers the perceived loss and lets people act; (d) Default to a subtraction prompt (Klotz: people neglect subtraction even when reminded) — so Khoj must actively push a removal quota, not wait to be asked.
Serendipity / discovery
  • Institutionalize the tail (barbell sleeve = serendipity engine). You cannot predict which wild bet pays off, so you buy many cheap tail options precisely so that rare, unpredictable hits can occur. Khoj rule: the wild sleeve is deliberately seeded from low-correlation, high-convexity books far from current clusters — serendipity as policy, not luck.
  • Convexity tolerates being wrong. Because downside is capped, most wild bets should miss. Khoj rule: never penalize the sleeve's hit-rate; measure it only by whether any bet reframed a thread.
Sharpest single idea

Every unread book is an option, and a library is a barbell portfolio of them. So Khoj should score both lists on one currency — forward option value (convexity × optionality, discounted by correlation) — ignoring sunk cost entirely. ADD buys cheap tail options; SUBTRACT sells the ones that expired worthless; and the freed attention is the premium that funds the next bet.

Sources: Antifragile / barbell, Klotz, Subtract, Lindy effect, Endowment effect / loss aversion (Kahneman-Knetsch-Thaler), Dixit-Pindyck, Investment under Uncertainty, Markowitz efficient frontier, Secretary problem / 37% rule.

Probe 06 · Ecological added lens

The library as a living system

Niche theory, succession, keystone species, biodiversity indices — and the librarians' real weeding protocol.

What this lens uniquely contributes. Every other framing treats the antilibrary as an information store to optimize. Ecology treats it as a living community with a carrying capacity, a successional stage, and a health that depends on diversity, not size. That reframes both lists: ADD is colonization of empty niches and keystone introduction; SUBTRACT is weeding and thinning for the health of the whole, not tidying. Crucially, ecology supplies the only field-tested, professionally-codified SUBTRACT methodology in existence — the librarians' CREW/MUSTIE protocol — plus a far richer diagnostic toolkit (evenness, richness, rarity) than plain Shannon entropy.

Add — colonize, seed, and introduce
  • Vacant-niche colonization (Hutchinson). Hutchinson modeled the niche as an n-dimensional hypervolume over resource gradients. Map each book/cluster into concept-space (your 321 clusters, 108 threads are the axes). Khoj mechanism: detect holes in the occupied hypervolume — regions bounded by owned clusters but empty inside — and rank ADD candidates by how precisely they fill one. Purists note Hutchinson's niche technically precludes truly "empty" niches, so operationalize it as low-density regions between dense clusters, which is defensible and computable (nearest-neighbor distance in embedding space).
  • Keystone introduction (Paine, 1966). Paine removed Pisaster ochraceus and diversity collapsed from 15 species to 8 — proof one node can hold up a whole community. Inverse it: Khoj's highest-leverage ADD is the single book whose introduction would raise structural connectivity most — the book that bridges the most currently-disconnected threads. Rank by predicted bridge-centrality gain (∆ betweenness if this node were inserted), not by the book's own popularity. The marquee "one book that reorganizes everything" output.
  • Evenness enrichment (Pielou's J′ = H′/ln S). Shannon alone conflates richness and evenness. Pielou's J′ isolates evenness on a 0–1 scale; J′→0 means one cluster dominates. Khoj mechanism: compute J′ across clusters, then target ADDs at the under-even regions — thin clusters that would most raise J′ per book added. A diversity-first ADD ranking distinct from bridge-centrality.
  • Ecotone / edge seeding. The richest life sits at boundaries between habitats (edge effect). Khoj mechanism: flag interdisciplinary edges — pairs of adjacent-but-distinct clusters — and prioritize books that sit on the seam. These are the fertile ADD zones.
  • r-mode vs. K-mode acquisition (MacArthur & Wilson). Two acquisition modes as a life-history trade-off: r-mode = many cheap, exploratory books seeding disturbed/new areas (breadth, quantity); K-mode = few deep, canonical works in mature areas (depth, quality). Khoj mechanism: let the user set an r/K slider, and route the succession stage of each region to the right mode (young area → r-seed; saturated area → K-canon).
Subtract — weed, thin, but protect the rare

The library-science backbone. CREW (Continuous Review, Evaluation & Weeding; Texas State Library, 1976) uses the MUSTIE test. Map each criterion to an antilibrary signal:

  • M — Misleading: factually outdated/superseded-by-reality → flag old editions, disproven science.
  • U — Ugly: worn/unusable → damaged, or here, un-processable (no metadata, can't be placed in concept-space).
  • S — Superseded: a better edition/title exists → same-niche competitor detected; keep the stronger, demote the other.
  • T — Trivial: no lasting significance → low idea-unit yield; contributes nothing to any thread.
  • I — Irrelevant: outside collection scope/interests → sits in no active thread and far from every cluster.
  • E — Elsewhere: available elsewhere → freely accessible / borrowable; low reason to own.

Two ecological overlays on top of MUSTIE:

  • Cull invasive monocultures (competitive exclusion). An over-represented cluster crowding out diversity is an invasive monoculture. Khoj mechanism: where one cluster's share exceeds a threshold and depresses J′, mark its marginal, redundant members for pruning — thin the stand, don't clear-cut.
  • Thin same-niche redundant competitors. Two books occupying one niche → competitive exclusion says one is redundant. Rank near-duplicates; recommend demoting the weaker.
  • PROTECT rarity/endemism — the override. Weeding's cardinal error is culling the "unproductive." Ecology values rare and endemic species precisely because they're rare. Khoj mechanism: a hard protection flag — any book that is the sole occupant of its niche is exempt from SUBTRACT even with low "productivity." Rarity is the point of an antilibrary.
Serendipity — dispersal and edges
  • Seed rain from adjacent fields. Serendipity = propagules blowing in from neighboring habitats. Khoj mechanism: surface recommendations one hop outside the user's densest cluster — near enough to germinate, far enough to surprise.
  • The ecotone surprise. Edges hold the unexpected. Deliberately recommend at cluster boundaries, where a book reads as belonging to neither parent field — the highest-serendipity ADD.
  • Seed-bank germination (dormant unread books). Unread books are dormant seeds. Khoj mechanism: periodically "germinate" a long-dormant owned-but-unread title when a newly-added book creates the conditions (a now-connected thread) for it to matter — resurfacing from within, not just acquiring.
  • Optimal foraging / patch depletion. When a cluster is exhausted (idea-unit yield per new book falls), signal move to a new patch — an anti-monoculture nudge that doubles as discovery.
Sharpest single idea

Weed by MUSTIE, but make rarity a veto. Ship the librarians' battle-tested six-point SUBTRACT test (MUSTIE) as Khoj's deaccession engine — then invert its usual logic with one ecological rule: a book that is the only occupant of its niche can never be weeded. That single inversion is what makes it an antilibrary engine rather than a shelf-space optimizer — it prunes the redundant monoculture and protects the lone endemic, exactly backwards from an ordinary library, exactly right for a map of what you don't yet know.

Sources: CREW/MUSTIE — Texas State Library CREW Manual; Hutchinson niche & competitive exclusion — Wikipedia "Vacant niche" / "Ecological niche"; keystone species — Paine 1966, Nature Scitable & American Naturalist; r/K — MacArthur & Wilson via Wikipedia "R/K selection theory"; indices — Bio LibreTexts "Diversity Indices," Pielou's J′.

Lenses not yet probed

These have not been run — no findings exist for them yet. They are candidates, listed so you can pick which to send a probe into. Say the word on any and I'll dispatch it the same way as the six above.

Not run

Contemplative / Indic-philosophical

Neti-neti ("not this, not that") as the original via-negativa; aparigraha (non-possession) as a subtract ethic; Manthan (the churning of the ocean) as the project's own metaphor; darshan and guru–shishya as discovery. On-brand with the Sanskrit naming.

Not run

Historical / evolution of ideas

How canons and lineages actually form; Harold Bloom's "anxiety of influence"; memetics and the sociology of the citation; what the last thinker in a school owes the first — acquisition as joining a live intellectual descent.

Not run

Memory & cognition / neuroscience

Forgetting curves, spaced repetition, memory consolidation, the generation effect. Bears directly on the hardest subtract question: what to reread, what to demote once internalized, what a personal library owes to a finite memory.

Not run

Adversarial / red-team

The spec's own proactive-gap question made into a lens: "what is the book that would most disturb the existing structure?" Steelman the opposite library; dialectic; buy your strongest disagreement, not your next agreement.

Not run

Social / collective

What other people's libraries reveal about the gaps in yours; collaborative filtering done ethically (not popularity); the library as gift, inheritance, and shared object rather than a solo optimization.

Not run

Narrative / autobiographical

The collection as memoir — the story the add and subtract lists tell about who the owner is becoming. Bibliomemoir, the shelf as self-portrait; acquisitions chosen for the arc, not only the coverage.

Six probes filed above; six more lenses waiting. None of it has been ranked, merged, or reconciled — that's yours to do when you've looked through.