We did not tidy the shelf or stage the shots. Real bookshelves are crowded, crooked and badly lit, and a tool that reads only a clean one is no use. The eight photos vary on purpose — one tidy, one crammed three shelves deep, one of repeated titles, one at an angle, one a jumble of orientations and languages — so that no two fail for the same reason.
| Model | Size | Recall | Precision | Author | Hallucinations | Fabrications | Time |
|---|---|---|---|---|---|---|---|
| Qwen 7B | 7B | 0.73 | 0.88 | 0.69 | 21 | 7 | 6m |
| Llama 11B | 11B | 0.72 | 0.76 | 0.70 | 62 | 18 | 22m |
| Gemma 12B | 12B | 0.58 | 0.67 | 0.41 | 66 | 17 | 7m |
| Qwen 3B | 3B | 0.46 | 0.68 | 0.53 | 29 | 5 | 4m |
| Smol 2B | 2.2B | 0.01 | 0.01 | 0.00 | 148 | 129 | 6m |
| model | T001 13 | T001 ↻ flip | T002 17 | T003 17 | T004 26 | T005 15 | T006 69 | T006 ↻ flip | T007 19 | T008 13 |
|---|---|---|---|---|---|---|---|---|---|---|
| Qwen 7B | 1.00 | 1.00 | 1.00 | 0.88 | 0.46 | 1.00 | 0.52 | 0.59 | 1.00 | 0.92 |
| Llama 11B | 1.00 | 0.85 | 0.82 | 0.82 | 0.69 | 0.93 | 0.61 | 0.32 | 0.58 | 0.85 |
| Gemma 12B | 0.85 | 0.61 | 0.59 | 0.71 | 0.61 | 0.80 | 0.41 | 0.30 | 0.68 | 0.54 |
| Qwen 3B | 1.00 | 0.69 | 0.94 | 0.00 | 0.54 | 0.60 | 0.20 | 0.22 | 0.53 | 0.85 |
| Smol 2B | 0.00 | 0.00 | 0.06 | 0.00 | 0.00 | 0.00 | 0.01 | 0.03 | 0.00 | 0.00 |
1. The picture matters more than the books. We built the eight shelves to be hard in different ways. One shelf held the same journal nine times over. Another held six books by a single author. A third packed in look-alike covers. These were the traps, and we expected them to break the models. They did not. When the type was large and the shelf was tidy, the models read the repeats and the look-alikes without trouble. The thing that decided success was not the books at all. It was the photograph: how it was framed, which way the books faced, the angle of the camera, and how clearly the spines could be read.
2. One model was clearly best. Qwen 7B led on every measure. It read the most books and invented the fewest. It also beat two larger rivals, an 11-billion model from Meta and a 12-billion model from Google. In this task, more parameters did not mean better sight. The gap ran the same way inside one family: the 7-billion model was far stronger than its 3-billion sibling. The smallest model, at 2 billion, could not do the task at all.
3. Recall falls as a shelf grows denser. Recall is the share of the real books that a model reads. It fell as the number of books in the frame rose. On the 13-book shelf, the best model read every title. On the 69-book shelf, it read about half. Every model showed the same slope. A crowded photograph gives the model more to read, less room per spine, and more chances to lose its place.
4. Framing creates false errors. Our answer key covers one shelf per photograph. But several photographs caught two or three shelves at once, and the models read all of them. The extra reads were correct: real books, read right, from the shelf above or below. They simply were not the books we asked about, so the scorer marked them wrong. On the busiest photograph, the models “found” titles like Data Science for Business and Trigonometry. Those books are real. They sit one shelf down. The fault was in the framing, not the model, which means the scoreboard is partly measuring the camera.
5. Orientation is a large effect. We took the easiest shelf and turned it upside-down. The books were identical; only the angle changed. The weaker models lost up to a third of their recall. The best model held its recall, but failed in a worse way. It still found all 13 real books. Then it added 15 that were not there. Upside-down text did not stop the model from reading the words. It stopped it from telling where one book ended and the next began. Turning the picture did not stop it from reading. It stopped it from stopping.
6. Flat books may read more easily than upright ones. One shelf held both kinds: some books standing up, some lying flat in a stack. The models read the flat books better. Across the models, the flat books scored about half right; the upright books, about a third. Four of the five models showed this. The reason is likely plain. A flat book’s title runs left-to-right, the way we read. An upright spine’s title runs sideways, so the model must read it turned on its side. We cannot call this settled, though, and here is the caveat. The upright books on that shelf were also small and faint, so part of the gap is legibility, not orientation. And the best model read the upright books slightly better, not worse. So we treat this as a hint, not a result. The clean test is still to come: photograph the same books once standing and once flat, and change nothing else.
7. Honesty matters more than accuracy. When a model cannot read a spine, it fails in one of two ways. It garbles a real title, or it invents a fake one. We call the invented kind a fabrication, and the difference between the two is everything. Garbling is recoverable: you can see the model reaching for a real book, and All the Light We Cannot See comes back as “We Cannot See.” Fabrication gives you no such warning. It is a clean, confident, wrong answer. The best model rarely fabricated. The smallest did little else. Shown a shelf of science books it could not make out, it reported The Great Gatsby, 1984, and The Lord of the Rings. None were there. It had stopped reading the shelf and started guessing what a shelf usually holds. The number that separates a safe model from a risky one is the fabrication rate, not the accuracy score.
8. You need a good-enough model first, then a good photograph. It is tempting to say the photograph matters more than the model. It does not, and one table shows why. These are the read-rates on the clean, upright, legible shelves — the ones where the picture is not the problem:
| Model | Read-rate on clean shelves |
|---|---|
| Qwen 7B | 1.00 |
| Llama 11B | 0.83 |
| Qwen 3B | 0.77 |
| Gemma 12B | 0.73 |
| Smol 2B | 0.01 |
Same easy photographs. The results run from perfect down to nothing, and that whole spread is the model, not the picture. Only the best model read the clean shelves in full. The others missed a fifth to a quarter of the books even when nothing was wrong with the shot, and the smallest read almost none of them. So there is an order to it. First the model has to be good enough to read a clean shelf. Below that bar, the model is the ceiling, and no photograph can lift it. Above that bar, the model reads any clean shelf, and from there the photograph decides the rest.
The practical rule follows from both halves. Pick a model good enough to read a clean shelf — here that was Qwen 7B, not the larger models. Then give it a good photograph: one shelf, so there are no neighbours to misread; straight on, so the small author text stays legible; in good light, for contrast; and right-side up. Get both right, and a free model on a laptop will read your shelf well.
Turning the easiest shelf upside-down, with the same books, is the cleanest test of orientation. The weaker models lost up to a third of their recall; the best model kept its recall but began inventing books — 0 became 15.