Background & Context§
The proliferation of AI-generated content—often derided as "slop"—has prompted a backlash among readers and researchers who value depth and rigor. A recent article on the Substack Res Obscura introduced a project that curates award-winning non-fiction books as a signal of quality, explicitly positioning them as the antithesis of AI slop. The project uses AI tools for data collection, coding, and semantic search, yet the creator frames it as a defense of human-curated, long-form knowledge. This tension has ignited a broader conversation on Hacker News about the value of book prizes, the role of AI in discovery, and the irreplaceability of deep reading.
The News: What Happened Exactly§
The article "Quality non-fiction books are the antithesis of AI slop" describes a database that aggregates books that have won or been shortlisted for major non-fiction prizes—including the Pulitzer, National Book Award, and PROSE awards—as a filter for quality. The project scrapes award lists, combines them with metadata from open library sources, and enables semantic search across the corpus. The author argues that these award-winning books represent a "better-than-average signal" in an era of information overload.
However, the Hacker News discussion reveals critical nuances. A commenter who volunteered with a book award noted that publishers mass-submit books to every remotely relevant prize, making it a "cost of doing business." The NCR Book Award scandal of 2013, where judges were revealed not to have read the books they judged, underscores the fragility of awards as quality signals. Similarly, the PROSE Awards are criticized for being so broad that a win or finalist placement carries inflated prestige.
The discussion also surfaced irony: the project decries AI-generated slop but uses AI for its own construction. One commenter observed: "There is really nothing 'AI' about this aside from the tool that collected the data and coded it, and, crucially, semantic search… So really, everything about it is AI. And that’s not a bad thing!" This paradox highlights the difficulty of separating tool from output in a world where AI is ubiquitous.
Another thread explored the changing nature of libraries and discovery. A commenter lamented seeing students in university libraries surrounded by shelves of books, yet all with ChatGPT open. The internet once felt like a library of serendipitous discovery; now it is algorithmically filtered. The project attempts to reintroduce curation, but questions remain: can a human-curated award system scale? And does using AI to find quality non-fiction books undermine the goal of resisting AI slop?
Historical Parallels & Similar Incidents§
The tension between human curation and algorithmic discovery has a long history in the digital era. A notable parallel is the rise and fall of Delicious (formerly del.icio.us), a social bookmarking service launched in 2003. Delicious allowed users to tag and share bookmarks, creating a folksonomy that many saw as a democratic alternative to expert-curated directories like Yahoo! or the Open Directory Project. The platform thrived for years, offering a human-powered discovery engine for web content. However, as the volume of submissions grew, the signal-to-noise ratio deteriorated. Spam, gaming of tags, and the sheer scale of content overwhelmed the community's ability to curate effectively. By the time Delicious was sold and later shuttered, it had become a cautionary tale: human curation alone cannot scale without robust mechanisms for quality control.
The parallel with the book awards project is clear. Both rely on human judgment—tags from Delicious users; jury selections from prize committees—to surface quality. But both face the same fundamental issues: gaming of the system (mass submissions to book prizes, spam tagging on Delicious) and limited bandwidth (how many books can a prize jury truly read? How many links can a community vet?). Delicious eventually failed because it couldn't solve these problems; the book awards project may face a similar fate if it relies solely on existing prize lists without further filtering.
A second, more recent parallel is the Google Books Ngram Viewer controversy of 2010–2011. Google digitized millions of books and created a tool to chart word frequencies over time, enabling quantitative analysis of culture. Scholars initially hailed it as a breakthrough for the digital humanities. However, critics soon pointed out that the corpus was flawed: it overrepresented scientific texts, contained OCR errors, and excluded many non-English or non-Western works. The tool's output was only as good as its input, and the input was biased by Google's digitization priorities. Similarly, the book awards project's quality signal is only as good as the prize lists it ingests—and as the NCR scandal shows, those prizes can be compromised.
Both parallels teach a lesson: datasets are never neutral. Whether it's user-generated tags, digitized books, or prize-winning titles, the act of curation introduces biases. The book awards project must confront this head-on, perhaps by weighting awards by prestige, cross-referencing multiple sources, or incorporating user reviews to offset institutional shortcomings. Without such measures, it risks replicating the same flaws that plagued earlier curation efforts while adding a new paradox: using AI to fight AI slop may inadvertently produce a dataset that is itself incomplete or skewed.