References

  1. Yehuda Koren, Robert Bell, Chris Volinsky. Matrix Factorization Techniques for Recommender Systems. IEEE Computer, 2009. The Netflix-Prize-era reference for MF (our SGD version).

  2. Yifan Hu, Yehuda Koren, Chris Volinsky. Collaborative Filtering for Implicit Feedback Datasets. ICDM 2008. Implicit ALS: preference + confidence (our ImplicitALS).

  3. Steffen Rendle et al. BPR: Bayesian Personalized Ranking from Implicit Feedback. UAI 2009. Pairwise learning-to-rank with negative sampling (our BPR).

  4. Greg Linden, Brent Smith, Jeremy York. Amazon.com Recommendations: Item-to-Item Collaborative Filtering. IEEE Internet Computing, 2003. The item-item neighborhood method at scale.

  5. Paul Covington, Jay Adams, Emre Sargin. Deep Neural Networks for YouTube Recommendations. RecSys 2016. The canonical two-stage (candidate generation + ranking) deep architecture.

  6. Xiangnan He et al. Neural Collaborative Filtering. WWW 2017. Neural generalization of matrix factorization.

  7. Maurizio Ferrari Dacrema, Paolo Cremonesi, Dietmar Jannach. Are We Really Making Much Progress? A Worrying Analysis of Recent Neural Recommendation Approaches. RecSys 2019. Why strong, well-tuned baselines matter (and often win).

  8. Cai-Nicolas Ziegler, Sean M. McNee, Joseph A. Konstan, Georg Lausen. Improving Recommendation Lists Through Topic Diversification. WWW 2005. Intra-list similarity: the diversity criterion in our slate rubric.

  9. Lianmin Zheng et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS 2023. LLM judges, plus the position and verbosity biases our rubric chapter guards against.

  10. Moses Charikar. Similarity Estimation Techniques from Rounding Algorithms. STOC 2002. SimHash, the fingerprint behind our ingest-time story clustering.

  11. Gurmeet Singh Manku, Arvind Jain, Anish Das Sarma. Detecting Near-Duplicates for Web Crawling. WWW 2007. SimHash + Hamming search at Google scale; the production version of our LSH banding.

  12. Adith Swaminathan et al. Off-Policy Evaluation for Slate Recommendation. NeurIPS 2017. The "slate" as the unit of recommendation and evaluation; where the term in our rubric chapters comes from.

The research arc (Chapters 30 to 35)

  1. Aditya Pal et al. PinnerSage: Multi-Modal User Embedding Framework for Recommendations at Pinterest. KDD 2020. The argument against a single averaged profile vector (cluster the user's items, keep a medoid per cluster) plus exponential time-decayed cluster importance; the finding our walkthrough rides into a PRD.

  2. Yi Ding, Xue Li. Time Weight Collaborative Filtering. CIKM 2005. Recency-decayed weighting of interactions; the ancestor of our EMA half-life.

  3. Ruining He, Julian McAuley. VBPR: Visual Bayesian Personalized Ranking from Implicit Feedback. AAAI 2016. The canonical "fold pretrained image features into ranking" result.

  4. Andrew Zhai et al. Learning a Unified Embedding for Visual Search at Pinterest. KDD 2019. One production image embedding, validated offline, in user studies, and in online A/B: the full evaluation ladder in one paper.

  5. Alec Radford et al. Learning Transferable Visual Models From Natural Language Supervision. ICML 2021. CLIP: images and text in one embedding space.

  6. Lihong Li, Wei Chu, John Langford, Robert E. Schapire. A Contextual-Bandit Approach to Personalized News Article Recommendation. WWW 2010. LinUCB on Yahoo's front page: bandits running inside a production news recommender.

  7. William R. Thompson. On the Likelihood that One Unknown Probability Exceeds Another in View of the Evidence of Two Samples. Biometrika, 1933. Thompson sampling, the scheduler in our sprint-allocation lab.

  8. Ronald A. Howard. Information Value Theory. IEEE Transactions on Systems Science and Cybernetics, 1966. EVPI: pricing an experiment before running it.

  9. Leslie Pack Kaelbling, Michael L. Littman, Anthony R. Cassandra. Planning and Acting in Partially Observable Stochastic Domains. Artificial Intelligence, 1998. POMDPs: the setting where gathering information is itself part of an optimal policy; the serious version of our research-MDP cartoon.

  10. Burr Settles. Active Learning Literature Survey. University of Wisconsin-Madison, TR 1648, 2009. Query what is most informative; directly useful for golden-set labeling.

  11. Jasper Snoek, Hugo Larochelle, Ryan P. Adams. Practical Bayesian Optimization of Machine Learning Algorithms. NIPS 2012. Expensive experiments chosen by expected improvement; see also Frazier's tutorial (arXiv:1807.02811) for the value-of-information view.

  12. Olivier Chapelle, Thorsten Joachims, Filip Radlinski, Yisong Yue. Large-Scale Validation and Analysis of Interleaved Search Evaluation. ACM TOIS, 2012. Interleaving: rank-sensitive online evaluation that needs far less traffic than A/B; Netflix's TechBlog describes using it to prune rankers in days.

  13. Ron Kohavi, Diane Tang, Ya Xu. Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing. Cambridge University Press, 2020. The book on not fooling yourself once research graduates to A/B.

  14. Darren Edge et al. From Local to Global: A Graph RAG Approach to Query-Focused Summarization. arXiv:2404.16130, 2024. GraphRAG: entity/claim graphs plus Leiden community summaries for corpus-wide questions.

  15. Shunyu Yao et al. Tree of Thoughts. NeurIPS 2023; Andy Zhou et al. Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models. ICML 2024. Tree search over model reasoning: MDP thinking inside deliberation.

  16. David Wadden et al. Fact or Fiction: Verifying Scientific Claims. EMNLP 2020. SciFact: SUPPORTS/REFUTES claim verification, the benchmarked form of adversarial claim-checking.

  17. Ihsan Gunes, Cihan Kaleli, Alper Bilge, Huseyin Polat. Shilling Attacks Against Recommender Systems: A Comprehensive Survey. Artificial Intelligence Review 42, 2014. Fake-profile attacks on collaborative filtering; why click feedback is not ground truth (Chapter 35).

  18. FutureSearch. Deep Research Bench: Evaluating AI Web Research Agents. arXiv:2506.06287, 2025. The frozen-web benchmark behind Chapter 34's finding that the "deep research" label is not a capability guarantee.

Methods, tools & engines (Chapters 30 to 35)

Regulation, licensing & security (Chapter 35)

Tools & libraries

  • implicit: fast ALS / BPR (Cython).
  • imagehash (Pillow), production perceptual hashes (pHash/dHash) for the thumbnail dedupe in the rubric chapter.
  • LightFM: hybrid content + collaborative (WARP/BPR).
  • FAISS / HNSW / ScaNN: ANN serving for candidate generation.
  • OpenSearch / Milvus / Qdrant / Pinecone: vector databases with kNN search.
  • TensorFlow Recommenders / TorchRec: two-tower and deep ranking models.

Companion books here

This book's code

All depend only on NumPy and the standard library (the judge optionally uses the Anthropic SDK when an API key is set).