A recommendation system I built to understand how recsys works end to end, not just the model part. Built it on Amazon review data across 3 product categories.
Most tutorials show you how to train a model. They don't show what happens when the model is too slow, or Redis goes down, or half your users have never been seen before. I wanted to build something that had all of that figured out.
Not papers. Actual blog posts by people who have built recommendation systems and written about what went wrong.
Understanding implicit feedback and ALS:
- Intro to Recommender Systems: Collaborative Filtering - Ethan Rosenthal, 2015. Explains the difference between explicit ratings and implicit signals like clicks and purchases. This is the post that made the problem make sense to me.
- Intro to Implicit Matrix Factorization with ALS - Ethan Rosenthal, 2016. Takes the above and shows how to actually train ALS on real data with the implicit library.
- Faster Implicit Matrix Factorization - Ben Frederickson, the person who built the implicit library. Shows the internals and why it is fast on CPU.
How production systems are actually built:
- System Design for Recommendations and Search - Eugene Yan, 2021. This is where I got the idea to split retrieval and ranking instead of doing everything in one step. He works at Amazon and describes real architectures from Alibaba, DoorDash, and others.
- Real-time Machine Learning For Recommendations - Eugene Yan, 2021. When to use real-time vs precomputed recommendations and how the serving layer works.
- System Architectures for Personalization and Recommendation - Netflix Tech Blog, 2013. The offline/online/nearline split they describe shaped how I thought about the pipeline.
Feature engineering for ranking:
- For Your Ears Only: Personalizing Spotify Home with Machine Learning - Spotify Engineering, 2020. How they think about candidate generation and ranking. The two-stage structure they describe is what I implemented.
- The Rise and Lessons Learned of ML Models to Personalize Content on Home - Spotify Engineering, 2021. Practical lessons about what breaks in production. Part I covers candidate generation, Part II covers ranking.
Takes a user ID and session context, returns 10 ranked product recommendations with scores and metadata. The whole thing runs on CPU, no GPU needed.
Two paths through the system:
-
New users (fewer than 5 interactions in training data): get popular items filtered for category diversity. Fast, no model needed.
-
Existing users: ALS retrieval pulls ~200 candidates from a faiss index, LightGBM re-ranks them using user and item features, constraint pipeline adjusts for seller diversity and price range, top 10 returned.
If the system is too slow or retrieval fails, it falls back to a cached feed from the last request, or global popularity if that is also unavailable. Something always comes back.
user request (user_id, session_items)
|
v
cold start check (< 5 interactions?)
yes no
| |
v v
popular items ALS + faiss retrieval
with category (top 200 candidates)
diversity cap |
| v
| LightGBM re-ranking
| (8 features, trained on
| rating as purchase proxy)
| |
| v
| constraint pipeline
| - max 2 items per seller
| - price band penalty
| - freshness boost
+----------+ |
v v
top 10 items + latency + source label
If total latency goes over 200ms, the system serves a cached feed instead of going through retrieval and ranking.
ALS (Alternating Least Squares) for collaborative filtering. I looked at matrix factorization approaches and ALS was the most practical for CPU. The implicit library implements it with sparse matrix operations and trains in a few minutes. Neural approaches need GPU and longer training time for a side project.
faiss for vector search. After ALS gives you user and item embeddings, you need to find the nearest items quickly. faiss IndexFlatIP does exact inner product search. On 1.2M items it takes about 15ms per query.
LightGBM for re-ranking. I needed something fast at inference time. LightGBM scores 200 candidates in 2-4ms. Training on 54M rows takes under 2 minutes. I also wanted feature importance to understand what the model is using, which tree models give you for free.
FastAPI for serving. Straightforward to use, handles async requests, pydantic validation on the request schema.
Redis for caching. User feeds cached for 30 minutes. If Redis is down, the system catches the ConnectionError and serves from the popularity baseline. No crash.
DuckDB for feature computation. The training set is 54M rows which does not fit in pandas easily on a 16GB machine. DuckDB can run SQL aggregations directly on parquet files using a few hundred MB of RAM.
Amazon Reviews 2023 from McAuley Lab (HuggingFace). Three categories:
- Clothing, Shoes and Jewelry
- Beauty and Personal Care
- Sports and Outdoors
I thought this was going to be 17M reviews based on the category pages. The dataset actually has all reviews since 1996, not just 2023. Final count was 108M reviews. After filtering to users with at least 5 interactions and items with at least 10, I got 54M training interactions, 6.1M users, 1.2M items.
Timestamps are in milliseconds (not seconds). Prices are strings like "$29.99" or "$10.00 - $39.99" (ranges, took the average). About 30% of items have no parseable price.
git clone https://github.com/sohamukute/feedrank && cd feedrank
pip install numpy==1.26.4 scipy==1.13.1 && pip install -r requirements.txt
# download the 6 parquet files from HuggingFace (run on Colab for speed)
# copy them to data/raw/
make all
docker-compose up
curl -X POST http://localhost:8000/recommend \
-H "Content-Type: application/json" \
-d '{"user_id": "AG73BVBKUOH22USSFJA5ZWL7AKXA", "session_items": [], "n": 10}'Tested on June-September 2023 interactions (held-out test period).
| approach | ndcg@10 | recall@50 | hr@5 | p99_ms | coverage |
|---|---|---|---|---|---|
| popularity | 0.0139 | 0.0113 | 0.0171 | 0.0 | 10 |
| als_only | 0.0080 | 0.0026 | 0.0103 | 15.8 | 34,721 |
| als+lightgbm | 0.0079 | 0.0025 | 0.0103 | 23.6 | 53,015 |
Popularity looks best in ndcg because it serves the same 10 popular items to everyone and those items happen to appear in many test interactions. That is not useful in practice. ALS+LightGBM recommends from 53,000 unique items across all users (5,300x more catalog coverage) at p99 of 23.6ms.
| bucket | n_users | ndcg@10 | recall@50 |
|---|---|---|---|
| cold (0-5 interactions) | 29,657 | 0.0088 | 0.0050 |
| warm (5-20 interactions) | 3,048 | 0.0002 | 0.0009 |
| hot (20+ interactions) | 13 | 0.0000 | 0.0000 |
Cold users outperform warm and hot here because the cold path serves popular items that happen to be in many test interactions. The warm/hot numbers are low because the test period is only 3 months so there is very sparse signal per user.
-
Dataset was 108M rows, not 17M. Had to rewrite the cleaning pipeline to process one category at a time instead of loading everything together.
-
Timestamps were in milliseconds. All my date calculations were off by 1000x.
-
ALS retrieval was 720ms per request because I was loading a 1.5GB numpy file from disk inside the retrieve function on every single request. Fixed to load once at startup.
-
pyarrow string type has 32-bit offsets and overflows when you concatenate 100M+ rows of string data. Had to switch to large_string type.
Details in the errors/ folder.