Skip to content

Latest commit

 

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

feedrank

A recommendation system I built to understand how recsys works end to end, not just the model part. Built it on Amazon review data across 3 product categories.


why I built this

Most tutorials show you how to train a model. They don't show what happens when the model is too slow, or Redis goes down, or half your users have never been seen before. I wanted to build something that had all of that figured out.


what I read before building this

Not papers. Actual blog posts by people who have built recommendation systems and written about what went wrong.

Understanding implicit feedback and ALS:

How production systems are actually built:

Feature engineering for ranking:


what it does

Takes a user ID and session context, returns 10 ranked product recommendations with scores and metadata. The whole thing runs on CPU, no GPU needed.

Two paths through the system:

  • New users (fewer than 5 interactions in training data): get popular items filtered for category diversity. Fast, no model needed.

  • Existing users: ALS retrieval pulls ~200 candidates from a faiss index, LightGBM re-ranks them using user and item features, constraint pipeline adjusts for seller diversity and price range, top 10 returned.

If the system is too slow or retrieval fails, it falls back to a cached feed from the last request, or global popularity if that is also unavailable. Something always comes back.


architecture

user request (user_id, session_items)
        |
        v
cold start check (< 5 interactions?)
   yes                    no
    |                      |
    v                      v
popular items        ALS + faiss retrieval
with category        (top 200 candidates)
diversity cap              |
    |                      v
    |              LightGBM re-ranking
    |              (8 features, trained on
    |               rating as purchase proxy)
    |                      |
    |                      v
    |              constraint pipeline
    |              - max 2 items per seller
    |              - price band penalty
    |              - freshness boost
    +----------+           |
               v           v
            top 10 items + latency + source label

If total latency goes over 200ms, the system serves a cached feed instead of going through retrieval and ranking.


tech stack and why

ALS (Alternating Least Squares) for collaborative filtering. I looked at matrix factorization approaches and ALS was the most practical for CPU. The implicit library implements it with sparse matrix operations and trains in a few minutes. Neural approaches need GPU and longer training time for a side project.

faiss for vector search. After ALS gives you user and item embeddings, you need to find the nearest items quickly. faiss IndexFlatIP does exact inner product search. On 1.2M items it takes about 15ms per query.

LightGBM for re-ranking. I needed something fast at inference time. LightGBM scores 200 candidates in 2-4ms. Training on 54M rows takes under 2 minutes. I also wanted feature importance to understand what the model is using, which tree models give you for free.

FastAPI for serving. Straightforward to use, handles async requests, pydantic validation on the request schema.

Redis for caching. User feeds cached for 30 minutes. If Redis is down, the system catches the ConnectionError and serves from the popularity baseline. No crash.

DuckDB for feature computation. The training set is 54M rows which does not fit in pandas easily on a 16GB machine. DuckDB can run SQL aggregations directly on parquet files using a few hundred MB of RAM.


dataset

Amazon Reviews 2023 from McAuley Lab (HuggingFace). Three categories:

  • Clothing, Shoes and Jewelry
  • Beauty and Personal Care
  • Sports and Outdoors

I thought this was going to be 17M reviews based on the category pages. The dataset actually has all reviews since 1996, not just 2023. Final count was 108M reviews. After filtering to users with at least 5 interactions and items with at least 10, I got 54M training interactions, 6.1M users, 1.2M items.

Timestamps are in milliseconds (not seconds). Prices are strings like "$29.99" or "$10.00 - $39.99" (ranges, took the average). About 30% of items have no parseable price.


how to run

git clone https://github.com/sohamukute/feedrank && cd feedrank
pip install numpy==1.26.4 scipy==1.13.1 && pip install -r requirements.txt

# download the 6 parquet files from HuggingFace (run on Colab for speed)
# copy them to data/raw/

make all
docker-compose up
curl -X POST http://localhost:8000/recommend \
  -H "Content-Type: application/json" \
  -d '{"user_id": "AG73BVBKUOH22USSFJA5ZWL7AKXA", "session_items": [], "n": 10}'

results

Tested on June-September 2023 interactions (held-out test period).

approach ndcg@10 recall@50 hr@5 p99_ms coverage
popularity 0.0139 0.0113 0.0171 0.0 10
als_only 0.0080 0.0026 0.0103 15.8 34,721
als+lightgbm 0.0079 0.0025 0.0103 23.6 53,015

Popularity looks best in ndcg because it serves the same 10 popular items to everyone and those items happen to appear in many test interactions. That is not useful in practice. ALS+LightGBM recommends from 53,000 unique items across all users (5,300x more catalog coverage) at p99 of 23.6ms.

cold start breakdown

bucket n_users ndcg@10 recall@50
cold (0-5 interactions) 29,657 0.0088 0.0050
warm (5-20 interactions) 3,048 0.0002 0.0009
hot (20+ interactions) 13 0.0000 0.0000

Cold users outperform warm and hot here because the cold path serves popular items that happen to be in many test interactions. The warm/hot numbers are low because the test period is only 3 months so there is very sparse signal per user.


things I ran into

  • Dataset was 108M rows, not 17M. Had to rewrite the cleaning pipeline to process one category at a time instead of loading everything together.

  • Timestamps were in milliseconds. All my date calculations were off by 1000x.

  • ALS retrieval was 720ms per request because I was loading a 1.5GB numpy file from disk inside the retrieve function on every single request. Fixed to load once at startup.

  • pyarrow string type has 32-bit offsets and overflows when you concatenate 100M+ rows of string data. Had to switch to large_string type.

Details in the errors/ folder.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages