Table of Contents
- 1. Links
- 2. Notes from TF Recommendation YT playlist
- 3. General notes
1. Links
- Tensorflow recommendation systems YT playlist
- Wide and deep blog: blog
- Youtube paper on deep recommendations: paper
- Netflix recommendation case study: paper
2. Notes from TF Recommendation YT playlist
Recommender system: Connect one item with another. Simplest example would be to
match a user with a movie they may perhaps be interested in watching next.
Challenges
- High cardinality sparse features (very few examples for cross features)
- Usually conflicting objectives that are difficult to filter down to one
- Long-term effects are challenging to elicit from A/B tests
- Biased data. The videos shown to a user influence the videos seen by a user. Also
imagine we only show highly rated videos and use the data to train a model. Predicting
on low-rated reviews would be out of domain extrapolation which can yield low-fidelity
results. We might have to deliberately display low probability of click videos to
collect diverse data.
- Scale
Usually broken down into stages:
- Retrieval: Fast retrieval from O(millions) -> O(thousands)
- Ranking: O(1000) -> O(100)
- Post-Ranking and Ordering: O(100) -> O(10)
Solutions
- For retrieval, it’s usually some sort of Approximate Nearest Neighbours technique
based on query and candidate embeddings which are usually made to be in the same
embedding space. TF provides for ScaNN (Scalable Nearest Neighbours).
- For privacy and speed, on-device smaller models are also common.
3. General notes
Dimension 1: Architecture
- Collaborative filtering and matrix/tensor factorisation
- Model as probability of interaction (click)
For probability of interaction we can build this as any binary classification problem
where for every user and a preselected set of items, we predict the probability of
click. The preselection comes from L1-1 ranking (recall phase).
- Model as pairwise ranking (compare alternatives and make the model select one)
Dimension 2: Type of model
Modelling as probability of click (or pairwise ranking) can be done through
any classification model such as logistic regression, xgboost or neural nets.
Neural net approaches
- Wide and deep: Wide has sparse cross features which are memorised (overfit) and then
the deep is a sequence of MLP layers or other carefully crafted layers that let the
model generalise better. The wide part is for exceptions and deep part is for
generalisations.
- Two tower model: One tower for user embeddings and meta features and another
tower for the item. The final dimensions of the two towers match, we simply
dot and take the sigmoid to compute the probability.
- Sequence model: User actions are first classified into different buckets. Each
of these gets an action embedding. Users have a general user embedding. We then concat
these and learn a sequence model such as an LSTM or transformers where we predict the
action in the next time step. This apparently was a game-changer for Netflix.