Skip to main content

A Film Recommendation Engine Package

Project description

Film Recommender Engine

[!IMPORTANT] It is highly recommended that you refer to the two below resources in order to get the best possible sense of the project!

  • End-to-end demonstration notebook: examples/demo.ipynb
  • Full technical write-up: docs/overview.docx

This repository represents a film recommendation engine built using collaborative filtering (matrix factorisation), and implemented within a full end-to-end simulation framework for evaluating model performance, monitoring key prediction metrics, detecting data distribution drift, and executing retraining strategies.

The engine is structured as a Python package (film_recommender_cg) that would conceptually be adapted for deployment across regions and markets. The PoC demonstrates modular architecture, offline batch scoring, model lifecycle management, and monitoring strategies aligned with real‑world MLOps practices.


Project Overview

This project implements a modular, production-inspired recommendation engine packaged as:

film_recommender_cg

It demonstrates:

  • Collaborative filtering using scikit-surprise SVD (Singular Value Decomposition) model
  • Cold-start handling through popularity-based fallback recommendations
  • Batch scoring, monitoring, and model lifecycle management
  • Simulation of user behaviour under shifting data distributions
  • Drift detection and challenger-vs-champion retraining logic
  • Usage of MLOps tools (Weights & Biases, MongoDB)

The design mirrors how recommendation engines operate in commercial environments with fixed content catalogs, evolving user cohorts, and monitoring-based triggering of retraining logic.


Core Methodology

Collaborative Filtering

Users with at least five interactions are served by a matrix-factorisation model. True cold-start users are routed to a popularity-based fallback.

Dataset

The system uses a subset of The Movies Dataset (Kaggle).
To avoid memory issues, the model is trained on the first 120,000 rows, covering ~1.8k users and ~4k films.

Simulation Engine

A chronological simulation loop generates synthetic user interactions:

  • Each user consumes 1–20 films per round
  • The model scores all unseen items
  • Users select films using weighted sampling (favours high-ranked items)
  • Ratings = predicted score + calibrated noise
  • Already-seen films are masked at inference

A calibrated noise factor ensures simulated R² resembles real-world performance instead of inflating metrics.


Monitoring & Evaluation

The system tracks several key recommender-quality dimensions:

  • Accuracy – R² on held-out/synthetic data
  • Coverage – proportion of catalog ever recommended
  • Diversity – average cosine distance between recommended items
  • Personalization – how different recommendations are across users

Per-genre segmentation is used for drift detection, ignoring segments with <100 interactions to reduce noise.

The simulation allows “what-if” drift experiments, such as penalising specific genres to trigger retraining.


Retraining Logic

Retraining is triggered when:

  • Global R² drops below the established baseline, or
  • Any genre-level metric deteriorates

Challenger models are only promoted if:

  • They improve degraded metrics
  • They do not harm stable segments
  • No new regressions are introduced

Rejected challengers mirror the human-review workflow used in real production systems.


Architecture & Tools

  • VSCode – development and orchestration
  • MongoDB – ratings storage
  • Weights & Biases – experiment tracking and model evaluation
  • Python package structure – clean separation of trainer, scorer, and utilities

Training uses interactions within the desired retraining window only.
Inference requires the full ratings table to correctly filter previously watched films.


Installation

The film_recommender_cg package is available on PyPI and can be installed using pip:

pip install film_recommender_cg

Future Directions

Potential next steps might include:

  • Hybrid models with metadata (genres, languages, keywords, embeddings)
  • Time-aware or rolling-window training
  • Real-time inference APIs

Full Documentation

Once again, a complete technical overview—including methodology, architecture, simulation design, and retraining logic—is available in:

docs/overview.docx


License

Apache 2.0 License


Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

film_recommender_cg-1.0.6.tar.gz (29.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

film_recommender_cg-1.0.6-py3-none-any.whl (26.1 kB view details)

Uploaded Python 3

File details

Details for the file film_recommender_cg-1.0.6.tar.gz.

File metadata

  • Download URL: film_recommender_cg-1.0.6.tar.gz
  • Upload date:
  • Size: 29.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.10.13

File hashes

Hashes for film_recommender_cg-1.0.6.tar.gz
Algorithm Hash digest
SHA256 26c7be609d548dfa77afb98168e27c80981f482d2d84346680f0e24fb760ac39
MD5 acfcf7d42a4ff54cc053cc27557824e2
BLAKE2b-256 6b0799fb87d313fa1f9ab6adfe2a78a9f71c7b990820f379699d5747a553670d

See more details on using hashes here.

File details

Details for the file film_recommender_cg-1.0.6-py3-none-any.whl.

File metadata

File hashes

Hashes for film_recommender_cg-1.0.6-py3-none-any.whl
Algorithm Hash digest
SHA256 4deab8a8f4066e3d856981e2a3e1e0cafa9bd36cae0454813bc0ca3389f54ab4
MD5 1b3b9bc4757f995da0b57c13ecfde0d0
BLAKE2b-256 47c8f02139a244c44a638d04b1e336cb1900acf4619b77e73a489e00b606e27f

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page