Skip to main content

Statistical analysis of video game sales data with preprocessing, visualization, and predictive modeling

Project description

Sales Forecast ML Package

A Python package for reading, preprocessing, visualizing, and modeling sales data. It provides utilities to:

  • Load tabular sales data from CSV (read.read_data)
  • Clean and engineer features (preprocess.process_data, preprocess.prepare_data)
  • Visualize sales distributions by genre or platform (viz.print_genre_distribution, viz.print_platform_distribution)
  • Train and tune a Random Forest model and generate predictions (model.rf_fit, model.predict)

Overview

This package is designed to help you work with video game sales and activity data. It includes tools for reading raw CSV files, preprocessing and feature engineering, visualizing distributions, and building predictive models using Random Forest.


Features

  • CSV Reader: Reads a CSV file and applies basic type cleaning for year, drops an extraneous index column.
  • Data Processing: Aggregates per Name and combines sales/metadata, removes duplicates.
  • Feature Preparation: Multi-label binarization for Platform and Genre; concatenates with numeric predictors, drops missing Year.
  • Visualization:
    • print_genre_distribution(sales, genre, area): Histogram of sales for a given genre and region.
    • print_platform_distribution(sales, platform, area): Histogram of sales for a given platform and region.
  • Modeling: Grid search over Random Forest hyperparameters; evaluation prints R², RMSE (log scale), and top feature importances.
  • Prediction: Scales inputs consistently and returns predictions on the original scale via expm1.

Included Datasets

The package includes two sample datasets in the data/ directory for quick experimentation:

  • data/vgsales.csv: Raw sales data with columns like Name, Platform, Year, Genre, Publisher, regional sales, and Global_Sales.
  • data/game_data.csv: Steam activity metrics (all_time_peak, last_30_day_avg) for integration with sales data aggrigated with the vgsales dataset.

Use get_data_path(filename) from read.py to access these files programmatically.


Installation

pip install -U pip
pip install -U scikit-learn pandas numpy matplotlib seaborn
# If packaged for PyPI:
# pip install your-package-name

Quickstart

from your_package import (
    read_data, process_data, prepare_data,
    print_genre_distribution, print_platform_distribution,
    rf_fit, predict
)

# 1) Read
data = read_data("data/sales.csv")

# 2) Process & Prepare
sales_combined = process_data(data)
final_df = prepare_data(sales_combined)

# 3) Visualize
print_genre_distribution(data, genre="Action", area="Global_Sales")
print_platform_distribution(data, platform="PS4", area="Global_Sales")

# 4) Train model
best_model = rf_fit(final_df, area="Global_Sales")

# 5) Predict
new_data = final_df.drop(columns=["Global_Sales"]).iloc[:5]
preds = predict(best_model, area="Global_Sales", new_data=new_data)
print(preds)

API Reference

Top-Level Functions

  • read_data(file_path)
  • process_data(df)
  • prepare_data(df)
  • print_genre_distribution(sales, genre, area)
  • print_platform_distribution(sales, platform, area)
  • rf_fit(final_df, area)
  • predict(model, area, new_data)

Dependencies

  • pandas
  • numpy
  • scikit-learn
  • matplotlib
  • seaborn

Contributing

Issues and PRs are welcome. Please include reproducible examples and data schema.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

stat386_final-0.1.7.tar.gz (13.1 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

stat386_final-0.1.7-py3-none-any.whl (9.9 kB view details)

Uploaded Python 3

File details

Details for the file stat386_final-0.1.7.tar.gz.

File metadata

  • Download URL: stat386_final-0.1.7.tar.gz
  • Upload date:
  • Size: 13.1 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.1

File hashes

Hashes for stat386_final-0.1.7.tar.gz
Algorithm Hash digest
SHA256 80f037be8bd74028dbc961fcadf93c46418958f1899a5a937a51d4052353cbe4
MD5 814845aa77d5502ff83bf528c7bca0b4
BLAKE2b-256 2c6cff2d1a185b97ef65c9347710d6708282feaba4e31b36de218db74a14b24e

See more details on using hashes here.

File details

Details for the file stat386_final-0.1.7-py3-none-any.whl.

File metadata

  • Download URL: stat386_final-0.1.7-py3-none-any.whl
  • Upload date:
  • Size: 9.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.1

File hashes

Hashes for stat386_final-0.1.7-py3-none-any.whl
Algorithm Hash digest
SHA256 a11882327774d4024a2cd6344d6df59759c1178b3ea3ff5e3517ebf61928bcdb
MD5 204fac7f4c56328b60b1d7ba502802d8
BLAKE2b-256 669c427ebe7ff466618a14545c76f07688694783809c213abf8ecd62f771444e

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page