Skip to main content

A collection of jupyter notebook to learn machine learning

Project description

Machine Learning for Climate and Energy

A machine learning course oriented on climate science and energy sector applications and based on Python notebooks.

Course Philosophy

The amount of data produced in environmental sciences (energy and climate prediction) paves the way to new applications. While traditionnal statistical methods are still essential, advanced machine-learning methods are more and more needed to make sense of such big data, whether it is for analysis or to make predictions. One goal of machine learning is to extract identifiable patterns from these complex data sets. These patterns can then be used to take informed decisions. Another is to model relationships between different variables and then use these models to predict one variable from information of the other. Examples of such big data sets are:

  • If we want to optimize the energy performance of a building, we can set sensors in different places of the building that will give us a good overview of the energy consumption and energy loss. Analyzing such data set with machine learning will help us predict or optimize our energy consumption.
  • The IPCC provides forecast of the average temperature for the next 100 years. This forecast is based on about 30 different predictions made by complex models. Machine learning provides a way to track reliable patterns in this complex data set.

The objective of this course is to provide an introduction to statistical analysis and machine learning in order to help the students apply relevant methods to analyze specific data sets. Different families of unsupervised and supervised methods will be covered, while also providing a general approach to validate and test results. We encourage students to develop their critical thinking skills when facing a new dataset and applying a method in order to draw robust conclusions. We illustrate this course with examples from environmental sciences. These datasets correspond to concrete cases of analysis presented in the form of ipython notebook.

Course Topics

  1. Introduction,
  2. Supervised Learning Problem and Least Squares,
  3. Overfitting/Underfitting and Bias/Variance Decomposition,
  4. Regularization, Model Selection and Evaluation,
  5. Classification I: Generative models,
  6. Classification II: Discriminative models,
  7. Introduction to Unsupervised Learning with a Focus on Principal Component Analysis,
  8. Ensemble Methods for Regression,
  9. Introduction to Neural Networks,
  10. Neural Network Regularization and Introduction to Deep Learning.

Tutorials

After each course session of about 45 min, the students work on a tutorial for about 45 min during which they code the applications on real data, analyse the results and discuss them. The professors assist the students during the tutorials and make sure that all students are able to achieve the objectives of the tutorials.

Projects

Students by groups of two choose projects based on applications to environmental problems. The objective of the projects is to apply the statistical learning methodology from processing the intput data to evaluating the prediction skills of a model. The focus is on mastering the overall approach rather than on applying the most sophisticated models.

Technical support: All examples, tutorials and projects are based on the scikit-learn Python package. The students are thus asked to code in Python, but they do not need to be skilled in scikit-learn to join the course (see Prerequisites below). Students work either on a JupyterHub (priority), a school computer or on their personal laptop.

Prerequisites

  • Elementary data analysis in Python with numpy, pandas and matplotlib,
  • Linear algebra (linear systems, inverse, eigenvalues and eigenvectors),
  • Elements of probabilities and statistics (probability distribution and density, random variable, conditional expectation, variance, covariance, sample estimates).

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

ml4climate-25.0.1.tar.gz (60.4 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

ml4climate-25.0.1-py3-none-any.whl (60.7 MB view details)

Uploaded Python 3

File details

Details for the file ml4climate-25.0.1.tar.gz.

File metadata

  • Download URL: ml4climate-25.0.1.tar.gz
  • Upload date:
  • Size: 60.4 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: python-requests/2.32.5

File hashes

Hashes for ml4climate-25.0.1.tar.gz
Algorithm Hash digest
SHA256 7ecfdc886bf73c7155d403a211be5e5b8facf7d00c2eb30e80b38efc56c091fa
MD5 07a77525c7be1c10d24ac3697e4f868e
BLAKE2b-256 0d7094861f602e50a1dc85b463ccce803659454720904e027ca07c60175e2de6

See more details on using hashes here.

File details

Details for the file ml4climate-25.0.1-py3-none-any.whl.

File metadata

  • Download URL: ml4climate-25.0.1-py3-none-any.whl
  • Upload date:
  • Size: 60.7 MB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: python-requests/2.32.5

File hashes

Hashes for ml4climate-25.0.1-py3-none-any.whl
Algorithm Hash digest
SHA256 31e091888179498a54db8aef6472d58142925d1f0f127aeb555334098c392d14
MD5 edb83417bbe3796394f0023080d6d4a3
BLAKE2b-256 1a865164035b0c999bd5146b0f488e1cd0504df7d36df0311a6879c824e96941

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page