Skip to main content

kagglehub

The kagglehub library provides a simple way to interact with Kaggle resources such as datasets, models, notebook outputs in Python.

This library also integrates natively with the Kaggle notebook environment. This means the behavior differs when you download a Kaggle resource with kagglehub in the Kaggle notebook environment:

  • In a Kaggle notebook:
    • The resource is automatically attached to your Kaggle notebook.
    • The resource will be shown under the "Input" panel in the Kaggle notebook editor.
    • The resource files are served from the shared Kaggle resources cache (not using the VM's disk).
  • Outside a Kaggle notebook:
    • The resource files are downloaded to a local cache folder.

Installation

Install the kagglehub package with pip:

pip install kagglehub

Usage

Authenticate

[!NOTE] kagglehub is authenticated by default when running in a Kaggle notebook.

Authenticating is only needed to access public resources requiring user consent or private resources.

First, you will need a Kaggle account. You can sign up here.

After login, you can download your Kaggle API token at https://www.kaggle.com/settings/api by clicking on the "Generate New Token" button.

You have several options to authenticate. Note that if you use kaggle-api (the kaggle command-line tool) you have already configured authentication and can skip this.

Option 1: kagglehub.login()

This will prompt you to enter your Kaggle API token:

import kagglehub

kagglehub.login()

Option 2: Environment variable

You can also choose to export your Kaggle token to the environment:

export KAGGLE_API_TOKEN=xxxxxxxxxxxxxx # Copied from the settings UI

Option 3: API token file

Store your Kaggle API token obtained from your Kaggle account API tokens settings page in a file at ~/.kaggle/access_token.

Option 4: Google Colab secret

Store your Kaggle API token obtained from your Kaggle account API tokens settings page in a Colab secret named KAGGLE_API_TOKEN.

Instructions on adding secrets in both Colab and Colab Enterprise can be found in this article.

Option 5: Legacy API credentials file

From your Kaggle account API tokens settings page, under "Legacy API Credentials", click on the "Create Legacy API Key" button to generate a kaggle.json file and store it at ~/.kaggle/kaggle.json.

Download Model

The following examples download the answer-equivalence-bem variation of this Kaggle model: https://www.kaggle.com/models/google/bert/tensorFlow2/answer-equivalence-bem

import kagglehub

# Download the latest version.
kagglehub.model_download('google/bert/tensorFlow2/answer-equivalence-bem')

# Download a specific version.
kagglehub.model_download('google/bert/tensorFlow2/answer-equivalence-bem/1')

# Download a single file.
kagglehub.model_download('google/bert/tensorFlow2/answer-equivalence-bem', path='variables/variables.index')

# Download a model or file, even if previously downloaded to cache.
kagglehub.model_download('google/bert/tensorFlow2/answer-equivalence-bem', force_download=True)

# Download to a custom local directory.
kagglehub.model_download('google/bert/tensorFlow2/answer-equivalence-bem', output_dir='./models')

# Overwrite an existing output directory.
kagglehub.model_download('google/bert/tensorFlow2/answer-equivalence-bem', output_dir='./models', force_download=True)

Upload Model

Uploads a new variation (or a new variation's version if it already exists).

import kagglehub

# For example, to upload a new variation to this model:
# - https://www.kaggle.com/models/google/bert/tensorFlow2/answer-equivalence-bem
# 
# You would use the following handle: `google/bert/tensorFlow2/answer-equivalence-bem`
handle = '<KAGGLE_USERNAME>/<MODEL>/<FRAMEWORK>/<VARIATION>'
local_model_dir = 'path/to/local/model/dir'

kagglehub.model_upload(handle, local_model_dir)

# You can also specify some version notes (optional)
kagglehub.model_upload(handle, local_model_dir, version_notes='improved accuracy')

# You can also specify a license (optional)
kagglehub.model_upload(handle, local_model_dir, license_name='Apache 2.0')

# You can also specify a list of patterns for files/dirs to ignore.
# These patterns are combined with `kagglehub.models.DEFAULT_IGNORE_PATTERNS` 
# to determine which files and directories to exclude. 
# To ignore entire directories, include a trailing slash (/) in the pattern.
kagglehub.model_upload(handle, local_model_dir, ignore_patterns=["original/", "*.tmp"])

Load Dataset

Loads a file from a Kaggle Dataset into a python object based on the selected KaggleDatasetAdapter:

NOTE: To use these adapters, you must install the optional dependencies (or already have them available in your environment)

  • KaggleDatasetAdapter.PANDASpip install kagglehub[pandas-datasets]
  • KaggleDatasetAdapter.HUGGING_FACEpip install kagglehub[hf-datasets]
  • KaggleDatasetAdapter.POLARSpip install kagglehub[polars-datasets]

KaggleDatasetAdapter.PANDAS

This adapter supports the following file types, which map to a corresponding pandas.read_* method:

File Extension pandas Method
.csv, .tsv1 pandas.read_csv
.json, .jsonl2 pandas.read_json
.xml pandas.read_xml
.parquet pandas.read_parquet
.feather pandas.read_feather
.sqlite, .sqlite3, .db, .db3, .s3db, .dl33 pandas.read_sql_query
.xls, .xlsx, .xlsm, .xlsb, .odf, .ods, .odt4 pandas.read_excel
  1. For TSV files, \t is automatically supplied for the sep parameter, but may be overridden with pandas_kwargs

  2. For JSONL files, True is supplied for the lines parameter

  3. For SQLite files, a sql_query must be provided to generate the DataFrame(s)

  4. dataset_load also supports pandas_kwargs which will be passed as keyword arguments to the pandas.read_* method. Some examples include:

    import kagglehub
    from kagglehub import KaggleDatasetAdapter
    
    # Load a DataFrame with a specific version of a CSV
    df = kagglehub.dataset_load(
        KaggleDatasetAdapter.PANDAS,
        "unsdsn/world-happiness/versions/1",
        "2016.csv",
    )
    
    # Load a DataFrame with specific columns from a parquet file
    df = kagglehub.dataset_load(
        KaggleDatasetAdapter.PANDAS,
        "robikscube/textocr-text-extraction-from-images-dataset",
        "annot.parquet",
        pandas_kwargs={"columns": ["image_id", "bbox", "points", "area"]}
    )
    
    # Load a dictionary of DataFrames from an Excel file where the keys are sheet names 
    # and the values are DataFrames for each sheet's data. NOTE: As written, this requires 
    # installing the default openpyxl engine.
    df_dict = kagglehub.dataset_load(
        KaggleDatasetAdapter.PANDAS,
        "theworldbank/education-statistics",
        "edstats-excel-zip-72-mb-/EdStatsEXCEL.xlsx",
        pandas_kwargs={"sheet_name": None},
    )
    
    # Load a DataFrame using an XML file (with the natively available etree parser)
    df = dataset_load(
        KaggleDatasetAdapter.PANDAS,
        "parulpandey/covid19-clinical-trials-dataset",
        "COVID-19 CLinical trials studies/COVID-19 CLinical trials studies/NCT00571389.xml",
        pandas_kwargs={"parser": "etree"},
    )
    
    # Load a DataFrame by executing a SQL query against a SQLite DB
    df = kagglehub.dataset_load(
        KaggleDatasetAdapter.PANDAS,
        "wyattowalsh/basketball",
        "nba.sqlite",
        sql_query="SELECT person_id, player_name FROM draft_history",
    )
    

    KaggleDatasetAdapter.HUGGING_FACE

    The Hugging Face Dataset provided by this adapater is built exclusively using Dataset.from_pandas. As a result, all of the file type and pandas_kwargs support is the same as KaggleDatasetAdapter.PANDAS. Some important things to note about this:

    1. Because Dataset.from_pandas cannot accept a collection of DataFrames, any attempts to load a file with pandas_kwargs that produce a collection of DataFrames will result in a raised exception
    2. hf_kwargs may be provided, which will be passed as keyword arguments to Dataset.from_pandas
    3. Because the use of pandas is transparent when pandas_kwargs are not needed, we default to False for preserve_index—this can be overridden using hf_kwargs

    Some examples include:

    import kagglehub
    from kagglehub import KaggleDatasetAdapter
    # Load a Dataset with a specific version of a CSV, then remove a column
    dataset = kagglehub.dataset_load(
        KaggleDatasetAdapter.HUGGING_FACE,
        "unsdsn/world-happiness/versions/1",
        "2016.csv",
    )
    dataset = dataset.remove_columns('Region')
    
    # Load a Dataset with specific columns from a parquet file, then split into test/train splits
    dataset = kagglehub.dataset_load(
        KaggleDatasetAdapter.HUGGING_FACE,
        "robikscube/textocr-text-extraction-from-images-dataset",
        "annot.parquet",
        pandas_kwargs={"columns": ["image_id", "bbox", "points", "area"]}
    )
    dataset_with_splits = dataset.train_test_split(test_size=0.8, train_size=0.2)
    
    # Load a Dataset by executing a SQL query against a SQLite DB, then rename a column
    dataset = kagglehub.dataset_load(
        KaggleDatasetAdapter.HUGGING_FACE,
        "wyattowalsh/basketball",
        "nba.sqlite",
        sql_query="SELECT person_id, player_name FROM draft_history",
    )
    dataset = dataset.rename_column('season', 'year')
    

    KaggleDatasetAdapter.POLARS

    This adapter supports the following file types, which map to a corresponding polars.scan_* or polars.read_* method:

    File Extension polars Method
    .csv, .tsv1 polars.scan_csv or polars.read_csv
    .json polars.read_json
    .jsonl polars.scan_ndjson or polars.read_ndjson
    .parquet polars.scan_parquet or polars.read_parquet
    .feather polars.scan_ipc or polars.read_ipc
    .sqlite, .sqlite3, .db, .db3, .s3db, .dl32 polars.read_database
    .xls, .xlsx, .xlsm, .xlsb, .odf, .ods, .odt3 polars.read_excel

    dataset_load also supports polars_kwargs which will be passed as keyword arguments to the polars.scan_* or polars_read_* method.

    LazyFrame vs DataFrame

    Per polars documentation, LazyFrame "allows for whole-query optimisation in addition to parallelism, and is the preferred (and highest-performance) mode of operation for polars." As such, scan_* methods are used by default whenever possible--and when not possible the result of the read_* method is returned after calling .lazy(). If a DataFrame is preferred, dataset_load supports an optional polars_frame_type and PolarsFrameType.DATA_FRAME may be passed in. This will force a read_* method to be used with no .lazy() call. NOTE: For file types that support scan_*, changing the polars_frame_type may affect which polars_kwargs are acceptable to the underlying method since it will force a read_* method to be used rather than a scan_* method.

    Some examples include:

    import kagglehub
    from kagglehub import KaggleDatasetAdapter, PolarsFrameType
    
    # Load a LazyFrame with a specific version of a CSV
    lf = kagglehub.dataset_load(
        KaggleDatasetAdapter.POLARS,
        "unsdsn/world-happiness/versions/1",
        "2016.csv",
    )
    
    # Load a LazyFramefrom a parquet file, then select specific columns
    lf = kagglehub.dataset_load(
        KaggleDatasetAdapter.POLARS,
        "robikscube/textocr-text-extraction-from-images-dataset",
        "annot.parquet",
    )
    lf.select(["image_id", "bbox", "points", "area"]).collect()
    
    # Load a DataFrame with specific columns from a parquet file
    df = kagglehub.dataset_load(
        KaggleDatasetAdapter.POLARS,
        "robikscube/textocr-text-extraction-from-images-dataset",
        "annot.parquet",
        polars_frame_type=PolarsFrameType.DATA_FRAME,
        polars_kwargs={"columns": ["image_id", "bbox", "points", "area"]}
    )
    
    # Load a dictionary of LazyFrames from an Excel file where the keys are sheet names 
    # and the values are LazyFrames for each sheet's data. NOTE: As written, this requires 
    # installing the default fastexcel engine.
    lf_dict = kagglehub.dataset_load(
        KaggleDatasetAdapter.POLARS,
        "theworldbank/education-statistics",
        "edstats-excel-zip-72-mb-/EdStatsEXCEL.xlsx",
        # sheet_id of 0 returns all sheets
        polars_kwargs={"sheet_id": 0},
    )
    
    # Load a LazyFrame by executing a SQL query against a SQLite DB
    lf = kagglehub.dataset_load(
        KaggleDatasetAdapter.POLARS,
        "wyattowalsh/basketball",
        "nba.sqlite",
        sql_query="SELECT person_id, player_name FROM draft_history",
    )
    

    Download Dataset

    The following examples download the Spotify Recommendation Kaggle dataset: https://www.kaggle.com/datasets/bricevergnou/spotify-recommendation

    import kagglehub
    
    # Download the latest version.
    kagglehub.dataset_download('bricevergnou/spotify-recommendation')
    
    # Download a specific version.
    kagglehub.dataset_download('bricevergnou/spotify-recommendation/versions/1')
    
    # Download a single file.
    kagglehub.dataset_download('bricevergnou/spotify-recommendation', path='data.csv')
    
    # Download a dataset or file, even if previously downloaded to cache.
    kagglehub.dataset_download('bricevergnou/spotify-recommendation', force_download=True)
    
    # Download a dataset to a custom output directory.
    kagglehub.dataset_download('bricevergnou/spotify-recommendation', output_dir='./data')
    
    # Download a single file to a custom output directory.
    kagglehub.dataset_download('bricevergnou/spotify-recommendation', path='data.csv', output_dir='./data')
    
    # Overwrite an existing output directory.
    kagglehub.dataset_download('bricevergnou/spotify-recommendation', output_dir='./data', force_download=True)
    

    Upload Dataset

    Uploads a new dataset (or a new version if it already exists).

    import kagglehub
    
    # For example, to upload a new dataset (or version) at:
    # - https://www.kaggle.com/datasets/bricevergnou/spotify-recommendation
    # 
    # You would use the following handle: `bricevergnou/spotify-recommendation`
    handle = '<KAGGLE_USERNAME>/<DATASET>'
    local_dataset_dir = 'path/to/local/dataset/dir'
    
    # Create a new dataset
    kagglehub.dataset_upload(handle, local_dataset_dir)
    
    # You can then create a new version of this existing dataset and include version notes (optional).
    kagglehub.dataset_upload(handle, local_dataset_dir, version_notes='improved data')
    
    # You can also specify a list of patterns for files/dirs to ignore.
    # These patterns are combined with `kagglehub.datasets.DEFAULT_IGNORE_PATTERNS` 
    # to determine which files and directories to exclude. 
    # To ignore entire directories, include a trailing slash (/) in the pattern.
    kagglehub.dataset_upload(handle, local_dataset_dir, ignore_patterns=["original/", "*.tmp"])
    

    Download Competition

    The following examples download the Digit Recognizer Kaggle competition: https://www.kaggle.com/competitions/digit-recognizer

    import kagglehub
    
    # Download the latest version.
    kagglehub.competition_download('digit-recognizer')
    
    # Download a single file.
    kagglehub.competition_download('digit-recognizer', path='train.csv')
    
    # Download a competition or file, even if previously downloaded to cache. 
    kagglehub.competition_download('digit-recognizer', force_download=True)
    
    # Download competition data to a custom output directory.
    kagglehub.competition_download('digit-recognizer', output_dir='./competition')
    
    # Overwrite an existing output directory.
    kagglehub.competition_download('digit-recognizer', output_dir='./competition', force_download=True)
    

    Download Notebook Outputs

    The following examples download the Titanic Tutorial notebook output: https://www.kaggle.com/code/alexisbcook/titanic-tutorial

    import kagglehub
    
    # Download the latest version.
    kagglehub.notebook_output_download('alexisbcook/titanic-tutorial')
    
    # Download a specific version of the notebook output.
    kagglehub.notebook_output_download('alexisbcook/titanic-tutorial/versions/1')
    
    # Download a single file.
    kagglehub.notebook_output_download('alexisbcook/titanic-tutorial', path='submission.csv')
    
    # Download notebook output to a custom output directory.
    kagglehub.notebook_output_download('alexisbcook/titanic-tutorial', output_dir='./output')
    
    # Overwrite an existing output directory.
    kagglehub.notebook_output_download('alexisbcook/titanic-tutorial', output_dir='./output', force_download=True)
    

    Install Utility Script

    The following example installs the utility script Physionet Challenge Utility Script Utility Script: https://www.kaggle.com/code/bjoernjostein/physionet-challenge-utility-script. Using this command allows the code from this script to be available in your python environment.

    import kagglehub
    
    # Install the latest version.
    kagglehub.utility_script_install('bjoernjostein/physionet-challenge-utility-script')
    

    Options

    Change the default cache folder

    By default, kagglehub downloads files to your home folder at ~/.cache/kagglehub/.

    You can override this path by setting the KAGGLEHUB_CACHE environment variable.

    Development

    Prequisites

    We use hatch to manage this project.

    Follow these instructions to install it.

    Tests

    # Run all tests for current Python version.
    hatch test
    
    # Run all tests for all Python versions.
    hatch test --all
    
    # Run all tests for a specific Python version.
    hatch test -py 3.11
    
    # Run a single test file
    hatch test tests/test_<SOME_FILE>.py
    

    Integration Tests

    To run integration tests on your local machine, you need to set up your Kaggle API credentials. You can do this in one of these two ways described in the earlier sections of this document. Refer to the sections:

    After setting up your credentials by any of these methods, you can run the integration tests as follows:

    # Run all tests
    hatch test integration_tests
    

    Run kagglehub from source

    Option 1: Execute a one-liner of code from the command line

    # Download a model & print the path
    hatch run python -c "import kagglehub; print('path: ', kagglehub.model_download('google/bert/tensorFlow2/answer-equivalence-bem'))"
    

    Option 2: Run a saved script from the /tools/scripts directory

    # This runs the same code as the one-liner above, but reads it from a 
    # checked in script located at tool/scripts/download_model.py
    hatch run python tools/scripts/download_model.py
    

    Option 3: Run a temporary script from the root of the repo

    Any script created at the root of the repo is gitignore'd, so they're just temporary scripts for testing in development. Placing temporary scripts at the root makes the run command easier to use during local development.

    # Test out some new changes
    hatch run python test_new_feature.py
    

    Lint / Format

    # Lint check
    hatch run lint:style
    hatch run lint:typing
    hatch run lint:all     # for both
    
    # Format
    hatch run lint:fmt
    

    Coverage report

    hatch test --cover
    

    Build

    hatch build
    

    Running hatch commands inside Docker

    This is useful to run in a consistent environment and easily switch between Python versions.

    The following shows how to run hatch run lint:all but this also works for any other hatch commands:

    # Use default Python version
    ./docker-hatch run lint:all
    
    # Use specific Python version (Must be a valid tag from: https://hub.docker.com/_/python)
    ./docker-hatch -v 3.10 run lint:all
    
    # Run test in docker with specific Python version
    ./docker-hatch -v 3.10 test
    
    # Run python from specific environment (e.g. one with optional dependencies installed)
    ./docker-hatch run extra-deps-env:python -c "print('hello world')"
    
    # Run commands with other root-level hatch options (everything after -- gets passed to hatch)
    ./docker-hatch -v 3.10 -- -v env create debug-env-with-verbose-logging
    

    VS Code setup

    Prerequisites

    Install the recommended extensions.

    Instructions

    Configure hatch to create virtual env in project folder.

    hatch config set dirs.env.virtual .env
    

    After, create all the python environments needed by running hatch test --all.

    Finally, configure vscode to use one of the selected environments: cmd + shift + p -> python: Select Interpreter -> Pick one of the folders in ./.env

    Support

    The kagglehub library has configured automatic logging for console. For file based logging, setting the KAGGLE_LOGGING_ENABLED=1 environment variable will output logs to a directory. The default log destination is resolved via the os.path.expanduser

    The table below contains possible locations:

    os log path
    osx /user/$USERNAME/.kaggle/logs/kagglehub.log
    linux ~/.kaggle/logs/kagglehub.log
    windows C:\Users\%USERNAME%\.kaggle\logs\kagglehub.log

    If needed, the root log directory can be overriden using the following environment variable: KAGGLE_LOGGING_ROOT_DIR

    Please include the log to help troubleshoot issues.

    Contributing

    If you'd like to contribute to kagglehub, please make sure to take a look at CONTRIBUTING.md.

  5. For TSV files, \t is automatically supplied for the separator parameter, but may be overridden with polars_kwargs 2

  6. For SQLite files, a sql_query must be provided to generate the DataFrame(s) 2

  7. The specific file extension may dictate which optional engine dependency needs to be installed to read the file 2

  8. The specific file extension will dictate which optional engine dependency needs to be installed to read the file

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

kagglehub-1.0.2.tar.gz (117.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

kagglehub-1.0.2-py3-none-any.whl (70.6 kB view details)

Uploaded Python 3

File details

Details for the file kagglehub-1.0.2.tar.gz.

File metadata

  • Download URL: kagglehub-1.0.2.tar.gz
  • Upload date:
  • Size: 117.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.0.1 CPython/3.10.13

File hashes

Hashes for kagglehub-1.0.2.tar.gz
Algorithm Hash digest
SHA256 08abae0c1e249904003ce6f23c189166e3032220478426a552c77c2edb08b93a
MD5 65237631a341c4e8fbecf087b37c0513
BLAKE2b-256 b5b9e1bdafdcdb98e8e5354b30959ed48f4a08ca2c2dc72c968c5b38c4baa1b5

See more details on using hashes here.

File details

Details for the file kagglehub-1.0.2-py3-none-any.whl.

File metadata

  • Download URL: kagglehub-1.0.2-py3-none-any.whl
  • Upload date:
  • Size: 70.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.0.1 CPython/3.10.13

File hashes

Hashes for kagglehub-1.0.2-py3-none-any.whl
Algorithm Hash digest
SHA256 ffad2d79cfda9b06848e98cb0a4a48e286285cfba4aeadadc0b8dbd9bcadd372
MD5 80c57def3353ad7757766218d50c836b
BLAKE2b-256 611a6b1293d0b091905e367f86771a7122e5f7f2729ab9f0f36a1617f7ad6c05

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page