retriever

Data Retriever

Project description

Retriever logo

Finding data is one thing. Getting it ready for analysis is another. Acquiring, cleaning, standardizing and importing publicly available data is time consuming because many datasets lack machine readable metadata and do not conform to established data structures and formats. The Data Retriever automates the first steps in the data analysis pipeline by downloading, cleaning, and standardizing datasets, and importing them into relational databases, flat files, or programming languages. The automation of this process reduces the time for a user to get most large datasets up and running by hours, and in some cases days.

Installing the Current Release

If you have Python installed you can install the current release using either pip:

pip install retriever

or conda after adding the conda-forge channel (conda config --add channels conda-forge):

conda install retriever

Depending on your system configuration this may require sudo for pip:

sudo pip install retriever

Precompiled binary installers are also available for Windows, OS X, and Ubuntu/Debian on the releases page. These do not require a Python installation.

List of Available Datasets

Installing From Source

To install the Data Retriever from source, you'll need Python 3.6.8+ with the following packages installed:

xlrd

The following packages are optionally needed to interact with associated database management systems:

PyMySQL (for MySQL)
sqlite3 (for SQLite)
psycopg2-binary (for PostgreSQL), previously psycopg2.
pyodbc (for MS Access - this option is only available on Windows)
Microsoft Access Driver (ODBC for windows)

To install from source

Either use pip to install directly from GitHub:

pip install git+https://git@github.com/weecology/retriever.git

or:

Clone the repository
From the directory containing setup.py, run the following command: pip install .. You may need to include sudo at the beginning of the command depending on your system (i.e., sudo pip install .).

More extensive documentation for those that are interested in developing can be found here

Using the Command Line

After installing, run retriever update to download all of the available dataset scripts. To see the full list of command line options and datasets run retriever --help. The output will look like this:

usage: retriever [-h] [-v] [-q]
                 {download,install,defaults,update,new,new_json,edit_json,delete_json,ls,citation,reset,help}
                 ...

positional arguments:
  {download,install,defaults,update,new,new_json,edit_json,delete_json,ls,citation,reset,help}
                        sub-command help
    download            download raw data files for a dataset
    install             download and install dataset
    defaults            displays default options
    update              download updated versions of scripts
    new                 create a new sample retriever script
    new_json            CLI to create retriever datapackage.json script
    edit_json           CLI to edit retriever datapackage.json script
    delete_json         CLI to remove retriever datapackage.json script
    ls                  display a list all available dataset scripts
    citation            view citation
    reset               reset retriever: removes configuration settings,
                        scripts, and cached data
    help

optional arguments:
  -h, --help            show this help message and exit
  -v, --version         show program's version number and exit
  -q, --quiet           suppress command-line output

To install datasets, use retriever install:

usage: retriever install [-h] [--compile] [--debug]
                         {mysql,postgres,sqlite,msaccess,csv,json,xml} ...

positional arguments:
  {mysql,postgres,sqlite,msaccess,csv,json,xml}
                        engine-specific help
    mysql               MySQL
    postgres            PostgreSQL
    sqlite              SQLite
    msaccess            Microsoft Access
    csv                 CSV
    json                JSON
    xml                 XML

optional arguments:
  -h, --help            show this help message and exit
  --compile             force re-compile of script before downloading
  --debug               run in debug mode

Examples

These examples are using the Iris flower dataset. More examples can be found in the Data Retriever documentation.

Using Install

retriever install -h   (gives install options)

Using specific database engine, retriever install {Engine}

retriever install mysql -h     (gives install mysql options)
retriever install mysql --user myuser --password ******** --host localhost --port 8888 --database_name testdbase iris

install data into an sqlite database named iris.db you would use:

retriever install sqlite iris -f iris.db

Using download

retriever download -h    (gives you help options)
retriever download iris
retriever download iris --path C:\Users\Documents

Using citation

retriever citation   (citation of the retriever engine)
retriever citation iris  (citation for the iris data)

Spatial Dataset Installation

Set up Spatial support

To set up spatial support for Postgres using Postgis please refer to the spatial set-up docs.

retriever install postgres harvard-forest # Vector data
retriever install postgres bioclim # Raster data
# Install only the data of USGS elevation in the given extent
retriever install postgres usgs-elevation -b -94.98704597353938 39.027001800158615 -94.3599408119917 40.69577051867074

Website

For more information see the Data Retriever website.

Acknowledgments

Development of this software was funded by the Gordon and Betty Moore Foundation's Data-Driven Discovery Initiative through Grant GBMF4563 to Ethan White and the National Science Foundation as part of a CAREER award to Ethan White.

Project details

Release history Release notifications | RSS feed

This version

3.1.0

Apr 27, 2022

3.0.0

Jul 16, 2020

2.4.0

Jun 11, 2019

2.3.1

May 1, 2019

2.3.0

May 1, 2019

2.2.0

Nov 6, 2018

2.1.0

Oct 27, 2017

2.0.0

Feb 24, 2017

1.8.3

Feb 12, 2016

1.7.0

Oct 5, 2014

1.6

Mar 23, 2014

1.4

Dec 3, 2012

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

retriever-3.1.0.tar.gz (90.0 kB view details)

Uploaded Apr 27, 2022 Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

retriever-3.1.0-py2.py3-none-any.whl (85.3 kB view details)

Uploaded Apr 27, 2022 Python 2Python 3

File details

Details for the file retriever-3.1.0.tar.gz.

File metadata

Download URL: retriever-3.1.0.tar.gz
Upload date: Apr 27, 2022
Size: 90.0 kB
Tags: Source
Uploaded using Trusted Publishing? No
Uploaded via: twine/4.0.0 CPython/3.9.12

File hashes

Hashes for retriever-3.1.0.tar.gz
Algorithm	Hash digest
SHA256	`d025ed69e006deefe7a4d98323d46a28f75dcddad47033d90cb4956aa01cd479`
MD5	`04803cbe6d26ac394868458cb652100c`
BLAKE2b-256	`b03b889cfd23e203ca2aaafb58e3fd2680f9987176f5561c43e0eeee7cb4529b`

See more details on using hashes here.

File details

Details for the file retriever-3.1.0-py2.py3-none-any.whl.

File metadata

Download URL: retriever-3.1.0-py2.py3-none-any.whl
Upload date: Apr 27, 2022
Size: 85.3 kB
Tags: Python 2, Python 3
Uploaded using Trusted Publishing? No
Uploaded via: twine/4.0.0 CPython/3.9.12

File hashes

Hashes for retriever-3.1.0-py2.py3-none-any.whl
Algorithm	Hash digest
SHA256	`ac3e7b234597d7a0f7963c84d7cf502158ba822070767aa14255cc453975fbb2`
MD5	`16ab85f7f2567babea39032e4d557823`
BLAKE2b-256	`ea699ec55346ffd184de708f2ebf23374a7a540b107cab75e5df578c08c9bb7d`

See more details on using hashes here.

retriever 3.1.0

Navigation

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Project description

Installing the Current Release

List of Available Datasets

Installing From Source

To install from source

Using the Command Line

Examples

Spatial Dataset Installation

Website

Acknowledgments

Project details

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

File details

File metadata

File hashes