Skip to main content
Travis
Coveralls
PyPi
SemVer
Gitter

Generate and load Pandas data frames based on JSON Table Schema descriptors.

Version v0.2 contains breaking changes:

  • removed Storage(prefix=) argument (was a stub)

  • renamed Storage(tables=) to Storage(dataframes=)

  • renamed Storage.tables to Storage.buckets

  • changed Storage.read to read into memory

  • added Storage.iter to yield row by row

Getting Started

Installation

$ pip install datapackage
$ pip install jsontableschema-pandas

Example

You can easily load resources from a data package as Pandas data frames by simply using datapackage.push_datapackage function:

>>> import datapackage

>>> data_url = 'http://data.okfn.org/data/core/country-list/datapackage.json'
>>> storage = datapackage.push_datapackage(data_url, 'pandas')

>>> storage.buckets
['data___data']

>>> type(storage['data___data'])
<class 'pandas.core.frame.DataFrame'>

>>> storage['data___data'].head()
             Name Code
0     Afghanistan   AF
1   Åland Islands   AX
2         Albania   AL
3         Algeria   DZ
4  American Samoa   AS

Also it is possible to pull your existing data frame into a data package:

>>> datapackage.pull_datapackage('/tmp/datapackage.json', 'country_list', 'pandas', tables={
...     'data': storage['data___data'],
... })
Storage

Storage

Package implements Tabular Storage interface.

We can get storage this way:

>>> from jsontableschema_pandas import Storage

>>> storage = Storage()

Storage works as a container for Pandas data frames. You can define new data frame inside storage using storage.create method:

>>> storage.create('data', {
...     'primaryKey': 'id',
...     'fields': [
...         {'name': 'id', 'type': 'integer'},
...         {'name': 'comment', 'type': 'string'},
...     ]
... })

>>> storage.buckets
['data']

>>> storage['data'].shape
(0, 0)

Use storage.write to populate data frame with data:

>>> storage.write('data', [(1, 'a'), (2, 'b')])

>>> storage['data']
id comment
1        a
2        b

Also you can use tabulator to populate data frame from external data file:

>>> import tabulator

>>> with tabulator.Stream('data/comments.csv', headers=1) as stream:
...     storage.write('data', stream)

>>> storage['data']
id comment
1        a
2        b
1     good

As you see, subsequent writes simply appends new data on top of existing ones.

API Reference

Snapshot

https://github.com/frictionlessdata/jsontableschema-py#snapshot

Detailed

Contributing

Please read the contribution guideline:

How to Contribute

Thanks!

Release files for jsontableschema-pandas 0.5.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for jsontableschema-pandas 0.5.0
File Size Uploaded
jsontableschema-pandas-0.5.0.tar.gz 9.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for jsontableschema-pandas 0.5.0
File Interpreter ABI Platform
jsontableschema_pandas-0.5.0-py2.py3-none-any.whl Python 3, Python 2 none any Details

Total release size: 18.7 kB

Release files / jsontableschema-pandas-0.5.0.tar.gz

Download URL jsontableschema-pandas-0.5.0.tar.gz
Size 9.6 kB
Tags Source
SHA-256 checksum
How to use checksums
cf5833ebe4ddcab29f3c652304aff6a4316fa9b92d52d9ee047b9cffbb18cebf
BLAKE2b-256 checksum
How to use checksums
fae7f73e2c77418819cb704e0d5ec08d7507df37a29eb5e06f5b403a39d669c5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No

Release files / jsontableschema_pandas-0.5.0-py2.py3-none-any.whl

Download URL jsontableschema_pandas-0.5.0-py2.py3-none-any.whl
Size 9.1 kB
Tags Python 2 Python 3
SHA-256 checksum
How to use checksums
32895c32d83d644ca017b5d376911a9ff018b00f550fdd70d967e34da4c29a25
BLAKE2b-256 checksum
How to use checksums
9204808b627c8d314bee0e45a296e8810328d5f7bc1449a77c63f6e79812ed59
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No

Release history Release notifications | RSS feed

This release

0.5.0 This release

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.0.0

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page