Skip to main content

DocReader

A utility for reading a CSV file and mapping it’s contents to python objects based on the field names in the document.

Contains classes for reading loosly structured documents containing rows of text (eg. CSV files) and turning them into slightly higher level lists of dictionaries.

Typical usage is to create a subclass of DocReader and define the cols attribute once. To use it, instantiate the subclass with an open file-like object (that supports the iterator protocol and returns a full line of text on each next() call).

Iterating over the DocReader subclass instance will then yield a dict where the keys are as defined in the cols attribute and the values are the cell’s logical values that have passed through the correct conversion function.

The cols attribute is itself is either a list of column names or a dictionary mapping keys (which will be used as the keys in the final returned data) to an options dictionary for that column.

An options dictionary can have the following keys:
  • column: The name of the column as it appears in the document

  • convert: A callable that will be called with the string associated with

    the data in this cell. Often a type (eg int, float) or a lambda.

>>> import decimal
>>> def _price(val):
...     txt = val.replace('$', '').strip().lower()
...     txt = txt or None
...     return decimal.Decimal(txt)
...
>>> class MyReader(DocReader):
...     cols = dict(
...         a = dict(column='a', convert=int),
...         q = dict(column='b'),
...         c = dict(column='c', convert=lambda x: int(x)),
...         price = dict(convert=_price),
...     )
...
>>> import StringIO
>>> #Ignore the use of chr(10), doctest doesn't like \n.
>>> list(MyReader(StringIO.StringIO("a,b,c,price" + chr(10) + "1,2,3,$45.6" + chr(10) + "4,5,6,$55")))
[{u'a': 1, u'q': u'2', u'c': 3, u'price': Decimal('45.6')}, {u'a': 4, u'q': u'5', u'c': 6, u'price': Decimal('55')}]
>>> list(MyReader(StringIO.StringIO("a,b" + chr(10) + "1,2")))
Traceback (most recent call last):
...
MissingColumnException: Columns missing from input: c, price

Release files for docreader 1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for docreader 1.0
File Size Uploaded
docreader-1.0.tar.gz 3.4 kB Details

Release files / docreader-1.0.tar.gz

Download URL docreader-1.0.tar.gz
Size 3.4 kB
Tags Source
SHA-256 checksum
How to use checksums
e770f8742de05d10d1c848592a90db1803cf2f52aaad92dab92b8d8fac221361
BLAKE2b-256 checksum
How to use checksums
18330f8e08bb3b0b3e4488b0eb9f610ee1111b667f676b4ff3a3d4fbf4177090
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No

Release history Release notifications | RSS feed

This release

1.0 This release

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page