publicdata-au
Query and download Australian government open data from publicdata.au. publicdata.au republishes datasets that governments already publish under open licences, keeps every version at a URL that never changes, and serves each one in twelve formats with a query API.
This package works for every dataset the site serves, named by its slug, so a dataset added to the site needs no new release.
pip install publicdata-au # queries and downloads, no dependencies
pip install "publicdata-au[pandas]" # adds read() into a pandas DataFrame
pip install "publicdata-au[duckdb]" # adds connect() and relation(), DuckDB on a version
pip install "publicdata-au[geo]" # adds read_geo() into a geopandas GeoDataFrame
Find a dataset
import publicdata_au as pd_au
pd_au.datasets("road crashes") # slug, title, publisher, licence and page of each match
pd_au.datasets(topic="roads", jurisdiction="Qld")
pd_au.datasets(publisher="Bureau of Meteorology")
pd_au.versions("au-road-deaths") # every version kept, newest first
pd_au.fields("au-road-deaths") # each field's type, description, range and values
pd_au.catalogue("water quality", jurisdiction="Queensland") # every portal dataset, served or not
Query rows and totals
from publicdata_au import gte, in_
deaths = pd_au.rows(
"au-road-deaths",
{"state": in_("QLD", "NSW"), "year": gte(2020)},
select=["state", "year", "road_user"],
order="year.desc",
all=True,
)
deaths.version # the version the rows came from
deaths.attribution # the attribution the publisher's licence requires
deaths.to_pandas()
pd_au.aggregate("au-road-deaths", group="state", metric="count", where={"year": 2025})
Values come back typed as their fields say: dates as datetime.date, timestamps as
datetime.datetime and booleans as bool. A plain value must match exactly, a list matches any
of its values and None matches a blank or suppressed cell. The filters are eq, neq, gt, gte, lt, lte, like, ilike, in_,
is_null and not_.
Without version= an answer comes from the newest version and changes when the publisher
releases again. Pass a date from versions() for an answer that never changes.
Whole tables
df = pd_au.read("au-road-deaths") # needs the [pandas] extra
df.attrs["publicdata"] # the version, licence and attribution the file itself carries
pd_au.download("au-road-deaths", "csv") # or parquet, csv.gz, json, ndjson, sqlite, duckdb, ...
A version never changes once published, so a downloaded file can be kept and reused. Nothing is
kept unless you ask, with cache=True on a call or PUBLICDATA_CACHE=1 for every call:
df = pd_au.read("au-road-deaths", cache=True) # downloaded once, read from disk after
pd_au.cache_list() # what is kept, in cache_dir()
pd_au.cache_clear("au-road-deaths")
Query a version in place
Every version has a DuckDB file, and connect() attaches it read-only over HTTPS. Only the
blocks a query touches are read, so a count over millions of rows runs without a download.
con = pd_au.connect("au-road-deaths") # needs the [duckdb] extra
con.sql("SELECT state, count(*) FROM records GROUP BY 1").df()
con.publicdata # the version, its URL, licence, attribution and citation
r = pd_au.relation("au-road-deaths") # one table as a lazy DuckDB relation
r.filter("year >= 2020").aggregate("state, count(*) AS n").df()
A database such as G-NAF is one DuckDB file holding every table, the keys between them and the publisher's views, with one Parquet file per table beside it.
pd_au.tables("gnaf") # every table with its fields, keys and references
con = pd_au.connect("gnaf")
con.sql("SELECT postcode, count(*) AS n FROM address_view GROUP BY 1 ORDER BY 2 DESC LIMIT 10").df()
pd_au.read("gnaf", table="locality") # one table as a pandas DataFrame
pd_au.download("gnaf", table="state") # one table as Parquet
For heavy work on a large database, connect("gnaf", cache=True) downloads the file once and
queries it from disk. G-NAF's DuckDB file is about 3 GB.
Maps
Rows of a point dataset carry the codes of the ABS areas their point falls in: council area, SA2,
suburb, postal area and state and federal electorate. Total by one and join_boundaries()
attaches the boundaries, in GDA2020 (EPSG:7844). Any DataFrame with a column of codes works, with
layer= and by=.
by_sa2 = pd_au.aggregate("act-road-crashes", group="sa2_2021_code")
gdf = pd_au.join_boundaries(by_sa2.to_pandas()) # needs the [geo] extra
gdf.plot(column="count")
pd_au.boundary_layers() # the layers and their code fields
Datasets with a location or a shape have their own GeoPackage, which read_geo() reads in the
reference system the publisher used.
What changed
Each release is compared with the one before it on the dataset's key.
pd_au.changes("rba-money-market-daily") # one entry per release: added, removed, changed
pd_au.diff("rba-money-market-daily") # the newest release in full, with the keys
pd_au.provenance("au-road-deaths") # the source file, its checksum and fetch time
file_url(slug, format, version) gives a file's fixed address for another tool. A Client closes
the connections relation() opened when used as a context manager or with close(). Failing to
reach the site raises SiteUnreachable, a PublicDataError.
Files have no rate limit. The query API allows 60 requests in 10 seconds from one address, and this package waits and retries when it answers 429 or a passing server error.
Licence and attribution
The data is under each publisher's own licence, which requires the attribution string that every
answer carries. Please also name publicdata.au and link to the version you used; cite(slug) gives
the citation and cite(slug, format="bibtex") the BibTeX entry. Where a licence sets a condition
beyond attribution, such as G-NAF's rule on mail compilation, the package warns once per dataset
with a LicenceCondition warning. publicdata.au is
an independent republication, and the publishers have not endorsed it.
The package itself is under the MIT licence.
Metadata
Release files for publicdata-au 0.4.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| publicdata_au-0.4.0.tar.gz | 24.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| publicdata_au-0.4.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 43.2 kB
Release files / publicdata_au-0.4.0.tar.gz
| Download URL | publicdata_au-0.4.0.tar.gz |
|---|---|
| Size | 24.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
fa18273da0d677a7366d5b4648f2148edd559bd365c81eb42ee69ff9b50c7335
|
|
BLAKE2b-256 checksum How to use checksums |
0c01a152fbbb706660c96726b6f4cae0a421343324b0c30a9dfcca2953248861
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.
Transparency logRelease files / publicdata_au-0.4.0-py3-none-any.whl
| Download URL | publicdata_au-0.4.0-py3-none-any.whl |
|---|---|
| Size | 18.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
4ed0c2c92bef2f4c1acf6f5b464e2e18c3a7dd534d0c7cc7c8002ea282057be3
|
|
BLAKE2b-256 checksum How to use checksums |
6cff6dec02c0f203f79996e20ec8e08a6ce2e2f064201aa309fec38afc5115f6
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.
Transparency log