ATO benchmark comparison
+----------------------------------------------------------------------+
| ATO benchmark comparison |
+----------------------------------------------------------------------+
| Offline variance analysis against ATO benchmarks |
+----------------------------------+-----------------------------------+
| DR what it gives you | CR what it needs |
+----------------------------------+-----------------------------------+
| ATO benchmark ratio compare | P and L figures per account |
| accurate turnover definition | account to label mapping |
| shows the working per account | - |
+----------------------------------+-----------------------------------+
Distribution ato-benchmark-compare, import package atobenchmark, command ato-benchmark-compare.
Compare a set of profit and loss figures against the ATO small business benchmarks, on your own machine, with the working shown.
The ATO publishes benchmark ranges for 100 industries and uses them to pick which small businesses to look at more closely. Checking a client against them is a sensible thing to do before lodgement, and it is usually done by hand: find the industry page, work out which turnover range applies, add up the right accounts, and divide. This does that, and it records which accounts went into which figure so the answer can be checked by someone else later.
Australian tax rules only. Every figure comes from the ATO's own published dataset, and the comparison is a comparison, not advice.
The published Python distribution and command remain ato-benchmark-compare, and the
import package remains atobenchmark.
What it gets right
The arithmetic is not 'expenses over income'. The ATO defines these ratios narrowly, and the differences change the answer:
- Turnover is the sales of goods and services label, not total income. It falls back to total business income only when sales are blank, zero, or less than 50% of total business income.
- Total expenses for the ratio is total expenses less payments to associated persons. Wages to a spouse or an associated entity come out before the division.
- Cost of sales for the ratio excludes salary and wages, so the wages sitting in a bakery's cost of sales are moved out of the numerator and stay in total expenses.
- The key range is cost of sales to turnover where the ATO publishes one for that industry, otherwise total expenses to turnover.
- Turnover bands are treated as adjoining. The ATO prints
$65,000 - $400,000then$400,001 - $750,000, which read literally leaves $400,000.50 in no band at all.
Source: Australian Taxation Office, How we calculate benchmark ratios (QC 37143).
Excel workbook
No Python? workbooks/ato-benchmark-compare.xlsx
is the same comparison in ordinary worksheet formulas: paste the profit and loss,
review the mapping, pick the industry and read the result. It is macro-free, needs
desktop Excel for Microsoft 365 or Excel 2024, and is held to this engine's answer
by tests/test_workbook.py. See workbooks/README.md.
Install
git clone https://github.com/ryanduguid/australian-accounting.git
cd australian-accounting/packages/ato-benchmark-compare
pip install .
Python 3.10 or later. The runtime has no dependencies at all: the benchmark data ships inside the package and nothing is fetched at run time.
pip install ato-benchmark-compare installs the same package from PyPI; clone the
repository when you want the example files used below.
Use it
The flow is 2 commands, because the middle step is a person reading the ledger.
1. Draft a mapping from the profit and loss.
ato-benchmark-compare map --profit-and-loss examples/bakery-pnl.csv --out mapping.csv
That writes one row per account with a suggested bucket, the reason it was suggested, and the amount it read. Suggestions come from account names alone.
2. Review it. Open mapping.csv, fix the buckets, and change the source column
to reviewed. ato-benchmark-compare buckets explains each bucket. This is the step
that decides whether the answer is worth anything: no account name tells you whether
wages went to an associate.
Open it so the spreadsheet keeps the account columns as text. In Excel, use Data >
From Text/CSV, set the account and account_key columns to Text in the preview,
then load and save as CSV UTF-8. Opening the file by double-click instead reads a
numeric-looking ledger code as a number: -00123, +00123 and 00123 are saved
back as -123, 123 and 123 even when no account cell was touched, and the next
run refuses the file because the account no longer identifies the same account as its
account_key. The refusal is deliberate, so regenerate the mapping and reapply the
reviewed values rather than editing the key. The double-click round trip was
observed in Excel 16.0. The text import above was checked in the same Excel 16.0 by
driving the equivalent text-typed import over COM: -00123, +00123 and 00123
came back unchanged and the next compare run read every file. The dialog itself and
other spreadsheet applications have not been checked here, so confirm the account and
account_key columns are unchanged before the next run.
The generated mapping includes an account_key immediately after account. It is a
SHA-256 digest of the tool's existing case-and-whitespace-insensitive account identity.
Leave both columns unchanged while reviewing bucket, source and note; amount is
shown for context and is not bound by the key. On the next run, the key lets the tool
recover a formula-guarded logical account without confusing =cmd|calc with the genuine
account '=cmd|calc when both look the same in a spreadsheet.
account_key is an identity integrity check, not authentication or tamper resistance.
Anyone who can edit the file can recompute it, and case- or whitespace-only account edits
remain valid by design. Older mappings without the column remain readable for ordinary,
unambiguous account names. A legacy formula-like name or one that could already contain a
spreadsheet guard must be regenerated; reapply the reviewed bucket, source and note values
to the new mapping. Profit and loss input parsing strips and normalises leading whitespace,
so leading tab, carriage-return and newline prefixes are not distinct raw-ledger identities.
3. Compare.
ato-benchmark-compare compare \
--profit-and-loss examples/bakery-pnl.csv \
--mapping examples/bakery-mapping.csv \
--industry "Bakeries and hot bread shops"
ATO small business benchmark comparison
=======================================
Business type: Bakeries and hot bread shops
Benchmark year: 2023-24
Turnover: $850,000.00 (sales of goods and services)
Turnover band: More than $750,000
Ratio This business ATO range Result
---------------------------------------------------------------------------------
Cost of sales to turnover (key) 31.76% 29% to 36% within
Total expenses to turnover 83.17% 82% to 90% within
Labour to turnover 32.68% - no benchmark in this dataset
Rent to turnover 7.29% - no benchmark in this dataset
Motor vehicle expenses to turnover 1.12% - no benchmark in this dataset
Figures used
Sales of goods and services $850,000.00
Other business income $1,200.00
Total business income $851,200.00
Total expenses $751,950.00
Less payments to associates $45,000.00
Total expenses for the ratio $706,950.00
Cost of sales excluding wages $270,000.00
Labour $277,800.00
Add --json result.json for the same result as structured data, including every
bucket total and the source metadata.
Library callers that distinguish an omitted figure from an evidenced zero can use
atobenchmark.to_evidenced_dict(comparison, supplied_fields); include w1 in that
collection when it was supplied.
The serialiser masks ratios and prose whose required inputs were not supplied and
returns supplied_buckets, omitted_buckets, and complete_buckets alongside the
ordinary comparison payload.
compare() applies the same rule to its own verdicts. compute() records which
buckets the totals you passed actually held, and compare() reads that record
unless you pass supplied_fields yourself, so a ratio resting on a bucket you
omitted comes back with status not_supplied rather than within, below or
above, and outside_key_range stays false. Pass a bucket as Decimal("0") when
the operator established it is nil.
Totals from route() are the exception to reading the keys, because routing
zero-fills every bucket and its keys would then vouch for buckets no account was
mapped to. They carry the routed set themselves, so routing a file and comparing
it withholds the ratios resting on a bucket no reviewed account reached, with no
extra argument. Pass supplied_fields explicitly where the set you can vouch for
differs again, which is what the command line does: it adds w1 and the
operator's --confirm-other-income-nil assertion to the routed set.
Other commands:
ato-benchmark-compare industries --search cleaning # find the ATO business type
ato-benchmark-compare show "Bakeries and hot bread shops"
ato-benchmark-compare buckets
Input formats
The format this tool guarantees is 2 columns, with an optional section column of
income, cost_of_sales or expense:
account,amount
Sales,850000
Purchases,290000
A report style export, with a title block, section headings and subtotal rows, is
also read. Subtotal rows are detected and written into the mapping marked excluded
rather than dropped, so a total can never be quietly added to the figures it totals,
and nothing vanishes without appearing in a file you can read. Amounts are taken from
the column with the most cells that parse as amounts, keeping the leftmost on a tie, so
a comparative export can pick the prior period; --amount-column takes a column number
or a column heading to name the intended period.
The report style layout is inferred. It was written against the shape these
exports normally take, not verified against a real export from any particular
accounting package. Check the mapping file against your own export the first time,
and use --amount-column if it picked the wrong period.
Buckets
| Bucket | What the ATO does with it |
|---|---|
turnover |
Sales of goods and services. The turnover denominator. |
other_income |
Business income that is not sales. Only reaches turnover through the fallback rule. |
cost_of_sales |
Cost of sales, excluding wages inside it. |
cost_of_sales_labour |
Wages inside cost of sales. Kept out of the cost of sales ratio, kept in total expenses and labour. |
salary_wages |
Salary and wages outside cost of sales. |
contractor_commission |
Contractor, subcontractor and commission expenses. |
associated_persons |
Payments to associated persons. Deducted from total expenses. The wage buckets already exclude these, so labour deducts nothing further. |
rent |
Rent expenses. |
motor_vehicle |
Motor vehicle expenses. |
other_expense |
Every other expense, including superannuation and depreciation. |
excluded |
Outside the ATO calculation: income tax expense, subtotal rows. |
Exit codes
| Code | Meaning |
|---|---|
| 0 | Comparison produced, nothing to flag: key ratio inside the ATO range, or no published range applies to this turnover |
| 1 | Could not produce a comparison, for example an account with no mapping entry |
| 2 | Comparison produced, key ratio outside the ATO range |
| 3 | Comparison produced, but accounts still carry suggested buckets |
What it does not do
- It is not tax advice, and sitting outside a range is not a finding that anything is wrong. The ATO publishes ranges precisely because businesses differ.
- The bulk dataset the ATO publishes carries the 2 key ratios only. Labour, rent and motor vehicle ratios are calculated and shown, but the ranges for them are on the ATO's individual industry pages and are not in this dataset yet.
- Activity statement benchmarks are not covered. The ATO has not produced them since 1 July 2017.
- It reads a profit and loss. It does not read a tax return, so it cannot see the
W1 label, the salary and wages code, or anything else that only exists at lodgement.
Pass
--w1if you want the ATO's W1 rule applied to the labour ratio.
Those are the things the tool never attempts. Separately, LIMITATIONS.md records the cases where it does produce a figure that you should not read at face value, naming what stays correct in each.
Client data
Nothing leaves the machine. There is no network call anywhere in the runtime.
The .gitignore blocks the file names a real ledger arrives under, including
pnl*.csv, mapping*.csv, spreadsheets, and client-data/. The example files are
invented. Do not commit a real one.
Benchmark data
| Year | Business types | Source |
|---|---|---|
| 2023-24 | 100 | ATO Small Business Benchmarks, data.gov.au |
| 2022-23 | 100 | same dataset |
Each shipped file records the resource URL, the date it was retrieved, and the SHA-256 of the ATO workbook it was built from. The comparison prints all of that, so a report can be traced back to a specific published file.
To add a year when the ATO publishes one:
uv run --with openpyxl python tools/build_dataset.py \
--xlsx small-business-benchmarks-2024-25-data.xlsx \
--year 2024-25 \
--resource-name "2024-25 Benchmarks" \
--resource-url https://data.gov.au/... \
--resource-last-modified "<observed resource modification timestamp>" \
--retrieved "<actual retrieval date YYYY-MM-DD>" \
--out atobenchmark/data/benchmarks-2024-25.json
Replace the placeholders with the actual observed resource metadata for the build.
The builder refuses a workbook whose columns are not where it expects them, rather than quietly producing a dataset with the ratios in the wrong places.
Attribution
The benchmark figures are derived from Australian Taxation Office data, Small Business Benchmarks, used under CC BY 2.5 AU. The data has been converted from the ATO's spreadsheet into JSON, and turnover band bounds have been made adjoining as described above. No published ratio has been altered. The ATO has not endorsed this tool and has nothing to do with it.
The code in this repository is MIT licensed. The data attribution is also recorded in NOTICE, which ships inside the wheel and the sdist.
Author
Written by Ryan Duguid, a provisional member of Chartered Accountants ANZ, independently, in his own time and on his own equipment. Nothing here is the work of any employer, and no client data was used to build or test it.
Every ATO rule it implements was checked against the ATO's own published pages, and the shipped benchmark figures were cross checked against the ATO's industry page for the same industry.
Metadata
Release files for ato-benchmark-compare 0.1.9
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| ato_benchmark_compare-0.1.9.tar.gz | 103.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| ato_benchmark_compare-0.1.9-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 159.7 kB
Release files / ato_benchmark_compare-0.1.9.tar.gz
| Download URL | ato_benchmark_compare-0.1.9.tar.gz |
|---|---|
| Size | 103.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
18b8fe34454a2e3527fbcad24bd0a4e9001f3b81d684078bc3cbb39200a96ce5
|
|
BLAKE2b-256 checksum How to use checksums |
319c9bda907e5be1a4afa9692e725b246d44568d44519a49543e4b70caf778f6
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 22, 2026.
Transparency logRelease files / ato_benchmark_compare-0.1.9-py3-none-any.whl
| Download URL | ato_benchmark_compare-0.1.9-py3-none-any.whl |
|---|---|
| Size | 56.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
2d570b478167ba7f11e8d8a4b28c667ffd63fd24245bc8108f135c36cf2ffc66
|
|
BLAKE2b-256 checksum How to use checksums |
1421662fe27d5569e0c7f148648e3111f303aac4d22b383da35d16ad08a0945f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 22, 2026.
Transparency log