Skip to main content

read-acq

Read EDGES ACQ spectrum files.

Tests codecov Ruff

Installation

In a new/existing python environment, run pip install git+https://github.com/edges-collab/read_acq.

If you wish to develop read_acq, install it in editable mode. read_acq is built with meson-python, whose editable installs rebuild the C extension when the package is imported. That rebuild uses the build tools of the environment you install into, so install them first and turn off build isolation:

git clone https://github.com/edges-collab/read_acq
cd read_acq
pip install meson-python meson ninja numpy setuptools setuptools-scm
pip install --no-build-isolation --group dev -e .

(--group needs pip >= 25.1.) Without --no-build-isolation the install succeeds, but importing read_acq fails, because the build tools it used were in a temporary environment that no longer exists.

Usage

read_acq can be used from Python or from the command line.

Which reader do I need?

You want Use Reads from disk
A GSData object (what most analysis code works with) read_acq_to_gsdata, or GSData.from_file the spectra you select
Only part of a file: a time, LST or frequency range, or one load read_acq_to_gsdata(..., selectors=...) only the selected spectra and channels
Only the metadata and per-cycle ancillary data (times, ADC max/min, data drops) read_metadata a few hundred bytes per spectrum
The raw decoded arrays decode_file the whole file

All of these handle a file in the same way (see Damaged and incomplete files), and give the same data for it: the cheaper readers skip over what they don't need rather than reading it.

Reading into a GSData object

from read_acq import read_acq_to_gsdata

data = read_acq_to_gsdata("my_data.acq", telescope="edges-low")

Several files can be given, and they are joined along the time axis:

data = read_acq_to_gsdata(["my_data.acq", "my_data_2.acq"], telescope="edges-low")

read_acq also registers itself as the reader of .acq files for pygsdata, so GSData.from_file("my_data.acq") works too. See help(read_acq_to_gsdata) for all the options, e.g. lst_setter=read_acq.gsdata.fast_lst_setter for a much faster approximate calculation of the LSTs (within ~1 s of astropy's for the test data).

Reading only part of a file

Pass selectors to select data while reading. Only the selected spectra, and only the selected channels of each, are decoded, so this is much faster (and uses much less memory) than reading everything and then selecting:

from astropy import units as un

data = read_acq_to_gsdata(
    "my_data.acq",
    telescope="edges-low",
    selectors={
        "lst_selector": {"lst_range": (6, 12)},
        "freq_selector": {"freq_range": (50 * un.MHz, 100 * un.MHz)},
    },
)

The keys are freq_selector, time_selector, lst_selector and load_selector, and each maps to the keyword arguments of the matching function in pygsdata.select (select_freqs, select_times, select_lsts and select_loads). The result is exactly what reading everything and then applying those functions (in that order) gives. GSData.from_file("my_data.acq", selectors=...) does the same.

Reading only the metadata

To get only the metadata and the per-cycle ancillary data, use read_metadata. It never reads the spectra: it opens the file once and seeks from one spectrum to the next, reading only a few hundred bytes around each. This makes it cheap even on a network filesystem, where each read costs a round trip, so it is a good way to scan many files (e.g. to find which ones cover a given time range):

from read_acq import read_metadata

meta, ancillary = read_metadata("my_data.acq")
ncycles = len(ancillary["times"])

meta is a dict of the file's header and settings (nfreq, freq_min, freq_max, freq_res, nblk, temperature, ...). ancillary has an array for each of times (as YYYY:DDD:HH:MM:SS byte strings), adcmax, adcmin and, for fastspec files, data_drops, each with shape (ncycles, 3): one column for each switch position (antenna, load, load + noise source).

Reading the raw arrays

decode_file decodes the whole file into numpy arrays:

from read_acq import decode_file

q, (p0, p1, p2), anc = decode_file("my_data.acq", progress=False)

p0, p1 and p2 are the powers of the three switch positions (antenna, load, load + noise source), each with shape (nfreq, ncycles), and q = (p0 - p1) / (p2 - p1). anc holds the metadata (anc.meta), the per-cycle ancillary data (anc.data, as from read_metadata) and the frequencies in MHz (anc.frequencies).

Damaged and incomplete files

ACQ files are often cut short, e.g. when an acquisition is interrupted. All the readers handle this in the same way:

  • Only complete switch cycles (swpos 0, 1 and 2, in order) are returned. Reading starts at the first swpos-0 entry, and a cycle with a missing or bad entry is dropped.
  • A bad entry after the first (e.g. a spectrum of the wrong length) gives a warning, and reading carries on.
  • If the rest of a file is NUL bytes (an interrupted write), you get a warning and the complete cycles read before it.
  • A bad first swpos-0 entry, or a malformed comment line before it, is an error, as is a file with no comment lines at all.

Writing ACQ files

encode(filename, [p0, p1, p2], meta, ancillary) writes an ACQ file from the powers of the three switch positions, each with shape (ncycles, nfreq) (i.e. transposed from what decode_file returns), and metadata and ancillary data like those from read_metadata. For example, to write a copy of a file (the powers are stored to about 1 part in 106, so they come back very slightly changed):

from read_acq import encode

encode("copy.acq", [p0.T, p1.T, p2.T], anc.meta, anc.data)

write_gsdata_to_acq writes a GSData object to an ACQ file.

CLI

The acq convert command converts ACQ files to a single GSH5 file (pygsdata's HDF5 format):

acq convert data/2023_070*.acq

The input files can include glob-style patterns. Matching files are sorted and joined along the time axis. By default the output is named by the year and day of the first time in the data (e.g. 2023_070.gsh5), and is written in the current directory. See acq convert --help for the options.

Performance

benchmarks/ has scripts to measure the readers on synthetic files:

  • python benchmarks/bench_read.py times full and partial reads with read_acq_to_gsdata.
  • python benchmarks/bench_metadata.py reports, for each reader, the wall time, the number of times the file is opened and (on Linux) the number of reads and bytes read. On a network filesystem these counts, not the wall time on local disk, are what decide how long reading takes.

Metadata

Release files for read-acq 1.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for read-acq 1.3.0
File Size Uploaded
read_acq-1.3.0.tar.gz 836.0 kB Details

Release files / read_acq-1.3.0.tar.gz

Download URL read_acq-1.3.0.tar.gz
Size 836.0 kB
Tags Source
SHA-256 checksum
How to use checksums
7f7909dacf4b6176748294f16398a754be1b5416a5870a4fc21946de89ef7b4d
BLAKE2b-256 checksum
How to use checksums
c6f430f9fa346d6a63d82ef06383c47167560d101cb3e67fca0f3aebd8d5ee4c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release history Release notifications | RSS feed

This release

1.3.0 This release

1 release file

1.2.0

1 release file

1.1.1

1 release file

1.1.0

1 release file

1.0.2

1 release file

1.0.1

1 release file

1.0.0

1 release file

0.6.3

1 release file

0.6.2

1 release file

0.5.2

1 release file

0.5.0

1 release file

0.4.4

1 release file

0.4.3

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page