Skip to main content

Arcane bigquery-storage

A thin wrapper around the BigQuery Storage Write API. The Storage Write API uses protocol buffers to describe and serialize rows, so to write to a table you bring a compiled protobuf message class that mirrors the table schema.

The flow is always the same:

  1. Write a .proto that mirrors your BigQuery table schema.
  2. Compile it with protoc to get a *_pb2.py module.
  3. Use it with Client.create_proto_descriptor / Client.create_row / Client.write_rows.

The next three sections walk through each step.

1. Write a .proto

Create a file like my_table.proto next to your code. Each field must match a column in your BigQuery table by name; the proto type must be compatible with the BigQuery column type (see the protobuf → BigQuery type mapping).

syntax = "proto2";

message MyRow {
    optional string name  = 1;
    optional int64  value = 2;
}

A couple of rules worth knowing:

  • Field names in the proto must equal column names in BigQuery.
  • Field numbers (the = 1, = 2 tags) identify fields on the wire. Never re-use a number for a different field — that breaks compatibility with any rows already written.
  • proto2 makes every field optional by default, which mirrors BigQuery's NULLABLE columns nicely. proto3 works too; just be aware of its default-value semantics.

2. Compile the .proto

You only need to do this once per message type — commit the generated *_pb2.py to your repo.

Install the compiler

Check the protoc downloads page for the current release and the Python compatibility matrix. At the time of writing, protoc 21.12 pairs with protobuf 3.20.3.

On macOS:

PB_REL="https://github.com/protocolbuffers/protobuf/releases"
curl -LO $PB_REL/download/v21.12/protoc-21.12-osx-universal_binary.zip
unzip protoc-21.12-osx-universal_binary.zip -d protoc-21.12
sudo mv protoc-21.12/bin/protoc /usr/local/bin/
sudo cp -r protoc-21.12/include/* /usr/local/include/
protoc --version

Run the compiler

From the directory containing my_table.proto:

protoc -I. -I/usr/local/include --python_out=. --pyi_out=. my_table.proto

This generates my_table_pb2.py (and a .pyi stub). Commit both.

3. Write rows to BigQuery

from arcane.bigquery_storage.client import Client
from my_package.proto import my_table_pb2  # the file you just compiled

client = Client()  # picks up Application Default Credentials

# Get a stream to write to.
# Default stream = at-least-once, no per-hour limit. For exactly-once,
# use client.create_application_stream(...) instead.
stream_name = client.create_default_stream_name(
    project_id="my-project",
    dataset_id="my_dataset",
    table_id="my_table",
)

# Build a descriptor of your message type.
descriptor = Client.create_proto_descriptor(my_table_pb2.MyRow)

# Serialize each row as bytes.
rows = [
    Client.create_row(my_table_pb2.MyRow, {"name": "alice", "value": 1}),
    Client.create_row(my_table_pb2.MyRow, {"name": "bob",   "value": 2}),
]

# Append.
client.write_rows(stream_name, descriptor, rows)

Keys in the dict passed to create_row must match the field names in your .proto exactly; unknown keys raise ValueError from the protobuf runtime.

See the stream-type docs for the trade-offs between the default stream and application-created streams.

Legacy helpers (deprecated)

Earlier versions of this package shipped table-specific helpers (create_feed_boost_result_proto_descriptor, create_feed_boost_statistic_proto_descriptor, create_feed_boost_result_row, create_feed_boost_statistic_row) and bundled the matching .proto files. They still work but now emit a DeprecationWarning and will be removed in a future major release. New code should use Client.create_proto_descriptor and Client.create_row with its own compiled protobufs.

Release history

See CHANGELOG.md.

Metadata

Release files for arcane-bigquery-storage 1.4.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for arcane-bigquery-storage 1.4.0
File Size Uploaded
arcane_bigquery_storage-1.4.0.tar.gz 6.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for arcane-bigquery-storage 1.4.0
File Interpreter ABI Platform
arcane_bigquery_storage-1.4.0-py3-none-any.whl Python 3 none any Details

Total release size: 15.9 kB

Release files / arcane_bigquery_storage-1.4.0.tar.gz

Download URL arcane_bigquery_storage-1.4.0.tar.gz
Size 6.2 kB
Tags Source
SHA-256 checksum
How to use checksums
6ffabccf55d253ffa01faf7882f2ded19f76676cfa8a5941cb0ac30b0cae925f
BLAKE2b-256 checksum
How to use checksums
6eb037486ccf02cdb5d646456456f2a24fcc64cdb7ad253658fc93bd00c1c8ca
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/2.4.3 CPython/3.12.14 Linux/6.17.0-1022-azure

Release files / arcane_bigquery_storage-1.4.0-py3-none-any.whl

Download URL arcane_bigquery_storage-1.4.0-py3-none-any.whl
Size 9.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
bd5421bca49dc3a969738b11aee0893eace622310857db8c3be2f8b270d3f1f9
BLAKE2b-256 checksum
How to use checksums
6fbb461a5b4b8d1237cdd0f6e0ecdaac1e6232f081aa3b17e711f7afdae79c35
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/2.4.3 CPython/3.12.14 Linux/6.17.0-1022-azure

Release history Release notifications | RSS feed

This release

1.4.0 This release

2 release files

1.3.2

2 release files

1.3.1

2 release files

1.3.0

2 release files

1.2.2

2 release files

1.2.1

2 release files

1.2.0

2 release files

1.1.3

2 release files

1.1.2

2 release files

1.1.1

2 release files

1.1.0

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page