sqlalchemy-bigquery

SQLAlchemy dialect for BigQuery

These details have not been verified by PyPI

Project links

Homepage

Project description

SQLALchemy Dialects

Quick Start

In order to use this library, you first need to go through the following steps:

Installation

Install this library in a virtualenv using pip. virtualenv is a tool to create isolated Python environments. The basic problem it addresses is one of dependencies and versions, and indirectly permissions.

With virtualenv, it’s possible to install this library without needing system install permissions, and without clashing with the installed system dependencies.

Supported Python Versions

Python >= 3.9, <3.14

Unsupported Python Versions

Python <= 3.7.

Mac/Linux

pip install virtualenv
virtualenv <your-env>
source <your-env>/bin/activate
<your-env>/bin/pip install sqlalchemy-bigquery

Windows

pip install virtualenv
virtualenv <your-env>
<your-env>\Scripts\activate
<your-env>\Scripts\pip.exe install sqlalchemy-bigquery

Installations when processing large datasets

When handling large datasets, you may see speed increases by also installing the bqstorage dependencies. See the instructions above about creating a virtual environment and then install sqlalchemy-bigquery using the bqstorage extras:

source <your-env>/bin/activate
<your-env>/bin/pip install sqlalchemy-bigquery[bqstorage]

Usage

SQLAlchemy

from sqlalchemy import *
from sqlalchemy.engine import create_engine
from sqlalchemy.schema import *
engine = create_engine('bigquery://project')
table = Table('dataset.table', MetaData(bind=engine), autoload=True)
print(select([func.count('*')], from_obj=table().scalar()))

Project

project in bigquery://project is used to instantiate BigQuery client with the specific project ID. To infer project from the environment, use bigquery:// – without project

Authentication

Follow the Google Cloud library guide for authentication.

Alternatively, you can choose either of the following approaches:

provide the path to a service account JSON file in create_engine() using the credentials_path parameter:

# provide the path to a service account JSON file
engine = create_engine('bigquery://', credentials_path='/path/to/keyfile.json')

pass the credentials in create_engine() as a Python dictionary using the credentials_info parameter:

# provide credentials as a Python dictionary
credentials_info = {
    "type": "service_account",
    "project_id": "your-service-account-project-id"
}
engine = create_engine('bigquery://', credentials_info=credentials_info)

Location

To specify location of your datasets pass location to create_engine():

engine = create_engine('bigquery://project', location="asia-northeast1")

Table names

To query tables from non-default projects or datasets, use the following format for the SQLAlchemy schema name: [project.]dataset, e.g.:

# If neither dataset nor project are the default
sample_table_1 = Table('natality', schema='bigquery-public-data.samples')
# If just dataset is not the default
sample_table_2 = Table('natality', schema='bigquery-public-data')

Batch size

By default, arraysize is set to 5000. arraysize is used to set the batch size for fetching results. To change it, pass arraysize to create_engine():

engine = create_engine('bigquery://project', arraysize=1000)

Page size for dataset.list_tables

By default, list_tables_page_size is set to 1000. list_tables_page_size is used to set the max_results for dataset.list_tables operation. To change it, pass list_tables_page_size to create_engine():

engine = create_engine('bigquery://project', list_tables_page_size=100)

Adding a Default Dataset

If you want to have the Client use a default dataset, specify it as the “database” portion of the connection string.

engine = create_engine('bigquery://project/dataset')

When using a default dataset, don’t include the dataset name in the table name, e.g.:

table = Table('table_name')

Note that specifying a default dataset doesn’t restrict execution of queries to that particular dataset when using raw queries, e.g.:

# Set default dataset to dataset_a
engine = create_engine('bigquery://project/dataset_a')

# This will still execute and return rows from dataset_b
engine.execute('SELECT * FROM dataset_b.table').fetchall()

Connection String Parameters

There are many situations where you can’t call create_engine directly, such as when using tools like Flask SQLAlchemy. For situations like these, or for situations where you want the Client to have a default_query_job_config, you can pass many arguments in the query of the connection string.

The credentials_path, credentials_info, credentials_base64, location, arraysize and list_tables_page_size parameters are used by this library, and the rest are used to create a QueryJobConfig

Note that if you want to use query strings, it will be more reliable if you use three slashes, so 'bigquery:///?a=b' will work reliably, but 'bigquery://?a=b' might be interpreted as having a “database” of ?a=b, depending on the system being used to parse the connection string.

Here are examples of all the supported arguments. Any not present are either for legacy sql (which isn’t supported by this library), or are too complex and are not implemented.

engine = create_engine(
    'bigquery://some-project/some-dataset' '?'
    'credentials_path=/some/path/to.json' '&'
    'location=some-location' '&'
    'arraysize=1000' '&'
    'list_tables_page_size=100' '&'
    'clustering_fields=a,b,c' '&'
    'create_disposition=CREATE_IF_NEEDED' '&'
    'destination=different-project.different-dataset.table' '&'
    'destination_encryption_configuration=some-configuration' '&'
    'dry_run=true' '&'
    'labels=a:b,c:d' '&'
    'maximum_bytes_billed=1000' '&'
    'priority=INTERACTIVE' '&'
    'schema_update_options=ALLOW_FIELD_ADDITION,ALLOW_FIELD_RELAXATION' '&'
    'use_query_cache=true' '&'
    'write_disposition=WRITE_APPEND'
)

In cases where you wish to include the full credentials in the connection URI you can base64 the credentials JSON file and supply the encoded string to the credentials_base64 parameter.

engine = create_engine(
    'bigquery://some-project/some-dataset' '?'
    'credentials_base64=eyJrZXkiOiJ2YWx1ZSJ9Cg==' '&'
    'location=some-location' '&'
    'arraysize=1000' '&'
    'list_tables_page_size=100' '&'
    'clustering_fields=a,b,c' '&'
    'create_disposition=CREATE_IF_NEEDED' '&'
    'destination=different-project.different-dataset.table' '&'
    'destination_encryption_configuration=some-configuration' '&'
    'dry_run=true' '&'
    'labels=a:b,c:d' '&'
    'maximum_bytes_billed=1000' '&'
    'priority=INTERACTIVE' '&'
    'schema_update_options=ALLOW_FIELD_ADDITION,ALLOW_FIELD_RELAXATION' '&'
    'use_query_cache=true' '&'
    'write_disposition=WRITE_APPEND'
)

To create the base64 encoded string you can use the command line tool base64, or openssl base64, or python -m base64.

Alternatively, you can use an online generator like www.base64encode.org <https://www.base64encode.org>_ to paste your credentials JSON file to be encoded.

Supplying Your Own BigQuery Client

The above connection string parameters allow you to influence how the BigQuery client used to execute your queries will be instantiated. If you need additional control, you can supply a BigQuery client of your own:

from google.cloud import bigquery

custom_bq_client = bigquery.Client(...)

engine = create_engine(
    'bigquery://some-project/some-dataset?user_supplied_client=True',
        connect_args={'client': custom_bq_client},
)

Creating tables

To add metadata to a table:

table = Table('mytable', ...,
    bigquery_description='my table description',
    bigquery_friendly_name='my table friendly name',
    bigquery_default_rounding_mode="ROUND_HALF_EVEN",
    bigquery_expiration_timestamp=datetime.datetime.fromisoformat("2038-01-01T00:00:00+00:00"),
)

To add metadata to a column:

Column('mycolumn', doc='my column description')

To create a clustered table:

table = Table('mytable', ..., bigquery_clustering_fields=["a", "b", "c"])

To create a time-unit column-partitioned table:

from google.cloud import bigquery

table = Table('mytable', ...,
    bigquery_time_partitioning=bigquery.TimePartitioning(
        field="mytimestamp",
        type_="MONTH",
        expiration_ms=1000 * 60 * 60 * 24 * 30 * 6, # 6 months
    ),
    bigquery_require_partition_filter=True,
)

To create an ingestion-time partitioned table:

from google.cloud import bigquery

table = Table('mytable', ...,
    bigquery_time_partitioning=bigquery.TimePartitioning(),
    bigquery_require_partition_filter=True,
)

To create an integer-range partitioned table

from google.cloud import bigquery

table = Table('mytable', ...,
    bigquery_range_partitioning=bigquery.RangePartitioning(
        field="zipcode",
        range_=bigquery.PartitionRange(start=0, end=100000, interval=10),
    ),
    bigquery_require_partition_filter=True,
)

Threading and Multiprocessing

Because this client uses the grpc library, it’s safe to share instances across threads.

In multiprocessing scenarios, the best practice is to create client instances after the invocation of os.fork by multiprocessing.pool.Pool or multiprocessing.Process.

Project details

These details have not been verified by PyPI

Project links

Homepage

Release history Release notifications | RSS feed

This version

1.16.0

Nov 6, 2025

1.15.0

Jun 23, 2025

1.14.1

May 12, 2025

1.14.0

Apr 28, 2025

1.13.0

Mar 11, 2025

1.12.1

Jan 21, 2025

1.12.0

Oct 2, 2024

1.11.0

Apr 18, 2024

1.11.0.dev2 pre-release

Feb 1, 2024

1.10.0

Feb 28, 2024

1.9.0

Dec 11, 2023

1.8.0

Aug 15, 2023

1.7.0

Jul 11, 2023

1.6.1

Feb 1, 2023

1.6.0

Jan 31, 2023

1.5.0

Nov 29, 2022

1.4.4

Jun 9, 2022

1.4.3

Mar 22, 2022

1.4.2

Mar 22, 2022

1.4.1

Mar 7, 2022

1.4.0

Feb 22, 2022

1.3.0

Jan 5, 2022

1.2.2

Nov 17, 2021

1.2.1

Oct 27, 2021

1.2.0

Sep 9, 2021

1.1.0

Aug 26, 2021

1.0.0

Aug 17, 2021

1.0.0a1 pre-release

Aug 12, 2021

0.0.7

Nov 2, 2015

0.0.6

Sep 2, 2015

0.0.5

Apr 2, 2015

0.0.4

Apr 2, 2015

0.0.3

Apr 2, 2015

0.0.2dev pre-release

Mar 30, 2015

0.0.1adev pre-release

Mar 30, 2015

0.0.1dev pre-release

Mar 30, 2015

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

sqlalchemy_bigquery-1.16.0.tar.gz (119.6 kB view details)

Uploaded Nov 6, 2025 Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

sqlalchemy_bigquery-1.16.0-py3-none-any.whl (40.6 kB view details)

Uploaded Nov 6, 2025 Python 3

File details

Details for the file sqlalchemy_bigquery-1.16.0.tar.gz.

File metadata

Download URL: sqlalchemy_bigquery-1.16.0.tar.gz
Upload date: Nov 6, 2025
Size: 119.6 kB
Tags: Source
Uploaded using Trusted Publishing? Yes
Uploaded via: twine/6.1.0 CPython/3.11.2

File hashes

Hashes for sqlalchemy_bigquery-1.16.0.tar.gz
Algorithm	Hash digest
SHA256	`fe937a0d1f4cf7219fcf5d4995c6718805b38d4df43e29398dec5dc7b6d1987e`
MD5	`b01da14e1e7c4bdf8f2dd5898c9bbe09`
BLAKE2b-256	`7e6ac49932b3d9c44cab9202b1866c5b36b7f0d0455d4653fbc0af4466aeaa76`

See more details on using hashes here.

File details

Details for the file sqlalchemy_bigquery-1.16.0-py3-none-any.whl.

File metadata

Download URL: sqlalchemy_bigquery-1.16.0-py3-none-any.whl
Upload date: Nov 6, 2025
Size: 40.6 kB
Tags: Python 3
Uploaded using Trusted Publishing? Yes
Uploaded via: twine/6.1.0 CPython/3.11.2

File hashes

Hashes for sqlalchemy_bigquery-1.16.0-py3-none-any.whl
Algorithm	Hash digest
SHA256	`0fe7634cd954f3e74f5e2db6d159f9e5ee87a47fbe8d52eac3cd3bb3dadb3a77`
MD5	`28ddc7affbaef8936464e919a2d12983`
BLAKE2b-256	`c08711e6de00ef7949bb8ea06b55304a1a4911c329fdf0d9882b464db240c2c5`

See more details on using hashes here.

sqlalchemy-bigquery 1.16.0

Navigation

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Project description

Quick Start

Installation

Supported Python Versions

Unsupported Python Versions

Mac/Linux

Windows

Installations when processing large datasets

Usage

SQLAlchemy

Project

Authentication

Location

Table names

Batch size

Page size for dataset.list_tables

Adding a Default Dataset

Connection String Parameters

Supplying Your Own BigQuery Client

Creating tables

Threading and Multiprocessing

Project details

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

File details

File metadata

File hashes