Skip to content

samples(bigquery-storage): add Arrow query results samples for query_and_wait and read_rows - #18126

Open
alextolpin wants to merge 2 commits into
googleapis:mainfrom
alextolpin:arrow_samples
Open

samples(bigquery-storage): add Arrow query results samples for query_and_wait and read_rows#18126
alextolpin wants to merge 2 commits into
googleapis:mainfrom
alextolpin:arrow_samples

Conversation

@alextolpin

Copy link
Copy Markdown
Contributor

Description

Adds documentation code snippets and system tests demonstrating high-performance query result retrieval in Apache Arrow format with LZ4 compression using the BigQuery Storage API:

  1. query_and_wait() with Arrow format & LZ4 frame compression (query_and_wait_arrow.py):

    • Demonstrates calling client.query_and_wait() with query_results_format=enums.QueryResultsFormat.ARROW and compression_codec=enums.QueryResultsCompressionCodec.LZ4_FRAME.
    • Returns results as an iterable of pyarrow.RecordBatch via results.to_arrow_iterable().
    • Wrapped in region tag: [START bigquerystorage_query_and_wait_arrow].
  2. Direct read_rows on query job default stream (read_rows_query_job.py):

    • Demonstrates directly reading query results using BigQueryReadClient.read_rows against the job stream projects/{project}/locations/{location}/jobs/{job_id}/streams/_default.
    • Deserializes schema and record batches safely via pyarrow.ipc.
    • Wrapped in region tag: [START bigquerystorage_read_rows_query_job].
  3. Tests & Dependencies:

    • Added query_and_wait_arrow_test.py and read_rows_query_job_test.py verifying batch iteration and schema types.
    • Added pyarrow dependency pins to samples/snippets/requirements.txt.

Follow-up to #18027
Related to #18047

Checklist

  • Ensure the tests and linter pass
  • Code coverage does not decrease (if any source code was changed)
  • Appropriate docs were updated (if necessary)

@alextolpin
alextolpin requested review from a team as code owners August 16, 2026 17:59
@alextolpin
alextolpin requested review from sindhuvy and removed request for a team August 16, 2026 17:59
@snippet-bot

snippet-bot Bot commented Aug 16, 2026

Copy link
Copy Markdown

Here is the summary of changes.

You are about to add 2 region tags.

This comment is generated by snippet-bot.
If you find problems with this result, please file an issue at:
https://github.com/googleapis/repo-automation-bots/issues.
To update this comment, add snippet-bot:force-run label or use the checkbox below:

  • Refresh this comment

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces two new code snippets and their corresponding tests to demonstrate reading BigQuery query results in Apache Arrow format. Specifically, query_and_wait_arrow.py uses the query_and_wait method to fetch Arrow RecordBatches directly, while read_rows_query_job.py initiates a query job and streams the results using the BigQuery Storage Read API. A critical issue was identified in read_rows_query_job.py where the query job is started asynchronously, but the code immediately attempts to read from the stream without waiting for the job to complete. It is recommended to call job.result() to ensure the query finishes before reading.

Comment on lines +45 to +46
# Start the query job.
job = client.query(query)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The client.query(query) method starts an asynchronous query job and returns immediately. Because the job runs asynchronously, attempting to construct the stream name and read from it immediately will fail since the job is still pending or running, and its results are not yet available (additionally, job.location may not be populated yet).

To ensure the query has finished and the results are ready to be read from the stream, you must wait for the job to complete by calling job.result().

Suggested change
# Start the query job.
job = client.query(query)
# Start the query job and wait for it to complete.
job = client.query(query)
job.result()

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant