Skip to content

Add column projection resolution benchmarks for the Parquet readers - #24193

Merged
rapids-bot[bot] merged 11 commits into
NVIDIA:mainfrom
qbacpey:bench/hybrid-scan-projection
Oct 2, 2026
Merged

rapids-bot[bot] merged 11 commits into
NVIDIA:mainfrom
qbacpey:bench/hybrid-scan-projection

Conversation

@qbacpey

@qbacpey qbacpey commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

Description

Split out of #24162

Adds a new axis for an existed benchmark, both of them are about column-name resolution

Checklist

  • I am familiar with the Contributing Guidelines.
  • New or existing tests cover these changes (benchmark-only change; no production code modified).
  • The documentation is up to date with these changes.

@copy-pr-bot

copy-pr-bot Bot commented Sep 16, 2026

Copy link
Copy Markdown

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@github-actions github-actions Bot added libcudf Affects libcudf (C++/CUDA) code. CMake CMake build issue labels Sep 16, 2026
@qbacpey qbacpey changed the title Bench/hybrid scan projection Add column projection resolution benchmarks for the Parquet readers Sep 16, 2026
@qbacpey qbacpey added improvement Improvement / enhancement to an existing function non-breaking Non-breaking change labels Sep 16, 2026
@qbacpey
qbacpey force-pushed the bench/hybrid-scan-projection branch 3 times, most recently from 94ee62d to eb6a7b2 Compare September 16, 2026 12:30
@qbacpey

qbacpey commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor Author

Benchmark results

Machine: NVIDIA GH200 480GB (132-SM Hopper, Neoverse-V2 CPU), driver 580.95.05, CUDA 13.3. All runs with CUDF_BENCHMARK_DROP_CACHE=1.

  • Base = main;
  • PR = this branch after merging upstream/main.

Raw nvbench output — PR branch, parquet_column_selection

| io_type  | selection_method | num_cols | Samples |  CPU Time  | Noise |  GPU Time  | Noise | cols_per_sec | peak_memory_usage |
|----------|------------------|----------|---------|------------|-------|------------|-------|--------------|-------------------|
| FILEPATH |             none |       64 |   5744x | 106.029 us | 2.67% |  87.200 us | 3.01% |       733941 |        44.561 KiB |
| FILEPATH |            names |       64 |   4304x | 136.041 us | 3.21% | 116.289 us | 3.48% |       550352 |        44.561 KiB |
| FILEPATH |             none |      512 |    880x | 673.656 us | 1.91% | 651.361 us | 1.98% |       786045 |       345.907 KiB |
| FILEPATH |            names |      512 |    384x |   1.377 ms | 1.18% |   1.354 ms | 1.19% |       378118 |       345.907 KiB |
| FILEPATH |             none |     2048 |    864x |   3.283 ms | 1.88% |   3.259 ms | 1.89% |       628432 |         1.385 MiB |
| FILEPATH |            names |     2048 |     50x |  10.090 ms | 0.39% |  10.064 ms | 0.39% |       203501 |         1.385 MiB |

@qbacpey qbacpey added the 2 - In Progress Currently a work in progress label Sep 16, 2026
@qbacpey
qbacpey force-pushed the bench/hybrid-scan-projection branch from eb6a7b2 to f0d6376 Compare September 16, 2026 15:11
rapids-bot Bot pushed a commit that referenced this pull request Sep 25, 2026
… produced (#24191)

Split out of #24162. 

This PR is the shared groundwork that #24192 and #24193 build on. It also sets `max_page_fragment_size` and raises `max_page_size_bytes` to ensure benchmark `parquet_read_file_shape` can actually generate requested Parquet shape.

Authors:
  - Qi Chen (https://github.com/qbacpey)

Approvers:
  - Vukasin Milovanovic (https://github.com/vuule)
  - Kyle Edwards (https://github.com/KyleFromNVIDIA)

URL: #24191
@qbacpey
qbacpey force-pushed the bench/hybrid-scan-projection branch from f0d6376 to a23494a Compare September 29, 2026 11:31
@copy-pr-bot

copy-pr-bot Bot commented Sep 29, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

qbacpey and others added 3 commits September 30, 2026 14:41
Moves the writer inlined in BM_parquet_filter_name_resolution into
write_named_resolution_parquet_file so the projection benchmarks can resolve
against the same deterministically named schema.
hybrid_scan_projection times the filter and payload column chunk byte range
calls across a side axis; parquet_read_column_projection is the naive reader
baseline, timing chunked reader construction with a full explicit projection.
Introduces the `named_resolution_column_names` function to generate deterministic column names for Parquet files. This function is utilized in `write_named_resolution_parquet_file` and various benchmarks to ensure consistent naming across different Parquet operations, enhancing code clarity and maintainability.
@qbacpey
qbacpey force-pushed the bench/hybrid-scan-projection branch from a23494a to bbe33c6 Compare September 30, 2026 15:28
@qbacpey
qbacpey requested a review from mhaseeb123 September 30, 2026 15:45
@qbacpey
qbacpey marked this pull request as ready for review September 30, 2026 15:45
@qbacpey
qbacpey requested review from a team as code owners September 30, 2026 15:45
@qbacpey
qbacpey requested a review from ttnghia September 30, 2026 15:45
@qbacpey qbacpey added 3 - Ready for Review Ready for review by team and removed 2 - In Progress Currently a work in progress labels Sep 30, 2026
@coderabbitai

coderabbitai Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: NVIDIA/cudf/.coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: a2016fe2-6161-4c66-b86a-9ac9e77d2d99

📥 Commits

Reviewing files that changed from the base of the PR and between df8428a and 0178f4b.

📒 Files selected for processing (2)
  • cpp/benchmarks/io/parquet/experimental/hybrid_scan/hybrid_scan_projection.cpp
  • cpp/benchmarks/io/parquet/parquet_reader_metadata.cpp

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review.


📝 Summary

Summary by CodeRabbit

  • Chores
    • Expanded Parquet benchmark coverage to include hybrid-scan payload-column selection across schemas with 64 to 4,096 columns.
    • Added comparisons of positional and name-based column selection in Parquet reader benchmarks.
    • Benchmarks report schema column counts and peak memory usage. Parquet filter-name-resolution results now report schema column counts instead of columns processed per second.

Walkthrough

Adds Parquet reader-selection and hybrid-scan projection benchmarks. The reader-selection benchmark covers positional and name-based selection. The filter-name-resolution benchmark now reports schema column count instead of throughput.

Changes

Parquet Projection Benchmarks

Layer / File(s) Summary
Reader selection and metadata benchmarks
cpp/benchmarks/io/parquet/parquet_reader_metadata.cpp
Adds positional and name-based selection cases to the column-selection benchmark. The filter-name-resolution benchmark reports schema_columns instead of cols_per_sec; peak-memory reporting remains.
Hybrid-scan projection benchmark
cpp/benchmarks/io/parquet/experimental/hybrid_scan/hybrid_scan_projection.cpp, cpp/benchmarks/CMakeLists.txt
Adds and registers a benchmark that writes a configurable-width Parquet table, prepares filter column chunks, and times payload column-chunk selection. It reports schema column count and peak memory usage.

Priority: ⬇️ Low

Estimated code review effort: 2 (Simple) | ~10 minutes

Change: Other

Suggested reviewers: vyasr, pointkernel

Merge Risk: ⚪ Minimal · up to 0178f

This change adds a registered projection benchmark and changes the filter benchmark’s reported metric. No concrete repository-internal breakage or other actionable merge risk is established.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 42.86% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 7 functions across 4 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly summarizes the main changes: it adds Parquet reader benchmarks for column projection resolution, including the new hybrid scan benchmark and the expanded existing benchmark.
Description check ✅ Passed The description is directly related to the changeset. It identifies the new benchmark axis and the focus on column-name resolution.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Comment @coderabbitai help to get the list of available commands.

Comment on lines +104 to +112
using projection_sides = nvbench::enum_type_list<projection_side::FILTER,
projection_side::PAYLOAD,
projection_side::PAYLOAD_EXPLICIT>;

NVBENCH_BENCH_TYPES(BM_hybrid_scan_projection, NVBENCH_TYPE_AXES(projection_sides))
.set_name("hybrid_scan_projection")
.set_type_axes_names({"side"})
.set_min_samples(4)
.add_int64_axis("num_cols", {64, 512, 2048, 4096});

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we really need enums here. We could just use a string axis.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Removed

{
auto const num_cols = static_cast<cudf::size_type>(state.get_int64("num_cols"));

auto source_sink = write_named_resolution_parquet_file(num_cols, io_type::FILEPATH);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

By default, the parquet writer writes column names as: _col0, _col1, .., ... Can we just use that instead of explicitly writing col0, col1... names? We can also select columns by index as well if needed. We don't need the two new helpers in parquet_common.xpp that way

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done

Comment on lines +60 to +64
// Caller-supplied payload list
if constexpr (Side == projection_side::PAYLOAD_EXPLICIT) {
read_opts_builder.column_names(named_resolution_column_names(num_cols));
}
auto const read_opts = read_opts_builder.build();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is PAYLOAD_EXPLICIT even needed? Isn't the column selection cost just proportional to the number of columns being selected? Also please check if any of the existing parquet benchmarks cover column selection. If so, then PAYLOAD_EXPLICIT as well as PAYLOAD modes are just duplicates and need to be removed

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Removed PAYLOAD_EXPLICIT.

None of the existing benchmarks cover the payload case (a filter and no projection):

  • No other benchmark calls payload_column_chunks_byte_ranges; the hybrid scan benchmarks time the filters or read in a single step through all_column_chunks_byte_ranges
  • parquet_column_selection passes no names, so it times the linear path
  • parquet_read_column_selection times a full read of a few dozen columns, where selection is negligible

Selection cost is only proportional to the number of columns when no names are passed. With names, select_columns looks up each one with a linear find_if over every schema path, so the cost is roughly selected columns × schema width. With a filter and no projection, select_payload_columns turns every non-filter column into a name and passes all n−1 of them to that lookup, while the regular reader in the same situation passes no names and selects all columns in one pass.

Same build of this branch, GH200:

num_cols hybrid payload selection regular reader, filter, no projection regular reader, all names projected
64 0.130 ms 0.079 ms 0.099 ms
512 1.45 ms 0.72 ms 1.21 ms
2048 10.58 ms 2.46 ms 9.08 ms
4096 38.52 ms 5.20 ms 32.64 ms

The columns come from hybrid_scan_projection, parquet_filter_name_resolution (case_sensitive=1, heavy_filter=0) and parquet_read_column_projection.

Should we further drop the the payload case (a filter and no projection)?

Comment on lines +51 to +55
cudf::ast::tree filter_tree;
auto const& col_ref = filter_tree.push(cudf::ast::column_name_reference("col0"));
auto const& lit = filter_tree.push(cudf::ast::literal(filter_literal));
auto const& filter_expr =
filter_tree.push(cudf::ast::operation(cudf::ast::ast_operator::GREATER_EQUAL, col_ref, lit));

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The filter AST tree seems trivial (single predicate) so the filter column selection cost would be trivial. Can we use something like here 4a59ca5#diff-c9f9907616a2121d2c0cd6dd2d50126a42964461acf05132bfe59db0b67b2a21 to build a varying depth/columns AST tree or skip this altogether as filter/payload columns ratio is usually small in real world.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed, dropped it since parquet_filter_name_resolution sweeps heavy filters for the regular reader.

qbacpey and others added 3 commits October 1, 2026 18:39
…ing metadata handling

- Removed the `named_resolution_column_names` and `write_named_resolution_parquet_file` functions from `parquet_common.cpp` and `parquet_common.hpp`.
- Updated benchmarks in `parquet_reader_metadata.cpp` and `hybrid_scan_projection.cpp` to directly handle column names and metadata without the removed functions.
- Adjusted CMake configuration to reflect the changes in benchmark source files.

This cleanup enhances code maintainability and reduces unnecessary complexity in the benchmark implementations.
@qbacpey
qbacpey requested a review from mhaseeb123 October 1, 2026 18:12
Comment on lines +71 to +73
// Select the filter columns first, as a real read does, so the timed call covers
// only the payload selection
std::ignore = reader->filter_column_chunks_byte_ranges(row_groups, read_opts);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this might not be needed anymore. The hybrid_scan_impl::select_columns() now computes the filter column names (if not already done and filter is present) before deducing payload columns. Please check and update 🙂

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Removed


state.exec(nvbench::exec_tag::sync | nvbench::exec_tag::timer,
[&](nvbench::launch& launch, auto& timer) {
drop_page_cache_if_enabled(source_info.filepaths());

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Don't think this is needed either

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Removed

mem_stats_logger.peak_memory_usage(), "peak_memory_usage", "peak_memory_usage");
}

// Benchmark full-projection column-name resolution during naive reader construction.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we just add an axis (selection method?) to the other benchmark instead where we either provide column selection or not instead of this separate benchmark?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done, parquet_column_selection now has a selection_method axis

qbacpey and others added 2 commits October 2, 2026 10:40
…mn name selection. Removed unused benchmark for full-projection column-name resolution. Updated benchmark parameters to include selection method options.
@qbacpey
qbacpey requested a review from mhaseeb123 October 2, 2026 11:35
Comment thread cpp/benchmarks/io/parquet/experimental/hybrid_scan/hybrid_scan_projection.cpp Outdated
anon and others added 2 commits October 2, 2026 19:39
…akeLists.txt. This cleanup eliminates unused code and simplifies the benchmark setup.
@qbacpey

qbacpey commented Oct 2, 2026

Copy link
Copy Markdown
Contributor Author

/ok to test 5f4e0f0

@qbacpey

qbacpey commented Oct 2, 2026

Copy link
Copy Markdown
Contributor Author

/merge

@rapids-bot
rapids-bot Bot merged commit baaaea9 into NVIDIA:main Oct 2, 2026
72 of 74 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

3 - Ready for Review Ready for review by team CMake CMake build issue improvement Improvement / enhancement to an existing function libcudf Affects libcudf (C++/CUDA) code. non-breaking Non-breaking change

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants