Count table rows without materializing every field during inspect - #371
Open
OskarEichler wants to merge 1 commit into
Open
Count table rows without materializing every field during inspect#371OskarEichler wants to merge 1 commit into
OskarEichler wants to merge 1 commit into
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Count non-header rows directly instead of building the full
to_arepresentation merely to read its size. Keep the existing five-row preview and header-inclusive count.Reproduction
Inspecting a 10,000-row, 20-column table currently creates arrays for every row even though only five rows are displayed.
Verification
24 exact-output checks pass across table sizes, header-row presence and all access modes. In seven local measurements of 20 inspections (10,000 rows x 20 columns), median allocations fell from 405,280 to 5,240; median elapsed time fell from 0.192385s to 0.006012s. These are local microbenchmark results, not production throughput claims.
This isolated patch passes the unchanged upstream suite in both normal and
CSV_PARSER_SCANNER_TEST=yesmodes: 527 tests / 4,050 assertions in each mode, zero failures/errors. Commands use Ruby 4.0.6 through rbenv (bundle exec ruby run-test.rb). Syntax andgit diff --checkpass. The combined release-based consumer branch also passes gem build andrake warning:error rdoc.No tests were added or changed, per the requesting project's explicit no-new-tests policy. Focused checks and measurements live outside this repository. The upstream no-RuboCop policy is respected. Verification was local on macOS/Ruby 4.0.6; the supported OS/Ruby matrix was not run locally. Current main and installed 3.3.6 share base commit
0873ab362d996f12796c3c3e8998b5be657a9b12; no unrelated changes are included.Breaking changes and limitations
None. Inspection text and encoding behavior remain unchanged. Counting still traverses rows; this removes field materialization, not the row scan.