Movatterモバイル変換


[0]ホーム

URL:


Skip to content

Navigation Menu

Sign in
Appearance settings

Search code, repositories, users, issues, pull requests...

Provide feedback

We read every piece of feedback, and take your input very seriously.

Saved searches

Use saved searches to filter your results more quickly

Sign up
Appearance settings

[Rows] sub-rowgroup loading using libviewer#3213

New issue

Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.

By clicking “Sign up for GitHub”, you agree to ourterms of service andprivacy statement. We’ll occasionally send you account related emails.

Already on GitHub?Sign in to your account

Draft
lhoestq wants to merge3 commits intomain
base:main
Choose a base branch
Loading
fromlibviewer-in-rows

Conversation

lhoestq
Copy link
Member

@lhoestqlhoestq commentedJul 7, 2025
edited
Loading

continuation of#3199

Pretty important PR since sub-rowgroup loading lets us:

  • load Pandas datasets that are >300MB, or from parquet files file PyArrow/Spark if row groups are >300MB
  • increase the row group size indatasets for better Xet deduplication

TODO:

  • create page index when required (row groups > 300MB)
    • it's a richer parquet metadata file than the one from from the "config-parquet-metadata" job and it replaces the existing metadata file
  • load rows using page index when required

It works using page pruning fromarrow-rs

Sign up for freeto join this conversation on GitHub. Already have an account?Sign in to comment
Reviewers
No reviews
Assignees
No one assigned
Labels
None yet
Projects
None yet
Milestone
No milestone
Development

Successfully merging this pull request may close these issues.

2 participants
@lhoestq@kszucs

[8]ページ先頭

©2009-2025 Movatter.jp