Problem
parquet-java has no public API to write Arrow VectorSchemaRoot to Parquet files. Callers must construct row objects and feed them through ParquetWriter<T>.write(T) one at a time. Arrow C++ and PyArrow support this natively via parquet::WriteTable() / pq.write_table().
Related prior discussion:
Motivation
Several downstream Java projects work with Arrow-columnar data internally and must materialize row objects solely to satisfy ParquetWriter's input API:
- Apache Iceberg (
FileAppender<Record>) — iceberg#17748
- Apache Fluss (Arrow-native streaming storage) — fluss#4047
- Apache Paimon (worked around this by building their own
paimon-arrow writer that bypasses parquet-java's row API entirely)
Key Insight
For PLAIN-encoded, fixed-width columns, Arrow's in-memory format (contiguous little-endian values) is identical to Parquet's PLAIN page encoding. A zero-copy path is possible by wrapping the Arrow data buffer directly as a BytesInput and passing it to PageWriter.writePage().
For nullable columns, Arrow's validity bitmap can be scanned for contiguous non-null runs, with each run bulk-copied and definition levels encoded as RLE runs of the same value — O(null_transitions) instead of O(N).
Proposed Design
A new ArrowParquetWriter in the parquet-arrow module that writes pages directly to PageWriter rather than going through RecordConsumer/ColumnWriter:
ArrowParquetWriter.writeBatch(VectorSchemaRoot)
→ per column: selects optimal write strategy
→ writes pages to PageWriter via getPageWriter(ColumnDescriptor)
→ manages row groups via ParquetFileWriter
Per-column strategy selection (best to worst):
- Zero-copy (non-null + fixed-width + PLAIN): wrap Arrow buffer as
BytesInput directly. RL/DL as single-value RLE runs. Stats from sequential buffer scan.
- Bulk-copy with nulls (nullable + fixed-width + PLAIN): scan validity bitmap for non-null runs, bulk-copy each, emit DL as RLE runs per null-transition.
- Variable-width rewrite (PLAIN + string/binary): single-pass transformation of Arrow offset+data buffers to Parquet's length-prefixed format.
- Dictionary: map Arrow's
DictionaryEncodedVector to Parquet dictionary pages, or build dictionary from plain vector.
- Fallback: per-value through
ValuesWriter (same cost as today, for unsupported encoding/type combos).
All verified public APIs exist for this:
PageWriteStore.getPageWriter(ColumnDescriptor) — get column page writer directly
PageWriter.writePage(BytesInput, valueCount, rowCount, stats, encodings) — write pre-encoded pages
BytesInput.from(ByteBuffer) — zero-copy buffer wrapping
RunLengthBitPackingHybridEncoder — encode RL/DL levels
ColumnChunkPageWriteStore.flushToFileWriter() — flush pages to file
Why Not Extend ParquetWriter?
ParquetWriter.write(T) calls InternalParquetRecordWriter.write() which increments recordCount by 1 per call and checks row-group boundaries based on that count. This is incompatible with batch semantics. The writer must manage ParquetFileWriter and row groups directly.
Scope
- Parquet file format unchanged
- Existing
ParquetWriter<T> API unchanged
parquet-arrow gains parquet-hadoop as a compile dependency (for ParquetFileWriter, ColumnChunkPageWriteStore)
- No Hadoop runtime dependency for uncompressed output (compression requires caller-supplied
CompressionCodecFactory)
- Flat schemas (primitive columns) first; nested types deferred
Implementation Phases
- Core infrastructure + zero-copy for non-null fixed-width PLAIN columns
- Nullable column support (validity bitmap scanning + run-based bulk copy)
- Variable-width types (string/binary offset rewriting)
- Dictionary encoding support
- Bloom filter, nested types, advanced features
Problem
parquet-javahas no public API to write ArrowVectorSchemaRootto Parquet files. Callers must construct row objects and feed them throughParquetWriter<T>.write(T)one at a time. Arrow C++ and PyArrow support this natively viaparquet::WriteTable()/pq.write_table().Related prior discussion:
Motivation
Several downstream Java projects work with Arrow-columnar data internally and must materialize row objects solely to satisfy
ParquetWriter's input API:FileAppender<Record>) — iceberg#17748paimon-arrowwriter that bypassesparquet-java's row API entirely)Key Insight
For PLAIN-encoded, fixed-width columns, Arrow's in-memory format (contiguous little-endian values) is identical to Parquet's PLAIN page encoding. A zero-copy path is possible by wrapping the Arrow data buffer directly as a
BytesInputand passing it toPageWriter.writePage().For nullable columns, Arrow's validity bitmap can be scanned for contiguous non-null runs, with each run bulk-copied and definition levels encoded as RLE runs of the same value — O(null_transitions) instead of O(N).
Proposed Design
A new
ArrowParquetWriterin theparquet-arrowmodule that writes pages directly toPageWriterrather than going throughRecordConsumer/ColumnWriter:Per-column strategy selection (best to worst):
BytesInputdirectly. RL/DL as single-value RLE runs. Stats from sequential buffer scan.DictionaryEncodedVectorto Parquet dictionary pages, or build dictionary from plain vector.ValuesWriter(same cost as today, for unsupported encoding/type combos).All verified public APIs exist for this:
PageWriteStore.getPageWriter(ColumnDescriptor)— get column page writer directlyPageWriter.writePage(BytesInput, valueCount, rowCount, stats, encodings)— write pre-encoded pagesBytesInput.from(ByteBuffer)— zero-copy buffer wrappingRunLengthBitPackingHybridEncoder— encode RL/DL levelsColumnChunkPageWriteStore.flushToFileWriter()— flush pages to fileWhy Not Extend ParquetWriter?
ParquetWriter.write(T)callsInternalParquetRecordWriter.write()which incrementsrecordCountby 1 per call and checks row-group boundaries based on that count. This is incompatible with batch semantics. The writer must manageParquetFileWriterand row groups directly.Scope
ParquetWriter<T>API unchangedparquet-arrowgainsparquet-hadoopas a compile dependency (forParquetFileWriter,ColumnChunkPageWriteStore)CompressionCodecFactory)Implementation Phases