Skip to content

Antalya 26.6: Added support for iceberg v3 unknown data type - #2363

Open
subkanthi wants to merge 15 commits into
antalya-26.6from
iceberg_unknown_data_type
Open

subkanthi wants to merge 15 commits into
antalya-26.6from
iceberg_unknown_data_type

Conversation

@subkanthi

@subkanthi subkanthi commented Sep 15, 2026 •

Copy link
Copy Markdown
Collaborator

Changelog category (leave one):

  • New Feature

Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md):

Support for iceberg v3 unknown datatype which is maps to Nullable(Nothing) in Clickhouse. Read and write path(Parquet).

CI/CD Options

Exclude tests:

  • Fast test
  • Integration Tests
  • Stateless tests
  • Stateful tests
  • Unit tests
  • Performance tests
  • Aarch64 tests
  • All with ASAN
  • All with TSAN
  • All with MSAN
  • All with UBSAN
  • All with Coverage
  • All Regression
  • Disable CI Cache

Regression jobs to run:

  • Fast suites (mostly <1h)
  • Aggregate Functions (2h)
  • Alter (1.5h)
  • Benchmark (30m)
  • CAS (content-addressed storage; Antalya only)
  • ClickHouse Keeper (1h)
  • Iceberg (2h)
  • LDAP (1h)
  • OAuth (5m)
  • Parquet (1.5h)
  • RBAC (1.5h)
  • SSL Server (1h)
  • S3 (2h)
  • S3 Export (2h)
  • Swarms (30m)
  • Tiered Storage (2h)

@github-actions

github-actions Bot commented Sep 15, 2026 •

Copy link
Copy Markdown

Workflow [PR], commit [ae5e168]

@subkanthi subkanthi mentioned this pull request Sep 16, 2026
8 of 15 tasks
@subkanthi
subkanthi marked this pull request as ready for review September 16, 2026 17:04
@subkanthi subkanthi changed the title Added support for iceberg v3 unknown data type Antalya 26.6: Added support for iceberg v3 unknown data type Sep 16, 2026
@subkanthi

subkanthi commented Sep 19, 2026 •

Copy link
Copy Markdown
Collaborator Author

Create table in ice

Altinity/ice#219

create-table ns1.t1 --format-version=3   --schema '[{"name":"id","type":"int","required":true},{"name":"payload","type":"unknown"}]'

CH

show create table ice.`ns1.t1`;

SHOW CREATE TABLE ice.`ns1.t1`

Query id: 482b62d6-9778-4392-b504-8d5e437292a8

   ┌─statement─────────────────────────────────────────────────┐
1. │ CREATE TABLE ice.`ns1.t1`                                ↴│
   │↳(                                                        ↴│
   │↳    `id` Int32,                                          ↴│
   │↳    `payload` Nullable(Nothing)                          ↴│
   │↳)                                                        ↴│
   │↳ENGINE = Iceberg('http://localhost:9000/bucket1/ns1/t1/') │
   └───────────────────────────────────────────────────────────┘

1 row in set. Elapsed: 0.012 sec. 

WRITE PATH


Ubuntu-2404-noble-amd64-base :) insert into ice.`ns1.t1` values(1, null);

INSERT INTO ice.`ns1.t1` FORMAT Values

Query id: 4169dc98-cebc-46be-81ab-41a0a45acfbe

Ok.

1 row in set. Elapsed: 1.169 sec. 

Ubuntu-2404-noble-amd64-base :) select * from ice.`ns1.t1`;

SELECT *
FROM ice.`ns1.t1`

Query id: 9df62c95-2ca1-4749-a98b-8728e6ac135b

   ┌─id─┬─payload─┐
1. │  1 │ ᴺᵁᴸᴸ    │
   └────┴─────────┘

1 row in set. Elapsed: 0.035 sec. 

@subkanthi
subkanthi requested a review from xieandrew September 21, 2026 15:48

@xieandrew xieandrew left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

An edge case and potential optimization I found, other than that it looks good.

Comment on lines +52 to +53
auto inner_type = removeNullable(sample_block->getByPosition(i).type);
if (isNothing(inner_type))

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It looks like this doesn't check for Nothing type inside nested types (tuple, array, or map), so those nested Nothing values are not filtered for the writer. It would be good to check if that causes the parquet writer to fail.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed, added logic to recursively remove Nothing

filtered_columns.reserve(columns.size() - nothing_column_indices.size());
for (size_t i = 0; i < columns.size(); ++i)
{
if (std::find(nothing_column_indices.begin(), nothing_column_indices.end(), i) == nothing_column_indices.end())

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The std::find on every column/chunk could removed if the constructor computes a list of column indices to keep instead of nothing_column_indices. Then you would only need to iterate the kept column indices and directly add columns[i] to filtered_columns.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

done.

@subkanthi

Copy link
Copy Markdown
Collaborator Author

AI audit note: This review comment was generated by AI.

Audit update for PR #2363 (Iceberg v3 unknown type: read mapping, metadata mapping, Parquet write path)

Confirmed defects

Medium: Complex column with an unknown subfield loses all of its real data on write, silently

  • Impact: A struct or map that mixes real fields with an unknown subfield is dropped from the data file as a whole column. For example, nested = (5, NULL) for Tuple(a Nullable(Int64), u Nullable(Nothing)) reads back as (NULL, NULL). There is no error or warning.

  • Anchor: src/Storages/ObjectStorage/DataLakes/Iceberg/MultipleFileWriter.cpp, the constructor, which uses containsNothing per top-level column.

  • Trigger: A v3 table containing struct<a: long, u: unknown>, followed by an INSERT with a non-NULL value for a.

  • Why defect: The only thing that cannot be serialized is the Nothing leaf. The sibling values are valid, the user supplied them, and they are discarded.

    if (!containsNothing(sample_block->getByPosition(i).type))
        kept_column_indices.push_back(i);   // whole Tuple excluded if any leaf is Nothing

@subkanthi
subkanthi requested a review from xieandrew September 28, 2026 16:27
/// Iceberg schema metadata and is read back as NULLs on the read path.
for (size_t i = 0; i < sample_block->columns(); ++i)
{
if (!containsNothing(sample_block->getByPosition(i).type))

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It looks like this excludes the entire column from being written, even if other nested fields are not the Nothing type. Only the unknown field in the nested should be stripped otherwise there might be data loss.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yes its this AI finding #2363 (comment)

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

fixed.


# This must NOT fail with UNKNOWN_TYPE.
instance.query(
f"INSERT INTO {table_name} (id, name) VALUES (3, 'charlie'), (4, 'dave')",

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This insert should write the struct containing the unknown e.g. (42, NULL) and the next SELECT should check that the non-unknown nested field is properly written.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

instance.query(
    f"INSERT INTO {table_name} (id, name, nested) VALUES (3, 'charlie', (42, NULL)), (4, 'dave', (NULL, NULL))",
    settings={"allow_insert_into_iceberg": 1},
)

# ... then verifies:
result = instance.query(
    f"SELECT id, name, nested.a, nested.u FROM {table_name} ORDER BY id",
).strip()
expected = (
    "1\talice\t\\N\t\\N\n"
    "2\tbob\t\\N\t\\N\n"
    "3\tcharlie\t42\t\\N\n"      # <-- 42 is preserved ✓
    "4\tdave\t\\N\t\\N"
)
assert result == expected

@Selfeer

Selfeer commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator

I don't think these are mentioned here anywhere. Can you please check?

Medium: Promoting unknown to a real type is rejected, and reads of files written before a promotion do not apply it.

Iceberg v3 allows unknown to be promoted to any type. checkValidSchemaEvolution only allows int to long, float to double, and a wider decimal with the same scale, so ALTER ... MODIFY COLUMN from Nullable(Nothing) to a concrete type throws Iceberg spec doesn't allow schema evolution. On the read path, allowPrimitiveTypeConversion has the same gap: getSchemaTransformationDag leaves the old Nullable(Nothing) node in the transform when the current schema type is different. A file written while the column was unknown is therefore not cast to the promoted type.

Medium: unknown is written into tables whose format-version is still 2.

unknown exists only in Iceberg format version 3. getIcebergType maps Nothing to "unknown" with no version check, and both createEmptyMetadataFile and generateAddColumnMetadata store that type on the table's current format-version. iceberg_format_version defaults to 2, so creating a table with Nullable(Nothing), or adding that column to an existing v2 table, produces metadata other v2 readers are not required to accept.

Medium: An insert into a table that has a top-level unknown column drops min/max bounds for every column, and records a column size for a column that was not written.

MultipleFileWriter::consume updates data-file statistics from the full chunk, including the Nullable(Nothing) column, and only afterwards strips that column from the Parquet file. The column is all nulls, so getExtremes stores POSITIVE_INFINITY. canDumpIcebergStats treats that as null and canWriteStatistics then skips lower_bounds and upper_bounds for the whole file, including columns such as Int64 and String that would otherwise have bounds. column_sizes is still written for the unknown field id, using the in-memory size of a column that the writer removed from the data file.

Medium: A write that strips every column produces a Parquet file the reader rejects.

If every top-level column contains Nothing (a table whose only column is unknown, or whose only column is a struct/list/map that containsNothing drops entirely), kept_column_indices is empty and the Parquet file is written with zero columns. SchemaConverter::checkHasColumns then throws INCORRECT_DATA (Parquet file has no columns / Schema root has no children) before missing columns can be filled with nulls. The insert succeeds and the following select fails.

@subkanthi

Copy link
Copy Markdown
Collaborator Author

Write Path:

When the writer writes the unknown field to parquet its stripped, this is part of the spec where the unknown columns are not persisted in data files.
Its however stored in metadata.json as part of the table schema.

Read Path:
Since the table schema is read from metadata.json by the reader, its the readers responsibility to map the unknown data type to Nullable(Nothing) in CH.

@subkanthi

Copy link
Copy Markdown
Collaborator Author

I don't think these are mentioned here anywhere. Can you please check?

Medium: Promoting unknown to a real type is rejected, and reads of files written before a promotion do not apply it.

Iceberg v3 allows unknown to be promoted to any type. checkValidSchemaEvolution only allows int to long, float to double, and a wider decimal with the same scale, so ALTER ... MODIFY COLUMN from Nullable(Nothing) to a concrete type throws Iceberg spec doesn't allow schema evolution. On the read path, allowPrimitiveTypeConversion has the same gap: getSchemaTransformationDag leaves the old Nullable(Nothing) node in the transform when the current schema type is different. A file written while the column was unknown is therefore not cast to the promoted type.

Medium: unknown is written into tables whose format-version is still 2.

unknown exists only in Iceberg format version 3. getIcebergType maps Nothing to "unknown" with no version check, and both createEmptyMetadataFile and generateAddColumnMetadata store that type on the table's current format-version. iceberg_format_version defaults to 2, so creating a table with Nullable(Nothing), or adding that column to an existing v2 table, produces metadata other v2 readers are not required to accept.

Medium: An insert into a table that has a top-level unknown column drops min/max bounds for every column, and records a column size for a column that was not written.

MultipleFileWriter::consume updates data-file statistics from the full chunk, including the Nullable(Nothing) column, and only afterwards strips that column from the Parquet file. The column is all nulls, so getExtremes stores POSITIVE_INFINITY. canDumpIcebergStats treats that as null and canWriteStatistics then skips lower_bounds and upper_bounds for the whole file, including columns such as Int64 and String that would otherwise have bounds. column_sizes is still written for the unknown field id, using the in-memory size of a column that the writer removed from the data file.

Medium: A write that strips every column produces a Parquet file the reader rejects.

If every top-level column contains Nothing (a table whose only column is unknown, or whose only column is a struct/list/map that containsNothing drops entirely), kept_column_indices is empty and the Parquet file is written with zero columns. SchemaConverter::checkHasColumns then throws INCORRECT_DATA (Parquet file has no columns / Schema root has no children) before missing columns can be filled with nulls. The insert succeeds and the following select fails.

I found four confirmed defects in PR 2363's changes. Only one affects users: data compaction can't run on a v3 table that has an unknown column. The other three are a spec gap (promoting unknown to a real type is rejected) and two test-hygiene problems. The new write, statistics and read paths themselves came out clean. One caveat: the compaction finding comes from reading the code; I didn't execute it.


AI audit note: This review comment was generated by AI (Claude).

Audit update for PR #2363 (Iceberg v3 unknown data type: write-path stripping, statistics exclusion, v2 guard, tuple-element read fixes):

Confirmed defects:

Medium: OPTIMIZE TABLE (data compaction) fails on a v3 table that has an unknown column

  • Impact: Compaction can't run on any such table. It fails with UNKNOWN_TYPE from the Parquet writer, so small files are never merged. The failure is fail-closed (an exception, with no metadata commit).
  • Anchor: src/Storages/ObjectStorage/DataLakes/Iceberg/Compaction.cpp, the per-file rewrite loop around lines 372-395, reached from IcebergMetadata::optimize. That function passes metadata_snapshot->getSampleBlock() straight through.
  • Trigger: A v3 table with nested Tuple(a Nullable(Int64), u Nullable(Nothing)) or a top-level unknown field, allow_experimental_iceberg_compaction = 1, then OPTIMIZE TABLE.
  • Why defect: This PR makes unknown columns writable, but compaction is a second data-file writer that doesn't use Iceberg::stripNothing. It hands the unstripped header to the writer:
    auto output_format = FormatFactory::instance().getOutputFormat(
        write_format, *write_buffer, *sample_block, context, format_settings, output_format_filter_info);
    Compaction only rejects format versions below 2 (line 206), so v3 tables reach this code, and PrepareForWrite.cpp:532 throws UNKNOWN_TYPE for Nothing. The compaction statistics (data_file->manifest_list->statistics.update(chunk)) also skip the new exclusion mask.
  • Fix direction: Move the strip-and-filter logic from MultipleFileWriter into a shared helper and use it in compaction, including the statistics exclusion. Alternatively, reject unknown up front with a clear NOT_IMPLEMENTED.
  • Regression test direction: An integration test that runs OPTIMIZE on a v3 table with nested and top-level unknown across two inserts, then checks the row count, nested.a values and NULLs for nested.u.

Low: Promoting unknown to a concrete type is rejected

  • Impact: ALTER TABLE ... MODIFY COLUMN u Nullable(Int64) on a v3 table fails with BAD_ARGUMENTS "Iceberg spec doesn't allow schema evolution". The Iceberg v3 spec allows promoting unknown to any type. This is fail-closed, with no data impact.
  • Anchor: src/Storages/ObjectStorage/DataLakes/Iceberg/MetadataGenerator.cpp / checkValidSchemaEvolution, lines 105-160, called from generateModifyColumnMetadata.
  • Trigger: A v3 table with a top-level unknown field (for example created by another engine), then MODIFY COLUMN to any real type.
  • Why defect: The function has rules only for int to long, float to double and decimal widening, and none for old_type == "unknown". This is review finding 1, which was deferred earlier.
  • Fix direction: Return true when old_type is the string unknown and the new field is optional.
  • Regression test direction: A gtest in gtest_iceberg_metadata_generator.cpp covering unknown to long and to a struct, plus an integration read-back of old files as NULL.

Low: test_writes_field_ids_spark_read.py fails on every run on this branch

  • Impact: CI is red. In the earlier run, the test failed at line 105 with AssertionError: {}.
  • Anchor: tests/integration/test_storage_iceberg_with_spark/test_writes_field_ids_spark_read.py:105 against IcebergWrites.cpp / generateManifestList.
  • Trigger: Any run of this test.
  • Why defect: The test expects field-id values in the manifest-list avro.schema. The full schema is written only on the empty-list path (lines 1068-1094). Non-empty lists go through avro::DataFileWriter, which strips field-id. The writer change that makes this test pass (the port of upstream PR 111786) isn't on this branch. It exists only on backport-116171-iceberg-v3-cdc-writes.
  • Fix direction: Remove the test from this PR, or include the manifest avro.schema field-id port.
  • Regression test direction: No new test needed; the existing test becomes the regression check once the port lands.

Low: A generated test config was committed

  • Impact: The repo gets a stray file that nothing references. It's harmless now, but it's confusing.
  • Anchor: tests/integration/test_storage_iceberg_with_spark/configs/config.d/cache_62964.xml.
  • Trigger: Not applicable; it's a repo artifact.
  • Why defect: conftest.py line 59 writes configs/config.d/{filesystem_cache_name}.xml with a random name on each run, so this file is leftover output from one run.
  • Fix direction: Delete the file.
  • Regression test direction: None.

Coverage summary:

  • Scope reviewed:
    • Every changed source file in PR 2363 at da969e7256a: MultipleFileWriter, DataFileStatistics, Utils, MetadataGenerator, SchemaProcessor, Constant.h and StorageObjectStorageSource.
    • The new gtests and integration tests.
    • Two other write paths: Compaction.cpp, and the Mutations.cpp update path. Mutations are not applicable here, because v3 mutations are rejected with SUPPORT_IS_DISABLED.
  • Categories failed:
    • alternative data-file write paths (compaction);
    • schema evolution of unknown;
    • test hygiene (a test that can't pass on this branch, and a stray config file).
  • Categories passed (11):
    • stripping of Nullable, Tuple, Array and Map types, including tuples that strip to nothing;
    • column stripping with const and sparse inputs, and assert_cast safety;
    • fail-closed rejection of non-default values and of all-unknown writes, with validation running before startNewFile;
    • the statistics mask and its merge;
    • the v2 guard on CREATE, ADD and MODIFY;
    • Parquet field ids for the kept leaves;
    • the dotted-name skip in the absent-column optimization;
    • the tuple-element remap after schema evolution;
    • materializing constant parents (the earlier segfault);
    • exception safety and rollback (no metadata commit on failure);
    • concurrency (no shared mutable state, and the transform-builder lambdas run synchronously).
  • Assumptions and limits:
    • The review is static.
    • Runtime evidence covers only the reader fixes: local clickhouse local runs with the optimization set to 0 and 1, and the earlier integration log. I did not run compaction or schema promotion.
    • Pre-existing issues outside the diff are excluded:
      • last-column-id ignores nested field ids;
      • any Tuple column suppresses whole-file bounds;
      • struct MODIFY COLUMN is generally unsupported.

@subkanthi
subkanthi requested a review from xieandrew October 2, 2026 16:58
xieandrew
xieandrew previously approved these changes Oct 5, 2026

@xieandrew xieandrew left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good to me

@subkanthi

Copy link
Copy Markdown
Collaborator Author

@blau-ai

@blau-ai

blau-ai commented Oct 5, 2026

Copy link
Copy Markdown
Collaborator

CI triage for iceberg_unknown_data_type @ fac315c

Verdict: 2 PR-caused failures, the rest (Grype + all RegressionTests/testflows) are not related to this PR.

Failure Classification
Unit tests (asan_ubsan) PR-caused — your own new gtest
Integration tests (amd_asan_ubsan, db disk, old analyzer, 4/8) PR-caused — your own new test
GrypeScanKeeper / GrypeScanServer Not PR-related (image CVE scan)
RegressionTests… aggregate_functions / cas_* / settings / s3_azure / cas_selects Not PR-related (pre-existing/flaky testflows)

1. Unit tests (asan_ubsan) — PR-caused, clear fix

[ RUN      ] IcebergMetadataGenerator.AddColumnUnknownTypeRecordsUnknownInSchema
C++ exception with description "Column 'placeholder' of type Nullable(Nothing) maps to the
Iceberg `unknown` type, which requires format version 3, but the table uses format version 2
(set `iceberg_format_version = 3`)" thrown in the test body.
[  FAILED  ] IcebergMetadataGenerator.AddColumnUnknownTypeRecordsUnknownInSchema
fail: 1, passed: 14882

This is an internal inconsistency introduced by commit 67cbdb61 ("…reject unknown on iceberg v2 tables"). That commit added checkUnknownTypeAllowed() (Utils.cpp:537), which generateAddColumnMetadata now calls (MetadataGenerator.cpp:729). But the test AddColumnUnknownTypeRecordsUnknownInSchema (gtest_iceberg_metadata_generator.cpp:517) builds its metadata with makeMetadataWithGap(), which hardcodes format_version = 2 (:60), then adds a Nullable(Nothing) column — which the new rule rejects. The test was not updated alongside the new validation.

Suggested fix (keeps the test's intent — that an unknown column records "unknown" — by putting the table on v3, where unknown is allowed):

 TEST(IcebergMetadataGenerator, AddColumnUnknownTypeRecordsUnknownInSchema)
 {
     auto metadata = makeMetadataWithGap();
+    // Adding an `unknown`-typed column now requires format version 3 (checkUnknownTypeAllowed).
+    metadata->set(f_format_version, 3);
     MetadataGenerator gen(metadata);

     gen.generateAddColumnMetadata("placeholder", makeNullable(std::make_shared<DataTypeNothing>()));

(Optionally also add a sibling test asserting that the same generateAddColumnMetadata call throws BAD_ARGUMENTS on a v2 table, to lock in the new rejection behavior.)


2. Integration tests (amd_asan_ubsan, db disk, old analyzer, 4/8) — PR-caused

test_storage_iceberg_with_spark/test_writes_field_ids_spark_read.py::test_writes_manifest_field_ids_spark_read (1/1184):

assert ml_ids.get("manifest_path") == 500, ml_ids
E   AssertionError: {}
E   assert None == 500

This test is new in this PR (added in da969e72). ml_ids came back as {}, and the earlier assert manifest_lists in _avro_metadata_schemas passed — so the snap-*.avro manifest-list was written and parsed, but its Avro writer schema carried no field-id attributes at all. That's exactly the "avro-cpp drops the field-id attributes" problem the test's own docstring describes.

Note the writer that embeds field-ids into the OCF header (IcebergWrites.cpp) is not modified by this PR, so this new test is exercising existing writer behavior for the first time. Two possibilities:

  • Most likely a real gap: the manifest-list (manifest_file) header schema isn't emitting field-ids under this build/config, so the assertion fails deterministically. If so, the manual-OCF-header path in IcebergWrites.cpp needs to cover the manifest-list schema the same way it covers the data-file/manifest schema.
  • Less likely: nondeterminism in avro-cpp attribute serialization.

Suggested next step: re-run this one test; if it reproduces, trace the manifest-list header serialization in IcebergWrites.cpp (the "write the OCF header directly … full JSON with field-ids intact" path) and confirm it's applied to snap-*.avro, not only to the manifests. I can dig into that code path and propose a concrete diff if you'd like — I just can't verify a writer fix without a build.


3. Not PR-related

  • GrypeScanKeeper / GrypeScanServer (-alpine) — container image CVE scans ("4 high/critical" in keeper, "1" in server-alpine). These flag vulnerabilities in base-image/dependency packages, independent of any ClickHouse source change; nothing in this Iceberg-only PR affects them. Handled via the usual image/dependency bumps, not here.
  • RegressionTests (testflows): aggregate_functions_1/2/3, cas_ variants, settings, s3_azure_1, cas_selects* — these suites exercise aggregate-function semantics, settings, and S3/Azure storage, none of which this PR touches (the diff is confined to src/Storages/ObjectStorage/DataLakes/Iceberg/** plus Iceberg tests). These are the known pre-existing/flaky testflows failures on the antalya-26.6 line (e.g. settings reports 2 failed scenarios out of 1715). I classified these from scope rather than a base-branch diff — worth a re-run / compare against a base-branch run to confirm, but they are not caused by this change.

Bottom line: fix the two PR-owned tests. #1 is a one-line change (above). #2 needs a quick look at the manifest-list header field-id emission. Want me to open a blau/* PR with the unit-test fix (and investigate #2), or commit the unit-test fix directly onto iceberg_unknown_data_type?

🤖 automated CI triage — I can't build/run ClickHouse in this container, so correctness of any fix is validated by CI on re-run.

subkanthi and others added 2 commits October 5, 2026 12:41
The bundled avro-cpp JSON compiler drops the Iceberg field-id/element-id
attributes, so the schema it serialized into the avro.schema header of a
ClickHouse-written manifest / manifest-list omitted them. External readers
(PyIceberg, Spark) reject such a schema during scan planning ("Cannot convert
field, missing field-id"), while ClickHouse itself reads via the Iceberg
schema metadata key and did not notice.

generateManifestFile and generateManifestList now write the original
id-carrying JSON schema string as the avro.schema header. Encoded data is
unchanged; only the header schema now carries the spec field-ids it should.

Also derive the manifest partition-struct field-id from the persisted
partition spec instead of the hardcoded 1000+i: ClickHouse numbers partition
fields from 1001, and Iceberg projects partition values by field-id, so the
previously-invisible mismatch would break external readers once the id is
emitted. Legacy v1 specs that do not track partition field-ids fall back to
the sequential 1000+i default.

Signed-off-by: Kanthi Subramanian <subkanthi@gmail.com>
@subkanthi
subkanthi force-pushed the iceberg_unknown_data_type branch from a69f948 to ae5e168 Compare October 5, 2026 17:32

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants