Skip to content

Antalya 26.6: Added support for iceberg v3 unknown data type - #2363

Open
subkanthi wants to merge 6 commits into
antalya-26.6from
iceberg_unknown_data_type
Open

subkanthi wants to merge 6 commits into
antalya-26.6from
iceberg_unknown_data_type

Conversation

@subkanthi

@subkanthi subkanthi commented Sep 15, 2026 •

Copy link
Copy Markdown
Collaborator

Changelog category (leave one):

  • New Feature

Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md):

Support for iceberg v3 unknown datatype which is maps to Nullable(Nothing) in Clickhouse. Read and write path(Parquet).

CI/CD Options

Exclude tests:

  • Fast test
  • Integration Tests
  • Stateless tests
  • Stateful tests
  • Unit tests
  • Performance tests
  • Aarch64 tests
  • All with ASAN
  • All with TSAN
  • All with MSAN
  • All with UBSAN
  • All with Coverage
  • All Regression
  • Disable CI Cache

Regression jobs to run:

  • Fast suites (mostly <1h)
  • Aggregate Functions (2h)
  • Alter (1.5h)
  • Benchmark (30m)
  • CAS (content-addressed storage; Antalya only)
  • ClickHouse Keeper (1h)
  • Iceberg (2h)
  • LDAP (1h)
  • OAuth (5m)
  • Parquet (1.5h)
  • RBAC (1.5h)
  • SSL Server (1h)
  • S3 (2h)
  • S3 Export (2h)
  • Swarms (30m)
  • Tiered Storage (2h)

@github-actions

github-actions Bot commented Sep 15, 2026 •

Copy link
Copy Markdown

Workflow [PR], commit [e77daf6]

@subkanthi subkanthi mentioned this pull request Sep 16, 2026
15 tasks
@subkanthi
subkanthi marked this pull request as ready for review September 16, 2026 17:04
@subkanthi subkanthi changed the title Added support for iceberg v3 unknown data type Antalya 26.6: Added support for iceberg v3 unknown data type Sep 16, 2026
@subkanthi

subkanthi commented Sep 19, 2026 •

Copy link
Copy Markdown
Collaborator Author

Create table in ice

Altinity/ice#219

create-table ns1.t1 --format-version=3   --schema '[{"name":"id","type":"int","required":true},{"name":"payload","type":"unknown"}]'

CH

show create table ice.`ns1.t1`;

SHOW CREATE TABLE ice.`ns1.t1`

Query id: 482b62d6-9778-4392-b504-8d5e437292a8

   ┌─statement─────────────────────────────────────────────────┐
1. │ CREATE TABLE ice.`ns1.t1`                                ↴│
   │↳(                                                        ↴│
   │↳    `id` Int32,                                          ↴│
   │↳    `payload` Nullable(Nothing)                          ↴│
   │↳)                                                        ↴│
   │↳ENGINE = Iceberg('http://localhost:9000/bucket1/ns1/t1/') │
   └───────────────────────────────────────────────────────────┘

1 row in set. Elapsed: 0.012 sec. 

WRITE PATH


Ubuntu-2404-noble-amd64-base :) insert into ice.`ns1.t1` values(1, null);

INSERT INTO ice.`ns1.t1` FORMAT Values

Query id: 4169dc98-cebc-46be-81ab-41a0a45acfbe

Ok.

1 row in set. Elapsed: 1.169 sec. 

Ubuntu-2404-noble-amd64-base :) select * from ice.`ns1.t1`;

SELECT *
FROM ice.`ns1.t1`

Query id: 9df62c95-2ca1-4749-a98b-8728e6ac135b

   ┌─id─┬─payload─┐
1. │  1 │ ᴺᵁᴸᴸ    │
   └────┴─────────┘

1 row in set. Elapsed: 0.035 sec. 

@subkanthi
subkanthi requested a review from xieandrew September 21, 2026 15:48

@xieandrew xieandrew left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

An edge case and potential optimization I found, other than that it looks good.

Comment on lines +52 to +53
auto inner_type = removeNullable(sample_block->getByPosition(i).type);
if (isNothing(inner_type))

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It looks like this doesn't check for Nothing type inside nested types (tuple, array, or map), so those nested Nothing values are not filtered for the writer. It would be good to check if that causes the parquet writer to fail.

filtered_columns.reserve(columns.size() - nothing_column_indices.size());
for (size_t i = 0; i < columns.size(); ++i)
{
if (std::find(nothing_column_indices.begin(), nothing_column_indices.end(), i) == nothing_column_indices.end())

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The std::find on every column/chunk could removed if the constructor computes a list of column indices to keep instead of nothing_column_indices. Then you would only need to iterate the kept column indices and directly add columns[i] to filtered_columns.

@subkanthi

Copy link
Copy Markdown
Collaborator Author

AI audit note: This review comment was generated by AI.

Audit update for PR #2363 (Iceberg v3 unknown type: read mapping, metadata mapping, Parquet write path)

Confirmed defects

Medium: Complex column with an unknown subfield loses all of its real data on write, silently

  • Impact: A struct or map that mixes real fields with an unknown subfield is dropped from the data file as a whole column. For example, nested = (5, NULL) for Tuple(a Nullable(Int64), u Nullable(Nothing)) reads back as (NULL, NULL). There is no error or warning.

  • Anchor: src/Storages/ObjectStorage/DataLakes/Iceberg/MultipleFileWriter.cpp, the constructor, which uses containsNothing per top-level column.

  • Trigger: A v3 table containing struct<a: long, u: unknown>, followed by an INSERT with a non-NULL value for a.

  • Why defect: The only thing that cannot be serialized is the Nothing leaf. The sibling values are valid, the user supplied them, and they are discarded.

    if (!containsNothing(sample_block->getByPosition(i).type))
        kept_column_indices.push_back(i);   // whole Tuple excluded if any leaf is Nothing

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants