Describe the situation
On a two-replica cluster with CAS as the default disk, GC rounds slowed from 4 minutes to 48 minutes over 8 days. Deletes per round are pinned at exactly 5000. Retired-but-undeleted blobs reached ~520k. No data damage: cas_log correctness invariants are clean, S3 upload error rate is ~0.03%, no lease loss.
Environment
- Version:
26.6.4.20001.altinityantalya
- 2 replicas, CAS disk
cas with a cache disk cas_cache on top, used as the default disk
- 94 live namespaces, ~400-475k ref-log (
_log) keys and ~320k blob retirements per day
- GC settings: defaults
Observed
From system.cas_gc_log (event_type = 'Phase'):
| date |
rounds/day |
round duration |
defer_decision.ref_log_keys_listed |
defer_decision |
fold_ref_intake |
| 09-14 |
278 |
98 s |
8k |
3 s |
35 s |
| 09-16 |
242 |
191 s |
58k |
19 s |
77 s |
| 09-17 |
122 |
407 s |
296k |
97 s |
154 s |
| 09-19 |
80 |
777 s |
1.15M |
369 s |
197 s |
| 09-21 |
49 |
1456 s |
2.16M |
711 s |
452 s |
| 09-24 |
19 |
2800 s |
4.28M |
1424 s |
903 s |
From system.blob_storage_log, _log keys per day:
| date |
uploads |
deletes |
rounds × 5000 |
| 09-18 |
446k |
453k |
455k |
| 09-20 |
379k |
339k |
340k |
| 09-22 |
420k |
189k |
190k |
| 09-24 |
299k |
95k |
95k |
Deletes per day equal 5000 × rounds exactly. Blob graduation shows the same shape: entries_graduated = 5000 on every round since 09-21.
Root cause
Three per-round caps are fixed numbers that do not depend on the backlog: cas_gc_round_ref_cleanup_budget, cas_gc_round_graduation_budget and cas_gc_round_redelete_budget, all 5000 by default (Gc::cleanupRefObjects and the fold in CasGc.cpp).
Every round starts with a full LIST of cas/ns/stream/, which costs O(all _log keys). Once _log uploads per day exceed 5000 × rounds per day, the key population grows, the LIST and the fold take longer, rounds per day fall, and the delete budget per day falls with them. The loop is self-reinforcing and does not recover on its own.
The loop started on 09-16 with a new writer epoch, while the budget was still being spent on ~2M keys left from older epochs. The 09-21 "inflection" is only when the blob side became visible.
The 5000-key cap is small compared to its own cost: 5000 keys is 5 batch-delete requests, while the LIST it causes today costs ~4300 requests per round.
Not the cause
- The 09-21
url() access-limitation change: its errors are executeQuery rows on 09-22 with no CAS effect.
- S3 errors, lease loss, restarts (09-22 and 09-23 only changed the epoch).
Expected
GC keeps pace with a steady write rate. Per-round budgets should scale with the observed backlog, or be unbounded for write-once covered keys, which are pure batch deletes with no HEAD. The round should not need a full prefix LIST: each namespace _ckpt knows its head sequence and the fold seal knows the covered floor.
Observability gap: the ref_object_cleanup phase row carries only namespaces_planned, suppressed, trim_enabled. It should also carry objects_deleted, objects_pending and budget_exhausted.
Note: system.cas_mounts.pending_reclaim is condemned minus executed deletes for the current process and resets on restart, so it understates the durable backlog.
Workaround
Raise or remove the caps in the CAS disk config and restart the server. Pool config is read at mount time, so a config reload is not enough. The first rounds after restart will be long while the backlog drains.
<clickhouse>
<storage_configuration>
<disks>
<cas>
<type>object_storage</type>
<object_storage_type>s3</object_storage_type>
<metadata_type>cas</metadata_type>
<!-- existing endpoint / credentials / server_root_id stay as they are -->
<!-- 0 = unbounded; or a value well above the daily key rate divided by rounds per day -->
<cas_gc_round_ref_cleanup_budget>0</cas_gc_round_ref_cleanup_budget>
<cas_gc_round_graduation_budget>0</cas_gc_round_graduation_budget>
<cas_gc_round_redelete_budget>0</cas_gc_round_redelete_budget>
</cas>
</disks>
</storage_configuration>
</clickhouse>
To confirm recovery, watch ref_log_keys_listed and duration_ms fall across rounds:
SELECT event_time, round, duration_ms, entries_graduated, objects_deleted
FROM system.cas_gc_log WHERE event_type = 'Finish' ORDER BY event_time DESC LIMIT 10;
SELECT event_time, phase_metrics['ref_log_keys_listed'], phase_duration_microseconds / 1e6
FROM system.cas_gc_log WHERE event_type = 'Phase' AND phase = 'defer_decision'
ORDER BY event_time DESC LIMIT 10;
Describe the situation
On a two-replica cluster with CAS as the default disk, GC rounds slowed from 4 minutes to 48 minutes over 8 days. Deletes per round are pinned at exactly 5000. Retired-but-undeleted blobs reached ~520k. No data damage:
cas_logcorrectness invariants are clean, S3 upload error rate is ~0.03%, no lease loss.Environment
26.6.4.20001.altinityantalyacaswith a cache diskcas_cacheon top, used as the default disk_log) keys and ~320k blob retirements per dayObserved
From
system.cas_gc_log(event_type = 'Phase'):defer_decision.ref_log_keys_listeddefer_decisionfold_ref_intakeFrom
system.blob_storage_log,_logkeys per day:Deletes per day equal
5000 × roundsexactly. Blob graduation shows the same shape:entries_graduated = 5000on every round since 09-21.Root cause
Three per-round caps are fixed numbers that do not depend on the backlog:
cas_gc_round_ref_cleanup_budget,cas_gc_round_graduation_budgetandcas_gc_round_redelete_budget, all 5000 by default (Gc::cleanupRefObjectsand the fold inCasGc.cpp).Every round starts with a full LIST of
cas/ns/stream/, which costs O(all_logkeys). Once_loguploads per day exceed5000 × rounds per day, the key population grows, the LIST and the fold take longer, rounds per day fall, and the delete budget per day falls with them. The loop is self-reinforcing and does not recover on its own.The loop started on 09-16 with a new writer epoch, while the budget was still being spent on ~2M keys left from older epochs. The 09-21 "inflection" is only when the blob side became visible.
The 5000-key cap is small compared to its own cost: 5000 keys is 5 batch-delete requests, while the LIST it causes today costs ~4300 requests per round.
Not the cause
url()access-limitation change: its errors areexecuteQueryrows on 09-22 with no CAS effect.Expected
GC keeps pace with a steady write rate. Per-round budgets should scale with the observed backlog, or be unbounded for write-once covered keys, which are pure batch deletes with no HEAD. The round should not need a full prefix LIST: each namespace
_ckptknows its head sequence and the fold seal knows the covered floor.Observability gap: the
ref_object_cleanupphase row carries onlynamespaces_planned,suppressed,trim_enabled. It should also carryobjects_deleted,objects_pendingandbudget_exhausted.Note:
system.cas_mounts.pending_reclaimis condemned minus executed deletes for the current process and resets on restart, so it understates the durable backlog.Workaround
Raise or remove the caps in the CAS disk config and restart the server. Pool config is read at mount time, so a config reload is not enough. The first rounds after restart will be long while the backlog drains.
To confirm recovery, watch
ref_log_keys_listedandduration_msfall across rounds: