Skip to content

feat(nvsnap): download each model once per cluster into a shared volume for Helm functions - #2106

Open
balajinvda wants to merge 120 commits into
nvsnap/helm-electionfrom
nvsnap/helm-artifacts
Open

balajinvda wants to merge 120 commits into
nvsnap/helm-electionfrom
nvsnap/helm-artifacts

Conversation

@balajinvda

@balajinvda balajinvda commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

Why

Customers deploy Helm charts of N identical GPU workers, and every worker
downloads and compiles on its own. The gate-and-promote election in #2104
cannot serve multi-node instances (LWS, StatefulSet groups, Dynamo
podgangs): holding any member deadlocks the group. The invariant we want
is simpler and storage-independent: the model is downloaded once per
cluster and every pod attaches it; compile caches are shared the same
way. Design and the decisions taken with the product owner:
docs/proposals/helm-shared-model-volume.md.

What changed

One model-volume lifecycle on every storage. The primary volume lives in
the nvsnap namespace; completion is a labelled, retained PV and a released
claim; each function namespace gets a per-namespace view of it (read-only
for readers, read-write for the download Job). The storage mode decides
only two things: the claim's access mode and how the download reaches
the primary (block: staged into a pod-local emptyDir and copied by the
agent, sized from the measured bytes; shared filesystem: the Job writes
through its namespace's read-write view while readers already hold their
views). Admission never waits for a volume to bind.

  • internal/modelid: identity and landing volume for any chart. Identity
    as a URI (hf://, ngc://, s3://, nim://) from engine args
    (including positional vllm serve <x> and $VAR expansion), engine
    env, download init containers (NGC CLI, hf download, S3, KServe
    storage-initializer), NIM images, and LeaderWorkerSet workers via the
    leader template. Landing volume kind decides substitutability; a
    customer's shared PVC, hostPath or image is left alone.
  • internal/modelvolume + webhook: one download Job per identity per
    cluster (nvsnap-model-dl-<key>, idempotent create, so no election),
    built from the chart's own download init wrapped to touch a marker, or
    from hf download on the engine image with credentials forwarded.
    Every workload pod is a reader: its download init becomes a wait for
    the marker (deadline, then self-download), the engine is started
    offline when it used to download itself. Charts whose engine script
    fetches an artifact nvsnap has no recipe for (an NGC model pulled in
    the main container) keep their download; the agent on that node
    captures the finished tree after Ready and later pods read the copy.
    Every container the feature creates gets requests and limits, dropped
    capabilities, no privilege escalation and the chart's own posture, so
    enforced Kyverno baselines pass.
  • Views are sized from the bytes the download measured. The Job reports
    the tree size in its termination message (kubelet unmounts a finished
    Job pod's volumes at once), the staged copy and the capture measure the
    same number, and the agent stamps it on the primary; a 62 GB model on a
    4Ti nominal filesystem claim gets 62Gi views.
  • Compile-cache volume, same shape: per configuration key and rank, the
    first Ready group is collected (startup plans, kernel and inductor
    caches, the transformers modules cache), later pods seed their local
    cachedir from it. Compile caches stay in a local cachedir on every
    storage; the model view is read-only. Capture-source pods collect a set
    too, so the first deployment hands the second one its startup plan.
    Credential paths under HOME (NGC CLI config, Hub tokens, and the like)
    never travel in a set.
  • Failure handling on the shared-filesystem Job path: a fresh filesystem
    root is root-owned and the Job runs as the function's user (uid 1000
    for Dynamo and NIM images), so the agent on the Job's node opens the
    primary root through the Job pod's kubelet mount and the writer waits
    until its landing is writable. A failed Job is recorded with its reason
    and the download container's last words, the writer view is retired,
    the primary claim is kept for the retry, and a failure marker next to
    the completion marker sends readers already waiting to their fallback
    at once instead of at their deadline. A successful retry clears the
    record.
  • Reaper on every agent (agent.modelVolume.reapInterval, default 10m)
    for model and cache volumes alike: unreferenced views are retired with
    their holders, view PVs whose claim is gone are removed, incomplete
    primaries released for 15 min are removed with their storage; complete
    primaries are kept.
  • nvsnap.io/cache-salt folds an operator-chosen value into the capture
    and cache identity for changes the identity cannot see (a wheel
    installed at start, a floating tag re-pointed).
  • The blob store tier is removed (binary, package, agent uploader, server
    endpoints, chart component, build targets); no current path used it.
  • Chart: agent.modelVolume.{enabled,waitDeadline,reapInterval}, storage
    profile modelVolume: {mode, storageClass, size, minSize, readerMode}
    (block by default on shared-volume strategies, rwx must be declared),
    agent.runtime with the runtime socket mounted at its host path for
    CRI-O, a fully qualified smoke-hook image for short-name enforcement,
    and smoke-hook and rollout timeouts that fit a serial agent roll on
    large clusters. The agent build fails loudly when the base image is
    missing from the target registry.

Customer Release Notes

Helm chart functions download each model once per cluster; every other
worker, later version, scale-up and redeploy attaches the downloaded
volume, and compile caches are shared between pods of one configuration.
No chart change. Requires NVMesh or a shared filesystem class such as
OCI File Storage.

Plan Summary

Helm: new values (off by default), one extra Bidirectional hostPath on
the agent DaemonSet, ClusterRole verbs (PVC update/patch, pods patch,
namespaces get, PV create/update/delete). Primaries and the agents' copy
holders live in the release namespace; views live in function
namespaces. No new controllers or CRDs.

Usage

helm upgrade nvsnap deploy/helm/nvsnap --set agent.modelVolume.enabled=true
kubectl get jobs,pvc -A -l nvsnap.io/model
kubectl get pv -l nvsnap.io/model-complete=true
kubectl get pods -A -l nvsnap.io/model-pending=true

Testing

NVMesh (block mode), stock vllm-workers chart, Qwen2.5-32B TP=4, two
replicas: first deploy one Job, 65 GB downloaded once, both Ready +591 s
with 0 downloads in the pods; reinstall both Ready +136 s; cold 325 s.
PVC reader mode, Qwen2.5-14B: first deploy both Ready +344 s, reinstall
+118 s. Real NVCF Helm function with the chart's own NGC download
(954 MB): claim sized at 2Gi from the measured bytes instead of 512Gi,
both Ready +230 s; compile-cache seeding took torch.compile from 23 s to
3 s per pod.

OCI File Storage (shared-filesystem mode), GB300 arm64 nodes, CRI-O,
NVCF storage class, 2026-10-02 and 2026-10-03:

  • Qwen2.5-32B vLLM, 62 GB: cold 6 min 37 s (Job writes through the
    writer view, cache set collected after Ready); warm 1 min 38 s;
    scaling 1 to 4 replicas on fresh nodes, each Ready 72 to 78 s after
    creation reading weights at about 2.3 GB/s per node from one
    filesystem, no downloads, no attach wait.
  • kimi-k3, 1.56 TB, two pods TP 8, engine-script download captured after
    Ready (image pulls excluded): cold admission to Ready 46 min 30 s
    (download 31 min, capture into the primary 24 min in the background);
    second deployment, unseeded, 20 min 42 s; third deployment, seeded,
    12 min 20 s (weights 425 s from the filesystem view, profiling 0 s
    with the saved plan, graph capture 80 s, no restarts). Against the
    same chart on NVMesh: 14 min 21 s warm.
  • Dynamo function with an HF model as a non-root user on a fresh
    filesystem primary: this is the run that found the root-permission and
    failed-Job gaps fixed above; the fixes are covered by unit tests and
    the function redeploy after the next agent roll is the e2e check.

Unit: identity forms and classifier; Job spec derivation and
idempotence; webhook matrix (block and shared-filesystem readers,
complete-at-admission minting, the same flow asserted on every storage,
plain pods keep every annotation, Job volume carry-over, hardening and
posture inheritance, capture sources, injected download, customer
volumes left alone, writer and reader scripts); controller completion,
release, detach gate, per-namespace view minting, bytes recording from
every fill path, capture from an engine download, root opening on the
Job's node, failed Job handling and the failure marker, cache-set
collection with credential exclusion, reaper rules. go test ./...
green.

Notes

Stacked on #2104 (base nvsnap/helm-election). Measured on the way and
recorded in the design doc: the CuTeDSL warmup cannot be cached from
outside the engine (cute.compile() forces no cache); CUDA graph capture
is not cacheable; on both NVMesh and NFS the weight read sits at about
2 GB/s per node against 4.4 GB/s from local NVMe, which points at a
node-local tier under the shared primary as the next step. The first
start of a multi-worker engine can race in transformers' dynamic-module
copy; seeded starts avoid it because the modules cache arrives populated.

References

Relates to #2099

Related Pull Requests

#2104, #2101

Dependencies

None

First step of docs/proposals/helm-shared-model-volume.md. For any GPU pod at
admission, internal/modelid answers which artifact the pod will download,
as a URI shared by every chart, version and namespace that names it, and
where the bytes land and what backs that path today.

Identity sources, as the field uses them: engine args (--model,
--model-path, --model=, positional `vllm serve <x>`, --revision), engine
env (HF_MODEL_ID, MODEL_ID, MODEL_PATH, with $VAR expansion from the
container's own env), download init containers (NGC CLI with
NGC_MODEL_NAME, huggingface-cli download with --local-dir, aws s3 sync,
KServe storage-initializer args), NIM images with NIM_MODEL_PROFILE, and
for LeaderWorkerSet workers that name no model the leader template of
their group. URIs: hf://org/repo[@rev], ngc://org/team/model:ver,
s3://bucket/key, nim://image@profile, path:///abs for pre-filled paths.

The landing volume is the mount at or above the download destination in
the main container: emptyDir or rootfs are substitutable, a customer's
PVC, hostPath or other volume means sharing is already solved and the pod
is left alone. KServe pvc:// is likewise not a downloader.

Tests are built from real specs: the NVCF Helm function observed on prd11
(NGC init download, positional MODEL_PATH), the Dynamo operator sample,
SGLang, NIM, KServe, huggingface-cli and S3 inits, the upstream LWS vLLM
example, and non-downloaders (Dynamo frontend, Ray worker, etcd).
Mutation-checked: PVC marked substitutable, revision split dropped and
env expansion dropped each turn tests red.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
Step 2 of docs/proposals/helm-shared-model-volume.md. For a GPU pod whose
model identity resolves (internal/modelid), the webhook replaces the
volume the download lands in with the per-identity model volume and
turns the download into a write-once step; no pod is gated.

internal/modelvolume names the claim per identity (nvsnap-model-<key>),
creates the writer claim (ReadWriteMany on a distributed filesystem,
ReadWriteOnce on NVMesh), looks the identity up cluster-wide and records
completion as a label. The webhook elects the writer with the existing
Lease elector keyed by identity; the writer's landing emptyDir becomes
the claim, its engine mounts the model read-only, and its download init
is wrapped to touch <volume>/.nvsnap-complete on success. Readers on a
distributed filesystem mount the same claim and their init waits for
the marker, downloading themselves after the deadline (decided: always
fall back). Readers on NVMesh keep their emptyDir, are labelled pending
and wait for the agent to bind the completed volume in (step 3). Charts
whose engine downloads itself get an injected huggingface-cli init that
lands the model in the volume, with the engine started offline and the
registry credentials forwarded. Compile caches are redirected into
<volume>/.nvsnap/cache/<engine key> on a distributed filesystem and into
the local cachedir on NVMesh; the model env of the cachedir template is
left alone because the model lives in the landing volume now.

Tests use the prd11 function shape and a stock vLLM Deployment: writer
on Block, pending reader on Block, reader on RWX, injected init with
forwarded token and offline engine, completed volume skipping the
election, and the customer-PVC / no-GPU / no-model pods left alone.
Mutation-checked: read-only mount dropped, reader fallback dropped,
Block reader substituting the volume, and writer marker dropped each
turn tests red.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
…Mesh

Step 3 of docs/proposals/helm-shared-model-volume.md. The agent gains a
model volume controller. For writers, on any node: when the download init
the webhook named on the pod exits 0, the writer claim is labelled
complete and, on block storage, SharedVolumePromoter.MintReadOnly exposes
it as a read-only claim (primary PV retained, secondary static PV with the
namespace-rewritten NVMesh handle, pre-bound claim). The writer claim
stays because the writer still runs on it and later writers must find it
complete. For pending readers on this node: once the identity is complete
the read-only claim is minted in the reader's namespace, attached to the
node through a mount-holder pod, bind-mounted read-only onto the reader's
hostPath landing under the Bidirectional overlays root
(<overlay-root>/models/<key>), and the pod is un-pended. The writer touched
the marker inside the volume, so the reader's wait init sees it as soon as
the bind lands. Every agent watches writers (idempotent marking and
minting); only the reader's own agent binds.

The webhook gives Block-mode readers a hostPath landing with
HostToContainer propagation on the engine and init mounts, and names the
writer's download init for the agent. Storage profiles gain
`modelVolume: {mode, storageClass, size}`; block is the default for
shared-volume strategies, rwx must be declared. Agent flags
--model-volume and --model-volume-wait-deadline, Helm
agent.modelVolume.{enabled,waitDeadline}, and pods patch in the agent
ClusterRole.

Tests with fake clients: writer completion marks and mints the read-only
PV and claim with the handle rewritten and the primary retained; a
pending reader on its node gets the claim in its own namespace, one
attach, one bind at the hostPath, and is un-pended, a second reader of
the same identity reuses the bind, readers on other nodes are ignored.
Mutation-checked: node filter dropped, completion on any exit code,
binding before completion, and un-pend dropped each turn tests red.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
…b in huggingface_hub 1.x

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
The first cluster run of the model volume on NVMesh exposed a constraint
the design had missed: a volume attached read-write by a running pod
cannot be attached read-only anywhere else (AttachVolume failed on the
read-only PV while the writer pod held the primary). The L2 promote never
met it because it deletes the writer claim before readers attach. So the
download step cannot live inside a pod that goes on to serve.

The download is now a Job per identity per cluster, nvsnap-model-dl-<key>,
created idempotently by the webhook on first sight: the chart's own
download init (image, command, env, pull secrets, tolerations) wrapped to
touch the marker, or `hf download` on the engine image when the engine
fetches the model itself. Its exit releases the volume. Every workload pod
is a reader; there is no writer pod and no Lease election on this path,
because Job create is atomic. The agent completes the identity when the
Job succeeds and mints the read-only claim on block storage; readers are
bound as before. The design doc records the constraint and the reason.

Tests: Job spec derived from the NGC init and from a stock vLLM pod (hf
download, forwarded token, claim mounted at the init's path), idempotent
create, no Job once the volume is complete, controller completes on Job
success only. Mutation-checked: Job never created, Job created when
complete, and completion on a running Job each turn tests red.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
…detaches

A Succeeded pod keeps its volumes attached. On NVMesh that let the
read-only attach succeed only on the Job's own node (dev1 2026-09-26);
every other node failed until the pod was deleted. The Job now carries
ttlSecondsAfterFinished=30; completion state lives on the claim label.
MarkComplete tolerates the update race between agents by re-reading.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
…a dedicated host root

Two findings from the second dev1 run of the model volume on NVMesh.

The retained volume must have no read-write attachment before any node can
attach it read-only: with the Job's claim still bound, the read-only attach
succeeded only on the Job's own node and failed everywhere else. On block
storage MarkComplete now labels the retained PV with identity and
completion, sets Retain, and deletes the download claim, the same release
the L2 promote performs; Lookup reads completion from the PV; readers'
read-only claims are minted from that PV, and only after no
VolumeAttachment references it (Detached). RWX mode keeps the shared claim
and labels it as before.

The bind root moves out of the overlays root: the overlay sweeper removes
entries it does not own and deleted the model binds under it. Completed
model volumes are bound under a dedicated Bidirectional hostPath
(agent.hostPaths.nvsnapModels, default /var/lib/containerd/nvsnap-models),
mounted at the same path in the agent and on the host.

Tests: block completion releases the claim and labels the PV, RWX labels
the claim, the reader is not served while a VolumeAttachment exists and is
served once it is gone (mutation-checked: dropping the detach gate turns
the test red), the read-only PV carries the reader namespace's handle.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
…on dev1

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
…d when stale

The controller kept "bound" in memory. After the volume for an identity
was replaced (and after a manual unmount during cleanup on dev1) it
un-pended readers without a bind and they waited on an empty hostPath.
The mount table is now the truth: a reader is served only when the
device mounted at the bind target belongs to the identity's primary PV
(NVMesh device csi-id against the volume handle); a missing mount is
bound again, a mount of a replaced volume is unbound and redone, and a
bind that leaves nothing mounted does not un-pend the pod.

Tests: rebinding after an unmount, replacing a stale bind, device/handle
matching. Mutation-checked: dropping the stale unbind and the missing
mount check each turn tests red.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
@balajinvda
balajinvda requested a review from a team as a code owner September 26, 2026 17:02
@balajinvda
balajinvda requested a review from vrv3814 September 26, 2026 17:02
@coderabbitai

coderabbitai Bot commented Sep 26, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

🗂️ Base branches to auto review (1)
  • main

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration
  • Configuration used: Repository: NVIDIA/nvcf/.coderabbit.yaml
  • Review profile: CHILL
  • Plan: Enterprise
  • Run ID: de7d48ad-9e18-4cc0-a042-6508dbe8383d

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

Three gaps against the shape of a real NVCF Helm function (an NGC init
download with a secrets file and a scripts ConfigMap, in a namespace
with Kyverno enforced):

- the download Job carries every volume the chart's init mounted, with
  only the landing volume redirected to the claim; before, the Job had
  only the landing mount and the download could not read its key
- block readers default to a `pvc` reader mode: the pod references the
  read-only claim in its namespace, the webhook mints it at admission
  when the volume is complete, and any agent mints it after the Job
  when the primary detaches; no hostPath, so disallow-host-path passes.
  The former bind-in path is `readerMode: hostPath` in the storage
  profile, for gang-scheduled charts
- Job pods get default requests and limits, a seccomp profile, dropped
  capabilities and no privilege escalation when the init set none

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
The engine-download Job container had requests without limits and the
injected reader init had neither, so both failed
require-pod-requests-limits where Kyverno is enforced (seen on dev1).
One helper now gives every container the model volume creates the same
baseline: requests and limits unless the chart sized it, no privilege
escalation, dropped capabilities. Mount propagation on the injected
init is set only in hostPath reader mode, the one case where the host
binds a volume in after the pod started.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
…iners

The injected init and the download Job still failed
require-run-as-non-root and the by-name CAP_NET_RAW check on dev1. The
download runs the engine's image and must own what it writes, so both
containers now inherit runAsNonRoot, runAsUser, runAsGroup and seccomp
from the chart's main container, and the Job pod inherits the pod
security context (fsGroup). NET_RAW is dropped by name next to ALL.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
A capture pod gets an empty emptyDir at the cachedir root with
NIM_CACHE_PATH and HF_HOME pointed at <root>/model and the compile
caches at <root>/cache. vLLM and Hugging Face create their own trees,
so this went unnoticed; NIM checks NIM_CACHE_PATH at startup and exits
("Unable to read from NIM_CACHE_PATH", first NIM function on dev1 with
the NVCA integration on, 2026-09-28). An init container on the
workload's own image now creates both subtrees, sticky world-writable,
on capture pods and on the model volume's cache-only emptyDir. It
inherits the workload's user posture and is hardened like the other
containers nvsnap injects.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
The prewarm became a storage-profile switch on 2026-09-24 but the
built-in NVMesh entry never used it, so every NVMesh restore still ran
the sweep. It is neutral to negative there: a single reader already
saturates the volume (70B TP=4: 247 s with, 257 s without) and the
sweep reads the whole tree, unused files included. The first NVCA-driven
NIM restore on dev1 spent 67 s reading 30 GB, half of it an original
.pth nothing opens, to save a 2 s weights load, and finished no faster
than a cold start. Hyperdisk ML keeps prewarm on; the ConfigMap overlay
and NVSNAP_PREWARM still override either way.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
The sweep fanned out per file with xargs -P 6 -n 16. A checkpoint tree
is a dozen files, four of them the weights, so one process read the
whole tree serially: 30 GB in 67 s on dev1 while the engine's own mmap
load of 15 GB took 33 s, and the restore finished no faster than cold.
Measured on that NVMesh volume, one stream reads 270 MB/s, four 1.0
GB/s, eight 2.1 GB/s. Files above 64 MiB are now split into 256 MiB
ranges read by dd under xargs -P; small files are still batched through
cat. The walk follows symlinks so a Hugging Face snapshot is read by
entry name, skipping original/ (the .pth copy nothing opens) and the
blobs directory (the same bytes again). The NVMesh profile keeps the
prewarm on and uses eight readers; this supersedes the previous commit's
off default, which was diagnosing the wrong cause.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
A block-mode model volume was created at admission at a fixed 512Gi
ceiling because nothing knew the model's size yet; on a RAID-10 NVMesh
class a 1 GB model reserved about 1 TiB of raw NVMe. Registries do not
all expose sizes and not every storage system can grow a volume, so the
download Job now stages into a pod-local emptyDir, the same shape as the
single-GPU cachedir capture. When it succeeds, the agent on that node
measures the bytes on disk, creates a claim of measured size plus ten
percent (whole GiB, profile modelVolume.minSize as the floor) in the
nvsnap namespace, copies the tree in through a mount-holder, labels the
retained PV complete, releases the claim and deletes the Job. The
primary no longer belongs to a function namespace. RWX mode is
unchanged.

A reaper on every agent removes read-only model PVs whose claim is gone
(Released, or bound in a deleted namespace) as objects only, and
abandoned incomplete primaries with their storage. The agent gains
namespaces/get so a lookup error can never read as a deleted namespace.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
@balajinvda
balajinvda force-pushed the nvsnap/helm-artifacts branch from 7927f5c to 65303c8 Compare September 29, 2026 04:45
Every Helm pod on block storage compiled into its own cachedir emptyDir:
23.5 s of torch.compile per pod on a 0.5B model, more on larger ones,
while a warm cache brings it to 2.9 s (measured 2026-09-29). The model
volume provisioner now has a second kind, cache, with its own names and
labels. The first pod of an engine configuration and ordinal that
becomes Ready is captured by the agent on its node: its cache subtree is
measured, copied into a claim of that size in the nvsnap namespace and
the PV labelled complete. Later pods with the same key get a read-only
claim minted at admission and a seed init copies it into their cachedir
before the engine starts. Atomic claim creation picks the source among
pods sharing a key; failures release the claim and, after the last
attempt, record a failure so admissions stop waiting.

Caches are keyed per ordinal, never merged: ranks of a tensor-parallel
group write rank-specific directories and the shared autotune results
differ in content between ranks.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
…re source

A pod whose engine downloads a model nvsnap has no recipe for was
stamped for capture of its landing volume and returned from admission
before the compile-cache injection, so the cold deployment of such a
chart never collected a set and the first warm deployment profiled
and compiled again (kimi-k3 on GB300, 2026-10-03: 4 min 38 s). The
capture source now gets the cachedir and the set stamping like any
reader; it still keeps the chart's own download step.

Those pods write their download credentials under HOME, which is the
collected root (the NGC CLI config holds the API key), so the set
collector and the rank tar server now exclude credential paths, and a
skipped directory is no longer descended by the tar writer.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
A download Job on storage that is shared while written runs with the
function's posture; a new filesystem is root-owned and 0755, so a
uid-1000 engine image could not create its first directory in the
primary and the Job crashed into its backoff limit (Qwen3-0.6B on an
FSS primary, 2026-10-03). The agent only reacted to succeeded Jobs, so
nothing was recorded and the readers waited out their deadline.

The agent on the Job's node now opens the primary root through the
Job pod's kubelet mount, the writer waits until its landing is
writable, and a failed Job records a failure with the Job's reason
and the download container's last words, retires the writer view and
keeps the primary claim for the retry.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
… job fails

Readers of a primary that is shared while written hold views and wait
on the completion marker; when the download Job gave up they waited
out their deadline (30 min on the cluster) before downloading for
themselves. The Job's node now writes a failure marker next to the
completion marker with the Job's reason, the reader's wait init runs
its fallback the moment it appears and logs the reason, and the next
writer removes the marker before it downloads. A successful retry
clears the failure record.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
…d ref

The base image was once built from a local CRIU checkout on a different
line than the fork and ref versions.sh pins for the OSS build, so two
builders produced two different CRIUs. The base build now fails when
the resolved source is not at NVSNAP_CRIU_REF (override only with
NVSNAP_CRIU_ALLOW_REF_MISMATCH=1), the app build warns when the last
recorded base commit is not the pinned ref, and check-criu-ref reports
both.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
The base image is now built from github.com/balajinvda/criu at
74b4170aa (tag nvsnap-base-v0.0.21): the criu-v2 io_uring engine line
plus the epoll shared-fd tolerance and the late-device-resume failure
propagation ported from the earlier line. The earlier line, which the
previous base was built from, is archived. The fork documentation
names the fork of record and what the line carries.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
…knows

Upstream CRIU renamed --compress-region to --compress-block when it
unified the page and region modes. The criu-v2 dump now probes the
bundled criu once and uses whichever spelling it advertises, so the
opt-in NVSNAP_CRIU_V2_COMPRESS setting keeps working across the base
image that carries the rename; "block" is accepted alongside "region".

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
… criu that will not exit

The criu-v2 restore of a model-volume reader failed on the first file
it tried to reopen: engines keep JIT artefacts open or mapped under the
webhook-injected cachedir (Triton launchers, the FlashInfer log), and
the placeholder has no copy of that emptyDir. The cachedir mount now
rides the same replay that already carries /dev/shm, so the placeholder
gets the files at their original paths before CRIU runs.

After that failure criu stayed alive waiting on its restorer tasks and
the agent waited with it. The restore now runs criu in its own process
group, kills the group on cancel, and cancels once the restore log has
reported a failed restore and a grace period has passed.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
@balajinvda
balajinvda force-pushed the nvsnap/helm-artifacts branch from 9d781b5 to 6f705dc Compare October 4, 2026 14:46
…xtracting into it

The cachedir replay extracts into /opt/nvsnap inside the placeholder's
mount namespace, but a placeholder only has the mountpoints its own spec
declares and tar refuses a missing directory. Create the mountpoint on
the placeholder's rootfs before extracting; /dev/shm is unaffected since
it always exists.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
…dev)

Base v0.0.22 is built from the official fork's criu-dev after the rebase onto
upstream criu-dev bf186f060. Upstream renamed --compress-region to
--compress-block; the agent already probes for whichever flag the bundled
criu accepts. Validated end to end on x86 (vllm-small: checkpoint 1m16s,
restore pod Ready 33 s, inference after restore OK).

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
…p/helm-artifacts

Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>

# Conflicts:
#	src/compute-plane-services/nvsnap/internal/webhook/cachedir.go
#	src/compute-plane-services/nvsnap/internal/webhook/cachedir_noshim_test.go

@FamousDirector FamousDirector left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the incremental change against declared base nvsnap/helm-election (7642bfc). Four P1 availability/data-integrity blockers are detailed inline: cross-namespace concurrent writers, releasing the writer claim after a failed completion update, fallback downloads into a read-only reader view, and writes escaping a partially extracted rank directory on retry.

Validation: existing modelid/modelvolume tests passed; webhook, hostlibs, and tarstream tests passed with a temporary Darwin portability overlay for unchanged treecopy code. Focused checkpointstore shared-volume/view and rootfsonly composer tests passed. Five isolated regression probes reproduced the four findings. Linux agent tests cross-compiled successfully; Helm lint and incremental diff whitespace checks passed. Full checkpointstore/rootfsonly filesystem tests could not run on macOS because the unchanged Linux sendfile path panics there. No live GPU/CSI/cluster validation was performed. All probes used temporary overlays; tracked files remain unchanged.

// election.
func (p *Provisioner) EnsureDownloadJob(ctx context.Context, uri, ns, claim string, step DownloadStep) (string, error) {
name := JobName(uri)
if _, err := p.Kube.BatchV1().Jobs(ns).Get(ctx, name, metav1.GetOptions{}); err == nil {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Elect the download owner across namespaces

When two Helm functions in different namespaces first request the same incomplete model URI, each namespace creates its own Job here. ensureDownload reuses the cluster-wide primary but gives both Jobs read-write views of the same CSI handle. A fake-client admission probe produced two Jobs sharing that handle. Running the generated writer scripts concurrently also left the completion marker present while the slower writer truncated a model file: readers can start against partially overwritten weights after the first Job reports completion. This breaks both the once-per-cluster requirement and model integrity.

Acquire one atomic, cluster-wide owner for the model identity before creating the namespaced Job. Only that owner should download/write; other namespaces should remain readers, with ownership released or recovered on failure.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 8e09512. The webhook now claims a cluster-wide download owner before creating any Job: a Lease in the nvsnap namespace whose create is atomic, so concurrent admissions agree on one owning namespace. Other namespaces create no Job and no writer view and read the primary like any other reader. An owner whose Job is gone past a 2 minute grace period is retired with UID and resourceVersion preconditions and the download claimed again; a recorded failure or the completion releases the record. I reused the Lease-create primitive rather than LeaseElector itself, because nvsnap-server lists and reconciles the capture-election Leases. Tests: TestClaimDownload_OneOwnerAcrossNamespaces (concurrent claims from two namespaces, grace period, stale takeover, release on failure) and TestModelVolume_SharedFilesystem_OneWriterAcrossNamespaces (two namespaces admitted, one Job, no second writer view; it failed with 2 Jobs before the fix).

pv.Annotations[k] = v
}
pv.Spec.PersistentVolumeReclaimPolicy = corev1.PersistentVolumeReclaimRetain
if _, err := p.Kube.CoreV1().PersistentVolumes().Update(ctx, pv, metav1.UpdateOptions{}); err != nil && !apierrors.IsConflict(err) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Persist completion and Retain before releasing the writer claim

A PV resource-version conflict is ignored here, so the following PVC deletion runs even though neither the completion labels nor Retain were persisted. This can occur when a CSI controller or another agent updates the PV between Get and Update. With a dynamically provisioned block PV whose reclaim policy is Delete, releasing the last writer claim then allows the completed model volume to be destroyed. The injected-conflict probe confirmed this method returns nil and deletes the claim while the PV remains incomplete with reclaim policy Delete; the completed artifact is also no longer discoverable through Lookup.

Retry conflicts against a freshly fetched PV, and delete the writer claim only after the required completion metadata and Retain policy have successfully persisted, or a reread confirms them.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in a1060fd. MarkCompleteClaim now retries PV update conflicts against a freshly read PV (bounded at 5 attempts) and deletes the writer claim only once the stored PV is labelled complete and set to Retain. Persistent conflicts return an error and leave the claim in place. Test: TestMarkComplete_ConflictKeepsClaimUntilPersisted, with one injected conflict (retried, then released) and permanent conflicts (error, claim kept, PV still Delete). Both cases failed before the fix.

fallback = download
}
failed := shellQuote(path.Join(path.Dir(marker), modelvolume.FailedMarkerFile))
return fmt.Sprintf("set -e\nd=0\nwhile [ ! -f %[1]s ]; do if [ -f %[4]s ]; then echo \"nvsnap: shared download failed: $(cat %[4]s 2>/dev/null)\"; %[3]s; exit 0; fi; if [ $d -ge %[2]d ]; then echo 'nvsnap: marker deadline passed; downloading locally'; %[3]s; exit 0; fi; sleep 5; d=$((d+5)); done\necho 'nvsnap: model complete, skipping download'\n", shellQuote(marker), deadline, fallback, failed)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Give failed readers a writable fallback landing

For PVC readers admitted while an RWX download is in flight, readerLanding replaces the original landing with a CSI view whose PV has CSI.ReadOnly=true. This script then runs the original download against that same landing when the failure marker appears or the wait deadline expires. Any fallback that creates model files fails on the read-only filesystem, leaving the function in an init-container restart loop rather than recovering. A generated-script probe with a failed marker and a non-writable landing confirmed the write fails. Recording the shared failure only bypasses mutation for later admissions; already admitted readers retain this mount.

Route fallback into a writable per-pod landing that the engine also consumes, or recreate the affected readers on their original download path after recording the failure. Reusing the read-only shared view cannot implement the advertised local fallback.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 2c36257, taking your second option. A writable per-pod landing would need the engine to read a second path, which this design cannot provide without rewriting the chart. PVC readers no longer fall back in place: on the failure marker or past the deadline their wait stops with the reason and a non-zero exit, instead of downloading into the read-only view. When the download Job fails, the agent on the Job node records the failure and deletes every waiting reader of that model that is not Ready. Their controllers recreate them, and admission sees the failure record and leaves them on their own download path. hostPath readers keep the in-place fallback because their landing is writable until the agent binds. Tests: TestReaderScript_NoInPlaceFallbackOnReadOnlyView, TestModelVolume_SharedFilesystem_ReaderHasNoInPlaceFallback and TestModelVolumeController_FailedSharedDownloadRecreatesWaitingReaders. The design doc failure table is updated to match. One limit remains: a writer that hangs without failing still leaves readers waiting, as before; the Job failing is what releases them.

mode = 0o777
}
}
f, err := os.OpenFile(target, os.O_CREATE|os.O_TRUNC|os.O_WRONLY, mode|0o600)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Confine extraction writes when retrying partial rank streams

fetchRank retries into the same destination after an interrupted response, retaining symlinks from the first extraction. A later regular-file entry uses os.OpenFile here, which follows those existing links. The lexical link check below does not account for intermediate symlink resolution, so writes can escape the rank directory and truncate other files. A temporary-directory probe using two archives emitted by the actual Write function reproduced this: an interrupted first stream left relative links; after the source entry changed to a regular file, the retry overwrote a sibling file outside the extraction destination. Cache collection must not corrupt neighboring cache or agent filesystem data when a transfer retries.

Use filesystem operations confined to the destination root for all extraction paths, rejecting escaping symlink resolution. Extract retries into fresh staging directories and publish only a complete tree.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in b0ea0ee. Extract now refuses any entry beneath a symlinked directory, replaces a symlink at an entry path instead of following it, and opens files with O_NOFOLLOW. The rank fetch now uses the new ExtractFresh: each attempt extracts into a fresh staging directory next to the destination and replaces it only when the whole stream extracted, so a retry never merges into what an interrupted attempt left. Tests: TestExtract_RefusesWritesThroughSymlinkedParents (your d -> . then d/f -> ../victim chain), TestExtract_ReplacesSymlinkAtFilePath (both overwrote the victim before the fix) and TestExtractFresh_PublishesOnlyCompleteTrees.

The Bazel BUILD files in nvsnap were not regenerated as packages and
files changed on this stacked branch; check-gazelle only runs on pull
requests to main, so the drift was not reported.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
Carry main's unqualified-storage change up the stack. The model volume
is built after the L2-disabled return; it already required a matched
profile and the shared-volume promoter, so behavior is unchanged.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
…default cache

An engine that sets neither HF_HOME nor HF_HUB_CACHE reads the Hugging
Face default under HOME, and nvsnap mounts the model volume at that
default path. The model-volume path also moves HOME into the pod-local
cachedir, so the engine looked for the model under the cachedir and,
with HF_HUB_OFFLINE set, exited with LocalEntryNotFoundError.

Set HF_HOME to the landing when the container names no cache location,
by value or by reference, and move HF_MODULES_CACHE off the read-only
volume in that case as well.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
…ined

MarkCompleteClaim ignored a conflicting PV update and deleted the writer
claim anyway. When a CSI controller or another agent wrote the PV between
the read and the update, neither the completion labels nor the Retain
policy were stored, and releasing the last claim of a Delete-policy volume
could destroy the completed model.

Retry conflicts against a freshly read PV, and release the claim only
once the stored PV is complete and retained. Persistent conflicts return
an error and keep the claim.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
Extract checked each symlink lexically but followed links when writing,
so links that resolve through one another, or a link left by an
interrupted transfer, let a later entry write outside the destination.
The rank fetch retried into the same directory, keeping those links.

Refuse entries under a symlinked directory, replace a symlink at an
entry's own path instead of following it, and open files with
O_NOFOLLOW. Fetch each rank attempt into a fresh staging directory and
publish it only when the whole stream extracted.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
…tion stack (#2261)

Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
@balajinvda
balajinvda requested a review from a team as a code owner October 6, 2026 18:10
Two Helm functions in different namespaces admitting the same incomplete
model each created a download Job, and on a shared filesystem both wrote
the one primary through their own read-write views. The completion marker
could appear while the slower writer was still truncating a file.

Record the downloading namespace in a Lease in the nvsnap namespace; the
create is atomic, so concurrent admissions agree on one owner. Other
namespaces create no Job and no writer view and read the primary like any
reader. An owner whose Job is gone past a grace period is retired with
preconditions and the download claimed again; a recorded failure or the
completion releases the record.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
A PVC reader admitted while a shared-filesystem download runs mounts the
read-only view of the primary at its landing. On the failure marker or
past its deadline its wait init ran the original download into that
view, which cannot be written, so the pod restarted its init forever.

PVC readers no longer fall back in place: the wait stops with the reason.
When the download Job fails, the agent on the Job's node records the
failure and deletes every waiting reader that is not Ready; their
controllers recreate them and admission, seeing the failure record,
leaves them on their own download. hostPath readers keep the in-place
fallback, since their landing is writable until the agent binds.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
balajinvda added a commit that referenced this pull request Oct 6, 2026
Bring in the squashed #2261 and the #2104 and #2106 review fixes. The
Makefile .PHONY list keeps check-gpushare.

Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants