Repository navigation
feat(nvsnap): download each model once per cluster into a shared volume for Helm functions - #2106
balajinvda wants to merge 120 commits into
Conversation
First step of docs/proposals/helm-shared-model-volume.md. For any GPU pod at admission, internal/modelid answers which artifact the pod will download, as a URI shared by every chart, version and namespace that names it, and where the bytes land and what backs that path today. Identity sources, as the field uses them: engine args (--model, --model-path, --model=, positional `vllm serve <x>`, --revision), engine env (HF_MODEL_ID, MODEL_ID, MODEL_PATH, with $VAR expansion from the container's own env), download init containers (NGC CLI with NGC_MODEL_NAME, huggingface-cli download with --local-dir, aws s3 sync, KServe storage-initializer args), NIM images with NIM_MODEL_PROFILE, and for LeaderWorkerSet workers that name no model the leader template of their group. URIs: hf://org/repo[@rev], ngc://org/team/model:ver, s3://bucket/key, nim://image@profile, path:///abs for pre-filled paths. The landing volume is the mount at or above the download destination in the main container: emptyDir or rootfs are substitutable, a customer's PVC, hostPath or other volume means sharing is already solved and the pod is left alone. KServe pvc:// is likewise not a downloader. Tests are built from real specs: the NVCF Helm function observed on prd11 (NGC init download, positional MODEL_PATH), the Dynamo operator sample, SGLang, NIM, KServe, huggingface-cli and S3 inits, the upstream LWS vLLM example, and non-downloaders (Dynamo frontend, Ray worker, etcd). Mutation-checked: PVC marked substitutable, revision split dropped and env expansion dropped each turn tests red. Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
Step 2 of docs/proposals/helm-shared-model-volume.md. For a GPU pod whose model identity resolves (internal/modelid), the webhook replaces the volume the download lands in with the per-identity model volume and turns the download into a write-once step; no pod is gated. internal/modelvolume names the claim per identity (nvsnap-model-<key>), creates the writer claim (ReadWriteMany on a distributed filesystem, ReadWriteOnce on NVMesh), looks the identity up cluster-wide and records completion as a label. The webhook elects the writer with the existing Lease elector keyed by identity; the writer's landing emptyDir becomes the claim, its engine mounts the model read-only, and its download init is wrapped to touch <volume>/.nvsnap-complete on success. Readers on a distributed filesystem mount the same claim and their init waits for the marker, downloading themselves after the deadline (decided: always fall back). Readers on NVMesh keep their emptyDir, are labelled pending and wait for the agent to bind the completed volume in (step 3). Charts whose engine downloads itself get an injected huggingface-cli init that lands the model in the volume, with the engine started offline and the registry credentials forwarded. Compile caches are redirected into <volume>/.nvsnap/cache/<engine key> on a distributed filesystem and into the local cachedir on NVMesh; the model env of the cachedir template is left alone because the model lives in the landing volume now. Tests use the prd11 function shape and a stock vLLM Deployment: writer on Block, pending reader on Block, reader on RWX, injected init with forwarded token and offline engine, completed volume skipping the election, and the customer-PVC / no-GPU / no-model pods left alone. Mutation-checked: read-only mount dropped, reader fallback dropped, Block reader substituting the volume, and writer marker dropped each turn tests red. Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
…Mesh
Step 3 of docs/proposals/helm-shared-model-volume.md. The agent gains a
model volume controller. For writers, on any node: when the download init
the webhook named on the pod exits 0, the writer claim is labelled
complete and, on block storage, SharedVolumePromoter.MintReadOnly exposes
it as a read-only claim (primary PV retained, secondary static PV with the
namespace-rewritten NVMesh handle, pre-bound claim). The writer claim
stays because the writer still runs on it and later writers must find it
complete. For pending readers on this node: once the identity is complete
the read-only claim is minted in the reader's namespace, attached to the
node through a mount-holder pod, bind-mounted read-only onto the reader's
hostPath landing under the Bidirectional overlays root
(<overlay-root>/models/<key>), and the pod is un-pended. The writer touched
the marker inside the volume, so the reader's wait init sees it as soon as
the bind lands. Every agent watches writers (idempotent marking and
minting); only the reader's own agent binds.
The webhook gives Block-mode readers a hostPath landing with
HostToContainer propagation on the engine and init mounts, and names the
writer's download init for the agent. Storage profiles gain
`modelVolume: {mode, storageClass, size}`; block is the default for
shared-volume strategies, rwx must be declared. Agent flags
--model-volume and --model-volume-wait-deadline, Helm
agent.modelVolume.{enabled,waitDeadline}, and pods patch in the agent
ClusterRole.
Tests with fake clients: writer completion marks and mints the read-only
PV and claim with the handle rewritten and the primary retained; a
pending reader on its node gets the claim in its own namespace, one
attach, one bind at the hostPath, and is un-pended, a second reader of
the same identity reuses the bind, readers on other nodes are ignored.
Mutation-checked: node filter dropped, completion on any exit code,
binding before completion, and un-pend dropped each turn tests red.
Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
…b in huggingface_hub 1.x Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
The first cluster run of the model volume on NVMesh exposed a constraint the design had missed: a volume attached read-write by a running pod cannot be attached read-only anywhere else (AttachVolume failed on the read-only PV while the writer pod held the primary). The L2 promote never met it because it deletes the writer claim before readers attach. So the download step cannot live inside a pod that goes on to serve. The download is now a Job per identity per cluster, nvsnap-model-dl-<key>, created idempotently by the webhook on first sight: the chart's own download init (image, command, env, pull secrets, tolerations) wrapped to touch the marker, or `hf download` on the engine image when the engine fetches the model itself. Its exit releases the volume. Every workload pod is a reader; there is no writer pod and no Lease election on this path, because Job create is atomic. The agent completes the identity when the Job succeeds and mints the read-only claim on block storage; readers are bound as before. The design doc records the constraint and the reason. Tests: Job spec derived from the NGC init and from a stock vLLM pod (hf download, forwarded token, claim mounted at the init's path), idempotent create, no Job once the volume is complete, controller completes on Job success only. Mutation-checked: Job never created, Job created when complete, and completion on a running Job each turn tests red. Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
…detaches A Succeeded pod keeps its volumes attached. On NVMesh that let the read-only attach succeed only on the Job's own node (dev1 2026-09-26); every other node failed until the pod was deleted. The Job now carries ttlSecondsAfterFinished=30; completion state lives on the claim label. MarkComplete tolerates the update race between agents by re-reading. Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
…a dedicated host root Two findings from the second dev1 run of the model volume on NVMesh. The retained volume must have no read-write attachment before any node can attach it read-only: with the Job's claim still bound, the read-only attach succeeded only on the Job's own node and failed everywhere else. On block storage MarkComplete now labels the retained PV with identity and completion, sets Retain, and deletes the download claim, the same release the L2 promote performs; Lookup reads completion from the PV; readers' read-only claims are minted from that PV, and only after no VolumeAttachment references it (Detached). RWX mode keeps the shared claim and labels it as before. The bind root moves out of the overlays root: the overlay sweeper removes entries it does not own and deleted the model binds under it. Completed model volumes are bound under a dedicated Bidirectional hostPath (agent.hostPaths.nvsnapModels, default /var/lib/containerd/nvsnap-models), mounted at the same path in the agent and on the host. Tests: block completion releases the claim and labels the PV, RWX labels the claim, the reader is not served while a VolumeAttachment exists and is served once it is gone (mutation-checked: dropping the detach gate turns the test red), the read-only PV carries the reader namespace's handle. Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
…on dev1 Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
…d when stale The controller kept "bound" in memory. After the volume for an identity was replaced (and after a manual unmount during cleanup on dev1) it un-pended readers without a bind and they waited on an empty hostPath. The mount table is now the truth: a reader is served only when the device mounted at the bind target belongs to the identity's primary PV (NVMesh device csi-id against the volume handle); a missing mount is bound again, a mount of a replaced volume is unbound and redone, and a bind that leaves nothing mounted does not un-pend the pod. Tests: rebinding after an unmount, replacing a stale bind, device/handle matching. Mutation-checked: dropping the stale unbind and the missing mount check each turn tests red. Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. 🗂️ Base branches to auto review (1)
Please check the settings in the CodeRabbit UI or the ⚙️ Run configuration
You can disable this status message by setting the Use the checkbox below for a quick retry:
Comment |
Three gaps against the shape of a real NVCF Helm function (an NGC init download with a secrets file and a scripts ConfigMap, in a namespace with Kyverno enforced): - the download Job carries every volume the chart's init mounted, with only the landing volume redirected to the claim; before, the Job had only the landing mount and the download could not read its key - block readers default to a `pvc` reader mode: the pod references the read-only claim in its namespace, the webhook mints it at admission when the volume is complete, and any agent mints it after the Job when the primary detaches; no hostPath, so disallow-host-path passes. The former bind-in path is `readerMode: hostPath` in the storage profile, for gang-scheduled charts - Job pods get default requests and limits, a seccomp profile, dropped capabilities and no privilege escalation when the init set none Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
The engine-download Job container had requests without limits and the injected reader init had neither, so both failed require-pod-requests-limits where Kyverno is enforced (seen on dev1). One helper now gives every container the model volume creates the same baseline: requests and limits unless the chart sized it, no privilege escalation, dropped capabilities. Mount propagation on the injected init is set only in hostPath reader mode, the one case where the host binds a volume in after the pod started. Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
…iners The injected init and the download Job still failed require-run-as-non-root and the by-name CAP_NET_RAW check on dev1. The download runs the engine's image and must own what it writes, so both containers now inherit runAsNonRoot, runAsUser, runAsGroup and seccomp from the chart's main container, and the Job pod inherits the pod security context (fsGroup). NET_RAW is dropped by name next to ALL. Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
A capture pod gets an empty emptyDir at the cachedir root with
NIM_CACHE_PATH and HF_HOME pointed at <root>/model and the compile
caches at <root>/cache. vLLM and Hugging Face create their own trees,
so this went unnoticed; NIM checks NIM_CACHE_PATH at startup and exits
("Unable to read from NIM_CACHE_PATH", first NIM function on dev1 with
the NVCA integration on, 2026-09-28). An init container on the
workload's own image now creates both subtrees, sticky world-writable,
on capture pods and on the model volume's cache-only emptyDir. It
inherits the workload's user posture and is hardened like the other
containers nvsnap injects.
Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
The prewarm became a storage-profile switch on 2026-09-24 but the built-in NVMesh entry never used it, so every NVMesh restore still ran the sweep. It is neutral to negative there: a single reader already saturates the volume (70B TP=4: 247 s with, 257 s without) and the sweep reads the whole tree, unused files included. The first NVCA-driven NIM restore on dev1 spent 67 s reading 30 GB, half of it an original .pth nothing opens, to save a 2 s weights load, and finished no faster than a cold start. Hyperdisk ML keeps prewarm on; the ConfigMap overlay and NVSNAP_PREWARM still override either way. Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
The sweep fanned out per file with xargs -P 6 -n 16. A checkpoint tree is a dozen files, four of them the weights, so one process read the whole tree serially: 30 GB in 67 s on dev1 while the engine's own mmap load of 15 GB took 33 s, and the restore finished no faster than cold. Measured on that NVMesh volume, one stream reads 270 MB/s, four 1.0 GB/s, eight 2.1 GB/s. Files above 64 MiB are now split into 256 MiB ranges read by dd under xargs -P; small files are still batched through cat. The walk follows symlinks so a Hugging Face snapshot is read by entry name, skipping original/ (the .pth copy nothing opens) and the blobs directory (the same bytes again). The NVMesh profile keeps the prewarm on and uses eight readers; this supersedes the previous commit's off default, which was diagnosing the wrong cause. Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
A block-mode model volume was created at admission at a fixed 512Gi ceiling because nothing knew the model's size yet; on a RAID-10 NVMesh class a 1 GB model reserved about 1 TiB of raw NVMe. Registries do not all expose sizes and not every storage system can grow a volume, so the download Job now stages into a pod-local emptyDir, the same shape as the single-GPU cachedir capture. When it succeeds, the agent on that node measures the bytes on disk, creates a claim of measured size plus ten percent (whole GiB, profile modelVolume.minSize as the floor) in the nvsnap namespace, copies the tree in through a mount-holder, labels the retained PV complete, releases the claim and deletes the Job. The primary no longer belongs to a function namespace. RWX mode is unchanged. A reaper on every agent removes read-only model PVs whose claim is gone (Released, or bound in a deleted namespace) as objects only, and abandoned incomplete primaries with their storage. The agent gains namespaces/get so a lookup error can never read as a deleted namespace. Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
7927f5c to
65303c8
Compare
Every Helm pod on block storage compiled into its own cachedir emptyDir: 23.5 s of torch.compile per pod on a 0.5B model, more on larger ones, while a warm cache brings it to 2.9 s (measured 2026-09-29). The model volume provisioner now has a second kind, cache, with its own names and labels. The first pod of an engine configuration and ordinal that becomes Ready is captured by the agent on its node: its cache subtree is measured, copied into a claim of that size in the nvsnap namespace and the PV labelled complete. Later pods with the same key get a read-only claim minted at admission and a seed init copies it into their cachedir before the engine starts. Atomic claim creation picks the source among pods sharing a key; failures release the claim and, after the last attempt, record a failure so admissions stop waiting. Caches are keyed per ordinal, never merged: ranks of a tensor-parallel group write rank-specific directories and the shared autotune results differ in content between ranks. Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
…re source A pod whose engine downloads a model nvsnap has no recipe for was stamped for capture of its landing volume and returned from admission before the compile-cache injection, so the cold deployment of such a chart never collected a set and the first warm deployment profiled and compiled again (kimi-k3 on GB300, 2026-10-03: 4 min 38 s). The capture source now gets the cachedir and the set stamping like any reader; it still keeps the chart's own download step. Those pods write their download credentials under HOME, which is the collected root (the NGC CLI config holds the API key), so the set collector and the rank tar server now exclude credential paths, and a skipped directory is no longer descended by the tar writer. Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
A download Job on storage that is shared while written runs with the function's posture; a new filesystem is root-owned and 0755, so a uid-1000 engine image could not create its first directory in the primary and the Job crashed into its backoff limit (Qwen3-0.6B on an FSS primary, 2026-10-03). The agent only reacted to succeeded Jobs, so nothing was recorded and the readers waited out their deadline. The agent on the Job's node now opens the primary root through the Job pod's kubelet mount, the writer waits until its landing is writable, and a failed Job records a failure with the Job's reason and the download container's last words, retires the writer view and keeps the primary claim for the retry. Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
… job fails Readers of a primary that is shared while written hold views and wait on the completion marker; when the download Job gave up they waited out their deadline (30 min on the cluster) before downloading for themselves. The Job's node now writes a failure marker next to the completion marker with the Job's reason, the reader's wait init runs its fallback the moment it appears and logs the reason, and the next writer removes the marker before it downloads. A successful retry clears the failure record. Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
…d ref The base image was once built from a local CRIU checkout on a different line than the fork and ref versions.sh pins for the OSS build, so two builders produced two different CRIUs. The base build now fails when the resolved source is not at NVSNAP_CRIU_REF (override only with NVSNAP_CRIU_ALLOW_REF_MISMATCH=1), the app build warns when the last recorded base commit is not the pinned ref, and check-criu-ref reports both. Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
The base image is now built from github.com/balajinvda/criu at 74b4170aa (tag nvsnap-base-v0.0.21): the criu-v2 io_uring engine line plus the epoll shared-fd tolerance and the late-device-resume failure propagation ported from the earlier line. The earlier line, which the previous base was built from, is archived. The fork documentation names the fork of record and what the line carries. Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
…knows Upstream CRIU renamed --compress-region to --compress-block when it unified the page and region modes. The criu-v2 dump now probes the bundled criu once and uses whichever spelling it advertises, so the opt-in NVSNAP_CRIU_V2_COMPRESS setting keeps working across the base image that carries the rename; "block" is accepted alongside "region". Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
… criu that will not exit The criu-v2 restore of a model-volume reader failed on the first file it tried to reopen: engines keep JIT artefacts open or mapped under the webhook-injected cachedir (Triton launchers, the FlashInfer log), and the placeholder has no copy of that emptyDir. The cachedir mount now rides the same replay that already carries /dev/shm, so the placeholder gets the files at their original paths before CRIU runs. After that failure criu stayed alive waiting on its restorer tasks and the agent waited with it. The restore now runs criu in its own process group, kills the group on cancel, and cancels once the restore log has reported a failed restore and a grace period has passed. Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
9d781b5 to
6f705dc
Compare
…xtracting into it The cachedir replay extracts into /opt/nvsnap inside the placeholder's mount namespace, but a placeholder only has the mountpoints its own spec declares and tar refuses a missing directory. Create the mountpoint on the placeholder's rootfs before extracting; /dev/shm is unaffected since it always exists. Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com>
…dev) Base v0.0.22 is built from the official fork's criu-dev after the rebase onto upstream criu-dev bf186f060. Upstream renamed --compress-region to --compress-block; the agent already probes for whichever flag the bundled criu accepts. Validated end to end on x86 (vllm-small: checkpoint 1m16s, restore pod Ready 33 s, inference after restore OK). Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com> Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
…p/helm-artifacts Signed-off-by: Balaji Ganesan <bganesan@nvidia.com> # Conflicts: # src/compute-plane-services/nvsnap/internal/webhook/cachedir.go # src/compute-plane-services/nvsnap/internal/webhook/cachedir_noshim_test.go
FamousDirector
left a comment
There was a problem hiding this comment.
Reviewed the incremental change against declared base nvsnap/helm-election (7642bfc). Four P1 availability/data-integrity blockers are detailed inline: cross-namespace concurrent writers, releasing the writer claim after a failed completion update, fallback downloads into a read-only reader view, and writes escaping a partially extracted rank directory on retry.
Validation: existing modelid/modelvolume tests passed; webhook, hostlibs, and tarstream tests passed with a temporary Darwin portability overlay for unchanged treecopy code. Focused checkpointstore shared-volume/view and rootfsonly composer tests passed. Five isolated regression probes reproduced the four findings. Linux agent tests cross-compiled successfully; Helm lint and incremental diff whitespace checks passed. Full checkpointstore/rootfsonly filesystem tests could not run on macOS because the unchanged Linux sendfile path panics there. No live GPU/CSI/cluster validation was performed. All probes used temporary overlays; tracked files remain unchanged.
| // election. | ||
| func (p *Provisioner) EnsureDownloadJob(ctx context.Context, uri, ns, claim string, step DownloadStep) (string, error) { | ||
| name := JobName(uri) | ||
| if _, err := p.Kube.BatchV1().Jobs(ns).Get(ctx, name, metav1.GetOptions{}); err == nil { |
There was a problem hiding this comment.
[P1] Elect the download owner across namespaces
When two Helm functions in different namespaces first request the same incomplete model URI, each namespace creates its own Job here. ensureDownload reuses the cluster-wide primary but gives both Jobs read-write views of the same CSI handle. A fake-client admission probe produced two Jobs sharing that handle. Running the generated writer scripts concurrently also left the completion marker present while the slower writer truncated a model file: readers can start against partially overwritten weights after the first Job reports completion. This breaks both the once-per-cluster requirement and model integrity.
Acquire one atomic, cluster-wide owner for the model identity before creating the namespaced Job. Only that owner should download/write; other namespaces should remain readers, with ownership released or recovered on failure.
There was a problem hiding this comment.
Fixed in 8e09512. The webhook now claims a cluster-wide download owner before creating any Job: a Lease in the nvsnap namespace whose create is atomic, so concurrent admissions agree on one owning namespace. Other namespaces create no Job and no writer view and read the primary like any other reader. An owner whose Job is gone past a 2 minute grace period is retired with UID and resourceVersion preconditions and the download claimed again; a recorded failure or the completion releases the record. I reused the Lease-create primitive rather than LeaseElector itself, because nvsnap-server lists and reconciles the capture-election Leases. Tests: TestClaimDownload_OneOwnerAcrossNamespaces (concurrent claims from two namespaces, grace period, stale takeover, release on failure) and TestModelVolume_SharedFilesystem_OneWriterAcrossNamespaces (two namespaces admitted, one Job, no second writer view; it failed with 2 Jobs before the fix).
| pv.Annotations[k] = v | ||
| } | ||
| pv.Spec.PersistentVolumeReclaimPolicy = corev1.PersistentVolumeReclaimRetain | ||
| if _, err := p.Kube.CoreV1().PersistentVolumes().Update(ctx, pv, metav1.UpdateOptions{}); err != nil && !apierrors.IsConflict(err) { |
There was a problem hiding this comment.
[P1] Persist completion and Retain before releasing the writer claim
A PV resource-version conflict is ignored here, so the following PVC deletion runs even though neither the completion labels nor Retain were persisted. This can occur when a CSI controller or another agent updates the PV between Get and Update. With a dynamically provisioned block PV whose reclaim policy is Delete, releasing the last writer claim then allows the completed model volume to be destroyed. The injected-conflict probe confirmed this method returns nil and deletes the claim while the PV remains incomplete with reclaim policy Delete; the completed artifact is also no longer discoverable through Lookup.
Retry conflicts against a freshly fetched PV, and delete the writer claim only after the required completion metadata and Retain policy have successfully persisted, or a reread confirms them.
There was a problem hiding this comment.
Fixed in a1060fd. MarkCompleteClaim now retries PV update conflicts against a freshly read PV (bounded at 5 attempts) and deletes the writer claim only once the stored PV is labelled complete and set to Retain. Persistent conflicts return an error and leave the claim in place. Test: TestMarkComplete_ConflictKeepsClaimUntilPersisted, with one injected conflict (retried, then released) and permanent conflicts (error, claim kept, PV still Delete). Both cases failed before the fix.
| fallback = download | ||
| } | ||
| failed := shellQuote(path.Join(path.Dir(marker), modelvolume.FailedMarkerFile)) | ||
| return fmt.Sprintf("set -e\nd=0\nwhile [ ! -f %[1]s ]; do if [ -f %[4]s ]; then echo \"nvsnap: shared download failed: $(cat %[4]s 2>/dev/null)\"; %[3]s; exit 0; fi; if [ $d -ge %[2]d ]; then echo 'nvsnap: marker deadline passed; downloading locally'; %[3]s; exit 0; fi; sleep 5; d=$((d+5)); done\necho 'nvsnap: model complete, skipping download'\n", shellQuote(marker), deadline, fallback, failed) |
There was a problem hiding this comment.
[P1] Give failed readers a writable fallback landing
For PVC readers admitted while an RWX download is in flight, readerLanding replaces the original landing with a CSI view whose PV has CSI.ReadOnly=true. This script then runs the original download against that same landing when the failure marker appears or the wait deadline expires. Any fallback that creates model files fails on the read-only filesystem, leaving the function in an init-container restart loop rather than recovering. A generated-script probe with a failed marker and a non-writable landing confirmed the write fails. Recording the shared failure only bypasses mutation for later admissions; already admitted readers retain this mount.
Route fallback into a writable per-pod landing that the engine also consumes, or recreate the affected readers on their original download path after recording the failure. Reusing the read-only shared view cannot implement the advertised local fallback.
There was a problem hiding this comment.
Fixed in 2c36257, taking your second option. A writable per-pod landing would need the engine to read a second path, which this design cannot provide without rewriting the chart. PVC readers no longer fall back in place: on the failure marker or past the deadline their wait stops with the reason and a non-zero exit, instead of downloading into the read-only view. When the download Job fails, the agent on the Job node records the failure and deletes every waiting reader of that model that is not Ready. Their controllers recreate them, and admission sees the failure record and leaves them on their own download path. hostPath readers keep the in-place fallback because their landing is writable until the agent binds. Tests: TestReaderScript_NoInPlaceFallbackOnReadOnlyView, TestModelVolume_SharedFilesystem_ReaderHasNoInPlaceFallback and TestModelVolumeController_FailedSharedDownloadRecreatesWaitingReaders. The design doc failure table is updated to match. One limit remains: a writer that hangs without failing still leaves readers waiting, as before; the Job failing is what releases them.
| mode = 0o777 | ||
| } | ||
| } | ||
| f, err := os.OpenFile(target, os.O_CREATE|os.O_TRUNC|os.O_WRONLY, mode|0o600) |
There was a problem hiding this comment.
[P1] Confine extraction writes when retrying partial rank streams
fetchRank retries into the same destination after an interrupted response, retaining symlinks from the first extraction. A later regular-file entry uses os.OpenFile here, which follows those existing links. The lexical link check below does not account for intermediate symlink resolution, so writes can escape the rank directory and truncate other files. A temporary-directory probe using two archives emitted by the actual Write function reproduced this: an interrupted first stream left relative links; after the source entry changed to a regular file, the retry overwrote a sibling file outside the extraction destination. Cache collection must not corrupt neighboring cache or agent filesystem data when a transfer retries.
Use filesystem operations confined to the destination root for all extraction paths, rejecting escaping symlink resolution. Extract retries into fresh staging directories and publish only a complete tree.
There was a problem hiding this comment.
Fixed in b0ea0ee. Extract now refuses any entry beneath a symlinked directory, replaces a symlink at an entry path instead of following it, and opens files with O_NOFOLLOW. The rank fetch now uses the new ExtractFresh: each attempt extracts into a fresh staging directory next to the destination and replaces it only when the whole stream extracted, so a retry never merges into what an interrupted attempt left. Tests: TestExtract_RefusesWritesThroughSymlinkedParents (your d -> . then d/f -> ../victim chain), TestExtract_ReplacesSymlinkAtFilePath (both overwrote the victim before the fix) and TestExtractFresh_PublishesOnlyCompleteTrees.
The Bazel BUILD files in nvsnap were not regenerated as packages and files changed on this stacked branch; check-gazelle only runs on pull requests to main, so the drift was not reported. Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com> Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
Carry main's unqualified-storage change up the stack. The model volume is built after the L2-disabled return; it already required a matched profile and the shared-volume promoter, so behavior is unchanged. Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com> Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
…default cache An engine that sets neither HF_HOME nor HF_HUB_CACHE reads the Hugging Face default under HOME, and nvsnap mounts the model volume at that default path. The model-volume path also moves HOME into the pod-local cachedir, so the engine looked for the model under the cachedir and, with HF_HUB_OFFLINE set, exited with LocalEntryNotFoundError. Set HF_HOME to the landing when the container names no cache location, by value or by reference, and move HF_MODULES_CACHE off the read-only volume in that case as well. Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com> Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
…ined MarkCompleteClaim ignored a conflicting PV update and deleted the writer claim anyway. When a CSI controller or another agent wrote the PV between the read and the update, neither the completion labels nor the Retain policy were stored, and releasing the last claim of a Delete-policy volume could destroy the completed model. Retry conflicts against a freshly read PV, and release the claim only once the stored PV is complete and retained. Persistent conflicts return an error and keep the claim. Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com> Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
Extract checked each symlink lexically but followed links when writing, so links that resolve through one another, or a link left by an interrupted transfer, let a later entry write outside the destination. The rank fetch retried into the same directory, keeping those links. Refuse entries under a symlinked directory, replace a symlink at an entry's own path instead of following it, and open files with O_NOFOLLOW. Fetch each rank attempt into a fresh staging directory and publish it only when the whole stream extracted. Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com> Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
…tion stack (#2261) Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
Two Helm functions in different namespaces admitting the same incomplete model each created a download Job, and on a shared filesystem both wrote the one primary through their own read-write views. The completion marker could appear while the slower writer was still truncating a file. Record the downloading namespace in a Lease in the nvsnap namespace; the create is atomic, so concurrent admissions agree on one owner. Other namespaces create no Job and no writer view and read the primary like any reader. An owner whose Job is gone past a grace period is retired with preconditions and the download claimed again; a recorded failure or the completion releases the record. Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com> Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
A PVC reader admitted while a shared-filesystem download runs mounts the read-only view of the primary at its landing. On the failure marker or past its deadline its wait init ran the original download into that view, which cannot be written, so the pod restarted its init forever. PVC readers no longer fall back in place: the wait stops with the reason. When the download Job fails, the agent on the Job's node records the failure and deletes every waiting reader that is not Ready; their controllers recreate them and admission, seeing the failure record, leaves them on their own download. hostPath readers keep the in-place fallback, since their landing is writable until the agent binds. Co-Authored-By: Balaji Ganesan <bganesan@nvidia.com> Signed-off-by: Balaji Ganesan <bganesan@nvidia.com>
…ap/helm-artifacts
Why
Customers deploy Helm charts of N identical GPU workers, and every worker
downloads and compiles on its own. The gate-and-promote election in #2104
cannot serve multi-node instances (LWS, StatefulSet groups, Dynamo
podgangs): holding any member deadlocks the group. The invariant we want
is simpler and storage-independent: the model is downloaded once per
cluster and every pod attaches it; compile caches are shared the same
way. Design and the decisions taken with the product owner:
docs/proposals/helm-shared-model-volume.md.What changed
One model-volume lifecycle on every storage. The primary volume lives in
the nvsnap namespace; completion is a labelled, retained PV and a released
claim; each function namespace gets a per-namespace view of it (read-only
for readers, read-write for the download Job). The storage mode decides
only two things: the claim's access mode and how the download reaches
the primary (block: staged into a pod-local emptyDir and copied by the
agent, sized from the measured bytes; shared filesystem: the Job writes
through its namespace's read-write view while readers already hold their
views). Admission never waits for a volume to bind.
internal/modelid: identity and landing volume for any chart. Identityas a URI (
hf://,ngc://,s3://,nim://) from engine args(including positional
vllm serve <x>and$VARexpansion), engineenv, download init containers (NGC CLI,
hf download, S3, KServestorage-initializer), NIM images, and LeaderWorkerSet workers via the
leader template. Landing volume kind decides substitutability; a
customer's shared PVC, hostPath or image is left alone.
internal/modelvolume+ webhook: one download Job per identity percluster (
nvsnap-model-dl-<key>, idempotent create, so no election),built from the chart's own download init wrapped to touch a marker, or
from
hf downloadon the engine image with credentials forwarded.Every workload pod is a reader: its download init becomes a wait for
the marker (deadline, then self-download), the engine is started
offline when it used to download itself. Charts whose engine script
fetches an artifact nvsnap has no recipe for (an NGC model pulled in
the main container) keep their download; the agent on that node
captures the finished tree after Ready and later pods read the copy.
Every container the feature creates gets requests and limits, dropped
capabilities, no privilege escalation and the chart's own posture, so
enforced Kyverno baselines pass.
the tree size in its termination message (kubelet unmounts a finished
Job pod's volumes at once), the staged copy and the capture measure the
same number, and the agent stamps it on the primary; a 62 GB model on a
4Ti nominal filesystem claim gets 62Gi views.
first Ready group is collected (startup plans, kernel and inductor
caches, the transformers modules cache), later pods seed their local
cachedir from it. Compile caches stay in a local cachedir on every
storage; the model view is read-only. Capture-source pods collect a set
too, so the first deployment hands the second one its startup plan.
Credential paths under HOME (NGC CLI config, Hub tokens, and the like)
never travel in a set.
root is root-owned and the Job runs as the function's user (uid 1000
for Dynamo and NIM images), so the agent on the Job's node opens the
primary root through the Job pod's kubelet mount and the writer waits
until its landing is writable. A failed Job is recorded with its reason
and the download container's last words, the writer view is retired,
the primary claim is kept for the retry, and a failure marker next to
the completion marker sends readers already waiting to their fallback
at once instead of at their deadline. A successful retry clears the
record.
agent.modelVolume.reapInterval, default 10m)for model and cache volumes alike: unreferenced views are retired with
their holders, view PVs whose claim is gone are removed, incomplete
primaries released for 15 min are removed with their storage; complete
primaries are kept.
nvsnap.io/cache-saltfolds an operator-chosen value into the captureand cache identity for changes the identity cannot see (a wheel
installed at start, a floating tag re-pointed).
endpoints, chart component, build targets); no current path used it.
agent.modelVolume.{enabled,waitDeadline,reapInterval}, storageprofile
modelVolume: {mode, storageClass, size, minSize, readerMode}(block by default on shared-volume strategies,
rwxmust be declared),agent.runtimewith the runtime socket mounted at its host path forCRI-O, a fully qualified smoke-hook image for short-name enforcement,
and smoke-hook and rollout timeouts that fit a serial agent roll on
large clusters. The agent build fails loudly when the base image is
missing from the target registry.
Customer Release Notes
Helm chart functions download each model once per cluster; every other
worker, later version, scale-up and redeploy attaches the downloaded
volume, and compile caches are shared between pods of one configuration.
No chart change. Requires NVMesh or a shared filesystem class such as
OCI File Storage.
Plan Summary
Helm: new values (off by default), one extra Bidirectional hostPath on
the agent DaemonSet, ClusterRole verbs (PVC update/patch, pods patch,
namespaces get, PV create/update/delete). Primaries and the agents' copy
holders live in the release namespace; views live in function
namespaces. No new controllers or CRDs.
Usage
Testing
NVMesh (block mode), stock
vllm-workerschart, Qwen2.5-32B TP=4, tworeplicas: first deploy one Job, 65 GB downloaded once, both Ready +591 s
with 0 downloads in the pods; reinstall both Ready +136 s; cold 325 s.
PVC reader mode, Qwen2.5-14B: first deploy both Ready +344 s, reinstall
+118 s. Real NVCF Helm function with the chart's own NGC download
(954 MB): claim sized at 2Gi from the measured bytes instead of 512Gi,
both Ready +230 s; compile-cache seeding took torch.compile from 23 s to
3 s per pod.
OCI File Storage (shared-filesystem mode), GB300 arm64 nodes, CRI-O,
NVCF storage class, 2026-10-02 and 2026-10-03:
writer view, cache set collected after Ready); warm 1 min 38 s;
scaling 1 to 4 replicas on fresh nodes, each Ready 72 to 78 s after
creation reading weights at about 2.3 GB/s per node from one
filesystem, no downloads, no attach wait.
Ready (image pulls excluded): cold admission to Ready 46 min 30 s
(download 31 min, capture into the primary 24 min in the background);
second deployment, unseeded, 20 min 42 s; third deployment, seeded,
12 min 20 s (weights 425 s from the filesystem view, profiling 0 s
with the saved plan, graph capture 80 s, no restarts). Against the
same chart on NVMesh: 14 min 21 s warm.
filesystem primary: this is the run that found the root-permission and
failed-Job gaps fixed above; the fixes are covered by unit tests and
the function redeploy after the next agent roll is the e2e check.
Unit: identity forms and classifier; Job spec derivation and
idempotence; webhook matrix (block and shared-filesystem readers,
complete-at-admission minting, the same flow asserted on every storage,
plain pods keep every annotation, Job volume carry-over, hardening and
posture inheritance, capture sources, injected download, customer
volumes left alone, writer and reader scripts); controller completion,
release, detach gate, per-namespace view minting, bytes recording from
every fill path, capture from an engine download, root opening on the
Job's node, failed Job handling and the failure marker, cache-set
collection with credential exclusion, reaper rules.
go test ./...green.
Notes
Stacked on #2104 (base
nvsnap/helm-election). Measured on the way andrecorded in the design doc: the CuTeDSL warmup cannot be cached from
outside the engine (
cute.compile()forces no cache); CUDA graph captureis not cacheable; on both NVMesh and NFS the weight read sits at about
2 GB/s per node against 4.4 GB/s from local NVMe, which points at a
node-local tier under the shared primary as the next step. The first
start of a multi-worker engine can race in transformers' dynamic-module
copy; seeded starts avoid it because the modules cache arrives populated.
References
Relates to #2099
Related Pull Requests
#2104, #2101
Dependencies
None