Brief
Amazon SageMaker HyperPod adds model caching to cut inference cold starts
Amazon SageMaker HyperPod now supports model caching for inference, pre-loading model weights and container images onto cluster nodes so pods read from local NVMe storage instead of downloading over the network.

The change targets cold starts, which the source says can drop from tens of minutes to seconds. Model caching pre-loads model weights and container images onto cluster nodes. Pods then read from local NVMe storage rather than downloading over the network. The source also explains how to enable it.
Source details and supporting facts
Each line is stated by the page named above it.
Stated by aws.amazon.com
- Amazon SageMaker HyperPod now supports model caching for inference.
- Model caching pre-loads model weights and container images onto cluster nodes.
- Pods read from local NVMe storage instead of downloading over the network.
- Model caching cuts cold starts from tens of minutes to seconds.
Sources
- AWS Machine Learning BlogText stored 15 September 2026
How this story was checked. Written from the 1 page listed above, stored 15 September 2026; claims checked against that stored text on 15 September 2026.
What that means
- 4 of 4 reported statements were confirmed against the page that carries them; the rest were removed rather than published.
- Figures in the text were required to appear in the stored source text: yes. Identifiers: yes.
- The check reads stored text only: no claim rests on a fresh look that did not happen.
- Where the reporting was silent, the text says so instead of filling the gap.