Brief
Ollama's new scheduler targets out-of-memory crashes and GPU utilization
Ollama has rolled out a new model scheduling system that it says reduces out-of-memory crashes and improves GPU performance, especially on multi-GPU machines.

Ollama says the new scheduling system reduces crashes caused by running out of memory. It also aims to maximize GPU utilization and performance, with the biggest benefits on multi-GPU systems.
For users, this could mean smoother operation when loading or switching between models, particularly on workstations with multiple GPUs.
Source details and supporting facts
Each line is stated by the page named above it.
Stated by ollama.com
- Ollama now includes a significantly improved model scheduling system
- The new system reduces crashes due to out of memory issues
- The new system maximizes GPU utilization and performance, especially on multi-GPU systems
Sources
- Ollama blogText stored 15 September 2026
How this story was checked. Written from the 1 page listed above, stored 15 September 2026; claims checked against that stored text on 15 September 2026.
What that means
- 3 of 3 reported statements were confirmed against the page that carries them; the rest were removed rather than published.
- Figures in the text were required to appear in the stored source text: yes. Identifiers: yes.
- The check reads stored text only: no claim rests on a fresh look that did not happen.
- Where the reporting was silent, the text says so instead of filling the gap.