Running AI locally?
Watch the whole Mac.
A local model shares resources with your other apps. If app switching becomes sluggish when a model loads, compare memory pressure and CPU activity before changing settings.
First, establish where the model runs
A local model runner performs inference on your computer. A hosted chat service usually performs inference remotely, while your browser handles the page. Local CPU or memory readings cannot explain a remote provider’s queue or server performance.
Compare idle and loaded states
- Open Activity Monitor’s Memory tab before loading a model and note the pressure graph.
- Load the model and repeat a typical prompt.
- Compare memory pressure and responsiveness, then inspect CPU activity.
- Unload the model or quit the runner normally and compare again.
WhySlow’s menu bar preview illustrates this sequence with sample data. Its expanded readings are planned for a future release. WhySlow does not measure GPU utilization, token throughput or inference quality.
Model size is only part of memory demand
The model, context window and simultaneous requests all matter. Ollama’s documentation explains that increasing context length increases memory requirements and that parallel requests can multiply the context allocation. Start with one model and one request, then increase the workload only if it remains comfortable.
If you use Ollama
Run ollama ps in Terminal to see loaded models and their processor allocation. If you no longer need a loaded model, ollama stop <model-name> unloads it; replace the placeholder with the name reported by ollama ps. Finish any active generation before stopping it.
Ollama also documents the keep_alive option for how long a model stays loaded. A smaller model or shorter context can reduce memory demand, but may change the result you get. Follow your runner’s documentation rather than copying settings intended for different hardware.
When the browser is the busy part
If you are using hosted AI, close unrelated heavy tabs after saving their state, check extensions and compare another browser. If local readings stay calm while responses arrive slowly, investigate the connection and service status. See the slow Mac checklist for separating these symptoms.
Sources and further reading
These guides explain common patterns. Confirm the cause on your own Mac before changing settings.