mtop

Elad Sertshuk·eladser.mtop

htop for your local AI

A terminal dashboard for local AI servers (Ollama, llama.cpp, LM Studio, vLLM). Shows loaded models and their VRAM, the GPU, and every request with its tok/s. Unloads models that overstay.

winget install --id eladser.mtop --exact --source winget

Latest 1.3.0·June 18, 2026

Release Notes

Watch more than one box, see more per request, and a couple of new numbers.

  • Multi-host: give -ollama a comma list and mtop stacks the models and GPUs from each machine, tagged by host. Handy if you run models on a couple of boxes.
  • GPU util and memory now draw as sparklines over time, next to the live numbers.
  • Request inspector: run with -inspect, press i, and you get the last request's prompt, completion, and a load/prompt/decode timing split. Off by default; the text it captures is stripped of control bytes so a model can't smuggle escape sequences into your terminal.
  • Session energy on the TOK/S line: watt-hours used and tokens per watt-hour. It's whole-GPU power, so read it as a rough efficiency number.
  • compare -openai runs the comparison against llama.cpp, LM Studio or vLLM, not just ollama.
  • -mem-alert and -temp-alert to set the alert thresholds instead of the built-in 93% and 87C. brew and scoop pick this up as usual; winget follows once Microsoft merges the bump.

Installer type: zip

x64D58ED2A2F12F9272F47A58872150DBF3F2B37165D2E9865CB5A586562EA95AA5

Details

Homepage
https://github.com/eladser/mtop
License
MIT
Publisher
Elad Sertshuk
Support
https://github.com/eladser/mtop/issues

Tags

gpullamacpplocal-llmollamatui

Older versions (1)

1.2.0
x64498feff1f6aa8914df8a27ff19ee8118851390de2d8392b20e459447a00f93f5