{"id":7,"date":"2026-09-28T22:43:48","date_gmt":"2026-09-28T20:43:48","guid":{"rendered":"https:\/\/linuxundich.de\/en\/?p=7"},"modified":"2026-10-03T16:58:37","modified_gmt":"2026-10-03T14:58:37","slug":"llmfit-which-ai-models-run-on-my-machine","status":"publish","type":"post","link":"https:\/\/linuxundich.de\/en\/gnu-linux\/llmfit-which-ai-models-run-on-my-machine\/","title":{"rendered":"llmfit: Which AI models run on my machine?"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">Anyone who wants to try local AI will eventually run into a pretty mundane question: which model actually runs on my machine? The selection of open-source LLMs is huge by now, and the candidates differ quite a bit in size, quantization, and memory requirements. If you go digging through model cards and gigabytes of GGUF files, that&#8217;s a whole evening gone.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/www.llmfit.org\/\">llmfit<\/a> takes this research off your hands. The open-source tool reads out your hardware and compares it against a database of language models. What you get is a pretty concrete answer to what can actually run sensibly on your own machine.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Which models fit my hardware?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">After launching, llmfit first detects your hardware: CPU, RAM, graphics card and available VRAM. It then compares those resources against the models it supports and sorts the list by a score.<\/p>\n\n\n\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6ac177021a24e&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6ac177021a24e\" class=\"wp-block-image wp-lightbox-container\"><img decoding=\"async\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/linuxundich.de\/wp-content\/uploads\/2026\/09\/llmfit-arch1.webp\" alt=\"\"\/><button\n\t\t\tclass=\"lightbox-trigger\"\n\t\t\ttype=\"button\"\n\t\t\taria-haspopup=\"dialog\"\n\t\t\tdata-wp-bind--aria-label=\"state.thisImage.triggerButtonAriaLabel\"\n\t\t\tdata-wp-init=\"callbacks.initTriggerButton\"\n\t\t\tdata-wp-on--click=\"actions.showLightbox\"\n\t\t\tdata-wp-style--right=\"state.thisImage.buttonRight\"\n\t\t\tdata-wp-style--top=\"state.thisImage.buttonTop\"\n\t\t>\n\t\t\t<svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"12\" height=\"12\" fill=\"none\" viewBox=\"0 0 12 12\">\n\t\t\t\t<path fill=\"#fff\" d=\"M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z\" \/>\n\t\t\t<\/svg>\n\t\t<\/button><figcaption class=\"wp-element-caption\">Terminal window showing llmfit&#8217;s model list on a laptop with a GeForce GTX 1050, sorted by score, with columns for parameters, tokens per second, quantization and memory requirements.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">This is useful mainly because raw model size isn&#8217;t the only thing that counts. A model can theoretically fit into the available memory and still run painfully slowly in practice. So llmfit estimates how well a model suits your hardware and what speed and memory requirements you can expect.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Just how much this differs from machine to machine shows up with two devices from my stash. On the laptop with a dedicated GTX 1050, the 4 GB of VRAM is a hard limit. On a computer with Iris Xe graphics, on the other hand, the CPU and the GPU share the same memory, and suddenly almost everything says \u201cPerfect\u201d.<\/p>\n\n\n\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6ac177021a85a&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6ac177021a85a\" class=\"wp-block-image wp-lightbox-container\"><img decoding=\"async\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/linuxundich.de\/wp-content\/uploads\/2026\/09\/llmfit-arch2.webp\" alt=\"\"\/><button\n\t\t\tclass=\"lightbox-trigger\"\n\t\t\ttype=\"button\"\n\t\t\taria-haspopup=\"dialog\"\n\t\t\tdata-wp-bind--aria-label=\"state.thisImage.triggerButtonAriaLabel\"\n\t\t\tdata-wp-init=\"callbacks.initTriggerButton\"\n\t\t\tdata-wp-on--click=\"actions.showLightbox\"\n\t\t\tdata-wp-style--right=\"state.thisImage.buttonRight\"\n\t\t\tdata-wp-style--top=\"state.thisImage.buttonTop\"\n\t\t>\n\t\t\t<svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"12\" height=\"12\" fill=\"none\" viewBox=\"0 0 12 12\">\n\t\t\t\t<path fill=\"#fff\" d=\"M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z\" \/>\n\t\t\t<\/svg>\n\t\t<\/button><figcaption class=\"wp-element-caption\">llmfit&#8217;s model list on a system with integrated Intel Iris Xe graphics, where almost all entries in the Fit column are rated Perfect.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">The program also takes different quantizations into account. With locally run models, you quickly run into terms like Q4, Q5 or Q8. Put simply, this reduces the precision of the model weights to save memory and computing power. Depending on the level, you have to accept some loss in quality in return.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This matters most for GGUF models, which are commonly used for local inference. llmfit factors the different quantization levels into its rating, so it can also compare models that come in several variants.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">If you&#8217;ve been using tools like <a href=\"https:\/\/linuxundich.de\/software\/mit-inxi-und-i-nex-informationen-zur-hardware-des-rechners-ausgeben\/\">Inxi<\/a> on Linux to get an overview of your hardware, you already know the principle. llmfit goes one step further, though, and puts the collected data directly in relation to possible AI models.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Installation on Arch Linux<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Update September 29, 2026:<\/strong> The wheel of time keeps turning, and on Arch it famously turns a little faster. When I started writing this article, llmfit was only available for Arch as an AUR package. Because of the turbulence and attacks on the AUR, I try to avoid it these days. In the meantime, the llmfit package has made it into the official Arch repositories, so all you need to install it is: <code>sudo pacman -S llmfit<\/code><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">On Arch Linux, uv is available directly from the official repositories. That makes the installation pleasantly unspectacular:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>sudo pacman -S uv<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">After that, you can install llmfit as a tool with <code>uv<\/code>:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>uv tool install -U llmfit<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Now <code>llmfit<\/code> is available right away as a command.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">If you just want to try the program first, you can use <code>uvx<\/code> instead. That runs llmfit without installing it permanently as a tool:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>uvx llmfit<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">So on Arch Linux, you don&#8217;t need a separate installer for uv.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Installation on Ubuntu<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">On Ubuntu, the route is a bit different. Depending on your Ubuntu version and the repositories you use, you can&#8217;t simply install uv with <code>apt<\/code>. The installer provided by the uv developers is therefore a straightforward option.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">With <code>curl<\/code> it works like this:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>curl -LsSf https:\/\/astral.sh\/uv\/install.sh | sh<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Alternatively, you can download and run the installation script with <code>wget<\/code>:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>wget -qO- https:\/\/astral.sh\/uv\/install.sh | sh<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">After that, first check whether <code>uv<\/code> is available:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>uv --version<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">From here on, installing llmfit works exactly like on Arch Linux:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>uv tool install -U llmfit<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">If you just want to try it out without a permanent installation, uvx is available on Ubuntu as well:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>uvx llmfit<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">If you&#8217;d rather not use <code>uv<\/code>, you can download the prebuilt binaries from the llmfit GitHub Releases instead.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Getting rid of llmfit again<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">If the tool doesn&#8217;t convince you after all, the way back is short, but on Arch there&#8217;s a catch. The tool itself goes away with<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>uv tool uninstall llmfit<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">and this tells you whether anything is left over:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>uv tool list<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">If you only tried llmfit with uvx, you never installed anything permanently in the first place. The downloaded packages still sit in uv&#8217;s cache, though. You empty it with<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>uv cache clean<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Now for the catch: if you installed uv through the package manager, you&#8217;re tempted to simply type<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>sudo pacman -Rns uv<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">But pacman only removes its own package. What uv itself created in your home directory is unknown to the package manager, and without uv you can no longer run <code>uv tool uninstall<\/code> either. So uninstall the tool first and <code>uv<\/code> second, not the other way around.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">These commands show you where the leftovers are:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>uv tool dir\nuv cache dir<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Usually that&#8217;s <code>~\/.local\/share\/uv\/tools<\/code> and <code>~\/.cache\/uv<\/code>. If uv is already gone and llmfit is still around, the only option is to clean up by hand:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>rm -r ~\/.local\/share\/uv<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Be careful here: this also wipes out all tools installed with uv and the Python versions managed by uv, not just llmfit.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">On Ubuntu, where uv comes from the install script, there is no <code>uv self uninstall<\/code>. There, after the <code>uv tool uninstall<\/code>, you remove the two launcher binaries yourself:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>rm ~\/.local\/bin\/uv ~\/.local\/bin\/uvx<\/code><\/pre>\n\n\n\n<h2 class=\"wp-block-heading\">The TUI is more than a simple model list<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The terminal interface is one of the more interesting parts of llmfit to me. Instead of just printing a list, it lets you move through the results, search, and pull up more information on each model.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Use <code>\/<\/code> to search for a specific model, the arrow keys or <code>j<\/code> and <code>k<\/code> to navigate, and <code>h<\/code> opens the help. With almost 10,000 entries in the list, the search is no luxury.<\/p>\n\n\n\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6ac177021c165&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6ac177021c165\" class=\"wp-block-image wp-lightbox-container\"><img decoding=\"async\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/linuxundich.de\/wp-content\/uploads\/2026\/09\/llmfit-txos.webp\" alt=\"\"\/><button\n\t\t\tclass=\"lightbox-trigger\"\n\t\t\ttype=\"button\"\n\t\t\taria-haspopup=\"dialog\"\n\t\t\tdata-wp-bind--aria-label=\"state.thisImage.triggerButtonAriaLabel\"\n\t\t\tdata-wp-init=\"callbacks.initTriggerButton\"\n\t\t\tdata-wp-on--click=\"actions.showLightbox\"\n\t\t\tdata-wp-style--right=\"state.thisImage.buttonRight\"\n\t\t\tdata-wp-style--top=\"state.thisImage.buttonTop\"\n\t\t>\n\t\t\t<svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"12\" height=\"12\" fill=\"none\" viewBox=\"0 0 12 12\">\n\t\t\t\t<path fill=\"#fff\" d=\"M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z\" \/>\n\t\t\t<\/svg>\n\t\t<\/button><figcaption class=\"wp-element-caption\">llmfit in the terminal on TUXEDO OS, the model list shows values between 2.9 and 65.7 in the tok\/s column, with the search running at the bottom.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">The built-in benchmark features are interesting, too. <code>b<\/code> brings up community benchmarks, while <code>I<\/code> starts an inference benchmark on your own hardware. That slowly turns a pure estimate into a measurement with real numbers.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">It reminds me a bit of classic system tools. If you use <a href=\"https:\/\/linuxundich.de\/software\/mission-center-1-0-der-systemmonitor-fuer-gnome-geht-in-die-vollen\/\">Mission Center<\/a> on GNOME, for example, you get CPU, RAM, GPU and other resources presented graphically. llmfit, on the other hand, looks specifically at the question of what you can do with those resources in terms of local AI.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Estimates instead of marketing promises<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Of course, you shouldn&#8217;t mistake llmfit&#8217;s numbers for a real benchmark. A theoretical calculation can only estimate how fast a certain model might run on a certain piece of hardware.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The hardware simulation is quite nice in this context. You can tweak RAM, VRAM, and core count just to see how the rating shifts. If you&#8217;re toying with the idea of buying more memory or a new graphics card, you at least get a ballpark figure.<\/p>\n\n\n\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6ac177021c7a7&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6ac177021c7a7\" class=\"wp-block-image wp-lightbox-container\"><img decoding=\"async\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/linuxundich.de\/wp-content\/uploads\/2026\/09\/llmfit-arch2-details.webp\" alt=\"\"\/><button\n\t\t\tclass=\"lightbox-trigger\"\n\t\t\ttype=\"button\"\n\t\t\taria-haspopup=\"dialog\"\n\t\t\tdata-wp-bind--aria-label=\"state.thisImage.triggerButtonAriaLabel\"\n\t\t\tdata-wp-init=\"callbacks.initTriggerButton\"\n\t\t\tdata-wp-on--click=\"actions.showLightbox\"\n\t\t\tdata-wp-style--right=\"state.thisImage.buttonRight\"\n\t\t\tdata-wp-style--top=\"state.thisImage.buttonTop\"\n\t\t>\n\t\t\t<svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"12\" height=\"12\" fill=\"none\" viewBox=\"0 0 12 12\">\n\t\t\t\t<path fill=\"#fff\" d=\"M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z\" \/>\n\t\t\t<\/svg>\n\t\t<\/button><figcaption class=\"wp-element-caption\">llmfit with the Hardware Simulation dialog open, where you can change RAM, VRAM, and the number of CPU cores.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">That&#8217;s exactly why the benchmarks that are now built in are interesting. llmfit can download a model, start it via a supported runtime provider, and then measure the actual speed on your own hardware. The results are stored locally.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">If you like, you can then share your measurements with the community via<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>llmfit bench --share<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">The results are submitted as a GitHub pull request. This way, a database of real measurements slowly builds up, instead of just theoretical calculations.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">With local AI in particular, I find this approach sensible. After all, the statement \u201cfits into VRAM\u201d doesn&#8217;t tell you much about whether you can actually work with the model in a reasonable way.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Also usable as a command-line tool<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">If a TUI isn&#8217;t your thing, you don&#8217;t have to use llmfit interactively. The program works entirely from the command line, too.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For example,<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>llmfit recommend<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">gives you a recommendation for your hardware. With<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>llmfit recommend --json<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">you can output the results as JSON. That&#8217;s obviously a lot more interesting if you want to build llmfit into your own scripts or other tools.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">You can also start a web interface or an API:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>llmfit serve --host 0.0.0.0 --port 8787<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">There is also a ready-made image for containers:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>ghcr.io\/alexsjones\/llmfit<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">That means llmfit runs on a server or in an existing Docker setup, too. If you keep your containers up to date regularly, you may already know the drill from <a href=\"https:\/\/linuxundich.de\/software\/docker-container-einfach-aktuell-halten-mit-dockcheck\/\">Dockcheck<\/a>.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">And what do you run the model with?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">One distinction you should keep in mind: llmfit is not an AI runtime itself. The program helps you choose and evaluate a model, while other tools do the actual inference.<\/p>\n\n\n\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6ac177021d3b2&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6ac177021d3b2\" class=\"wp-block-image wp-lightbox-container\"><img decoding=\"async\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/linuxundich.de\/wp-content\/uploads\/2026\/09\/llmfit-arch2-sim.webp\" alt=\"\"\/><button\n\t\t\tclass=\"lightbox-trigger\"\n\t\t\ttype=\"button\"\n\t\t\taria-haspopup=\"dialog\"\n\t\t\tdata-wp-bind--aria-label=\"state.thisImage.triggerButtonAriaLabel\"\n\t\t\tdata-wp-init=\"callbacks.initTriggerButton\"\n\t\t\tdata-wp-on--click=\"actions.showLightbox\"\n\t\t\tdata-wp-style--right=\"state.thisImage.buttonRight\"\n\t\t\tdata-wp-style--top=\"state.thisImage.buttonTop\"\n\t\t>\n\t\t\t<svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"12\" height=\"12\" fill=\"none\" viewBox=\"0 0 12 12\">\n\t\t\t\t<path fill=\"#fff\" d=\"M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z\" \/>\n\t\t\t<\/svg>\n\t\t<\/button><figcaption class=\"wp-element-caption\">Detail view in llmfit for DeepSeek-R1-Distill-Qwen-7B with a score breakdown, memory requirements per quantization level, and a GGUF download hint from bartowski.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Supported tools include Ollama, llama.cpp, MLX, Docker Model Runner, and LM Studio. That means llmfit can, for example, benchmark a model through an existing Ollama or llama.cpp stack.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">On Linux, <a href=\"https:\/\/ollama.com\/\">Ollama<\/a> is certainly one of the easier ways to get local models up and running. If you want more control over the inference itself, you quickly end up with <a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\/\">llama.cpp<\/a>. The two projects take somewhat different approaches, but both combine well with llmfit&#8217;s hardware check.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Interesting for local AI<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">To me, llmfit closes a gap between \u201cI&#8217;d like to try local AI\u201d and \u201cI first have to work my way through the hardware requirements of dozens of models.\u201d That&#8217;s far more helpful than a generic list of the \u201cbest\u201d local models. In the end, what matters isn&#8217;t which model scores well on some test system, but which one runs properly on your hardware for your use case.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Of course, the estimates don&#8217;t replace your own testing. That&#8217;s exactly what the built-in benchmarks are for, and they&#8217;re what I want to tackle next: my own measurements on different machines, and then a comparison of how well llmfit&#8217;s predictions match reality.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The tool should be especially interesting for developers. If you want to use a local AI assistant for programming, you first find out which models fit your hardware, and then experiment with the speeds you can actually reach.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">You&#8217;ll find more information, the current documentation, and the source code on <a href=\"https:\/\/github.com\/AlexsJones\/llmfit\">GitHub<\/a>.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Want to run a local AI but not sure which model really fits your hardware? llmfit reads out your specs, scores thousands of options, and estimates speed and VRAM requirements. This open-source tool takes the guesswork out and gives you concrete recommendations. In short: no more agonizing over quantization, you just know what runs.<\/p>\n","protected":false},"author":2,"featured_media":8,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"lui_source_id":47801,"lui_source_hash":"06857770d0924aeb2d8c4199a6f7464eac1949af9285f2d54015b39bd02783d4","lui_source_translated":"2026-10-03","lui_source_reviewed":true,"lui_via_url":"","lui_via_label":"","lui_source_url":"","lui_source_label":"","footnotes":""},"categories":[2],"tags":[],"class_list":["post-7","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-gnu-linux"],"_links":{"self":[{"href":"https:\/\/linuxundich.de\/en\/wp-json\/wp\/v2\/posts\/7","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/linuxundich.de\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/linuxundich.de\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/linuxundich.de\/en\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/linuxundich.de\/en\/wp-json\/wp\/v2\/comments?post=7"}],"version-history":[{"count":1,"href":"https:\/\/linuxundich.de\/en\/wp-json\/wp\/v2\/posts\/7\/revisions"}],"predecessor-version":[{"id":9,"href":"https:\/\/linuxundich.de\/en\/wp-json\/wp\/v2\/posts\/7\/revisions\/9"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/linuxundich.de\/en\/wp-json\/wp\/v2\/media\/8"}],"wp:attachment":[{"href":"https:\/\/linuxundich.de\/en\/wp-json\/wp\/v2\/media?parent=7"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/linuxundich.de\/en\/wp-json\/wp\/v2\/categories?post=7"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/linuxundich.de\/en\/wp-json\/wp\/v2\/tags?post=7"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}