Real, verified pull commands for small local AI models — every use case, organized by the hardware you actually have. Nothing hosted here, every link goes straight to the real source (Ollama's registry or Hugging Face).
A curated list of small local AI models covering every common AI use case — chat,
code, reasoning, images, video, voice, translation, and more — that run on modest
consumer hardware. Every command below is real and verified: Ollama tags checked
against Ollama's live registry, LocalAI ids checked against LocalAI's live model
gallery (checked 2026-08-02). No files are hosted here — these commands download
directly from Ollama's registry or Hugging Face.
How to read this: organized by System RAM needed (8/12/16/24 GB), then by
VRAM needed (1/2/3/4 GB) within each. If your machine has more RAM/VRAM than a
given section, everything in the smaller sections below it will run on your machine
too — not just the section matching your exact specs.
The all-purpose assistant type — ask it questions, have a conversation, get help writing something. This is what most people mean by "chatting with AI."
ollama pull qwen2.5:0.5b (~379 MB)local-ai run qwen3-0.6b (~462 MB) — Qwen3 0.6B replaces Qwen2.5 0.5B — newer generation, same tiny-chat role, already verified in galleryhttps://huggingface.co/MaziyarPanahi/Qwen3-0.6B-GGUF/resolve/main/Qwen3-0.6B.Q4_K_M.ggufollama pull qwen3:0.6b (you already have this)local-ai run qwen3-0.6b (~462 MB)https://huggingface.co/MaziyarPanahi/Qwen3-0.6B-GGUF/resolve/main/Qwen3-0.6B.Q4_K_M.ggufollama pull smollm2:360m (you already have this)local-ai run qwen3-0.6b (~462 MB) — no SmolLM2-360M in gallery; nearest ultra-tiny chat modelhttps://huggingface.co/MaziyarPanahi/Qwen3-0.6B-GGUF/resolve/main/Qwen3-0.6B.Q4_K_M.ggufollama pull qwen2:0.5b (~336 MB)local-ai run qwen3-0.6b (~462 MB) — Qwen2 (old gen) not in gallery; Qwen3 0.6B is the current tiny generalisthttps://huggingface.co/MaziyarPanahi/Qwen3-0.6B-GGUF/resolve/main/Qwen3-0.6B.Q4_K_M.ggufollama pull gemma3:1b (you already have this)local-ai run gemma-3-1b-it (~769 MB)https://huggingface.co/ggml-org/gemma-3-1b-it-GGUF/resolve/main/gemma-3-1b-it-Q4_K_M.ggufollama pull llama3.2:1b (you already have this)local-ai run llama-3.2-1b-instruct:q4_k_m (~770 MB)https://huggingface.co/hugging-quants/Llama-3.2-1B-Instruct-Q4_K_M-GGUF/resolve/main/llama-3.2-1b-instruct-q4_k_m.ggufollama pull tinyllama (you already have this)local-ai run qwen3-0.6b (~462 MB) — TinyLlama not in gallery; Qwen3 0.6B is the modern equivalent-tier chat modelhttps://huggingface.co/MaziyarPanahi/Qwen3-0.6B-GGUF/resolve/main/Qwen3-0.6B.Q4_K_M.ggufWrites and explains programming code. Good for a developer who wants suggestions, bug fixes, or a starting point — not a replacement for understanding what the code does.
ollama pull qwen2.5-coder:0.5b (~379 MB)local-ai run opencoder-1.5b-base (~1.3 GB) — no 0.5B coder in gallery; OpenCoder 1.5B is the smallest real code modelhttps://huggingface.co/QuantFactory/OpenCoder-1.5B-Base-GGUF/resolve/main/OpenCoder-1.5B-Base.Q4_K_M.ggufTurns written text into spoken audio — a synthetic voice reading your text out loud.
local-ai run piper-en_US-lessac-medium-crispasr (~30 MB)https://huggingface.co/LocalAI-Community/piper-voices-GGUF/resolve/main/piper-en_US-lessac-medium-f16.ggufDoesn't chat at all — turns text into a mathematical "fingerprint" used for search. This is the engine behind "search my own documents" features (RAG).
ollama pull all-minilm (~44 MB)local-ai run all-MiniLM-L6-v2 (~86 MB)ollama pull bge-m3 (~1.1 GB)local-ai run bge-m3-colbert (~2.1 GB)ollama pull nomic-embed-text (you already have this)local-ai run nomic-embed-text-v1.5 (~261 MB)https://huggingface.co/mradermacher/nomic-embed-text-v1.5-GGUF/resolve/main/nomic-embed-text-v1.5.f16.ggufThe all-purpose assistant type — ask it questions, have a conversation, get help writing something. This is what most people mean by "chatting with AI."
ollama pull qwen2.5:1.5b (you already have this)local-ai run qwen3-1.7b (~1.2 GB) — Qwen3 1.7B replaces Qwen2.5 1.5B, same tier, already verifiedhttps://huggingface.co/MaziyarPanahi/Qwen3-1.7B-GGUF/resolve/main/Qwen3-1.7B.Q4_K_M.ggufollama pull qwen3:1.7b (you already have this)local-ai run qwen3-1.7b (~1.2 GB)https://huggingface.co/MaziyarPanahi/Qwen3-1.7B-GGUF/resolve/main/Qwen3-1.7B.Q4_K_M.ggufollama pull qwen2:1.5b (~892 MB)local-ai run qwen3-1.7b (~1.2 GB) — old-gen Qwen2 not in gallery; Qwen3 1.7B is the current equivalenthttps://huggingface.co/MaziyarPanahi/Qwen3-1.7B-GGUF/resolve/main/Qwen3-1.7B.Q4_K_M.ggufollama pull smollm2:1.7b (you already have this)local-ai run smollm2-1.7b-instruct (~1007 MB)https://huggingface.co/HuggingFaceTB/SmolLM2-1.7B-Instruct-GGUF/resolve/main/smollm2-1.7b-instruct-q4_k_m.ggufWrites and explains programming code. Good for a developer who wants suggestions, bug fixes, or a starting point — not a replacement for understanding what the code does.
ollama pull deepseek-coder:1.3b (~740 MB)ollama pull qwen2.5-coder:1.5b (you already have this)local-ai run opencoder-1.5b-instruct (~1.3 GB) — real, current small code-instruct modelhttps://huggingface.co/QuantFactory/OpenCoder-1.5B-Instruct-GGUF/resolve/main/OpenCoder-1.5B-Instruct.Q4_K_M.ggufBuilt to "think through" harder problems step by step — math, logic puzzles, multi-step decisions — instead of just pattern-matching an answer. Often slower because it works through the problem first.
ollama pull deepseek-r1:1.5b (you already have this)local-ai run deepseek-r1-distill-qwen-1.5b (~1.0 GB)https://huggingface.co/bartowski/DeepSeek-R1-Distill-Qwen-1.5B-GGUF/resolve/main/DeepSeek-R1-Distill-Qwen-1.5B-Q4_K_M.ggufCan look at an image you give it and describe, answer questions about, or read text from it. Doesn't generate new images — just understands existing ones.
ollama pull moondream (~1.6 GB)local-ai run moondream2 (~3.5 GB)https://huggingface.co/moondream/moondream2-gguf/resolve/main/moondream2-text-model-f16.ggufhttps://huggingface.co/moondream/moondream2-gguf/resolve/main/moondream2-mmproj-f16.ggufTurns spoken audio (a recording, podcast, video) into written text — transcription.
local-ai run whisper-small (~465 MB)https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-small.binTranslates text (or sometimes speech) from one language to another. Some are conversational generalists that translate okay; these are dedicated specialists.
ollama pull translategemma (~3.1 GB) — substituted for 'madlad400:3b' (doesn't exist in Ollama's registry)Condenses a long document into a short summary. Dedicated summarizers do this more reliably than asking a general chat model to "summarize this."
local-ai run llama-chat-summary-3.2-3b (~1.9 GB) — BART-Large-CNN not in gallery; Llama-Chat-Summary 3.2 3B is a real, current dedicated summarization finetunehttps://huggingface.co/bartowski/Llama-Chat-Summary-3.2-3B-GGUF/resolve/main/Llama-Chat-Summary-3.2-3B-Q4_K_M.gguflocal-ai run llama-chat-summary-3.2-3b (~1.9 GB) — Pegasus-X not in gallery; same real summarization substitute as BARThttps://huggingface.co/bartowski/Llama-Chat-Summary-3.2-3B-GGUF/resolve/main/Llama-Chat-Summary-3.2-3B-Q4_K_M.ggufThe all-purpose assistant type — ask it questions, have a conversation, get help writing something. This is what most people mean by "chatting with AI."
ollama pull gemma2:2b (you already have this)ollama pull ministral-3 (~5.6 GB)local-ai run mistralai_ministral-3-3b-instruct-2512-multimodal (~3.6 GB) — real 2512 release, already appliedhttps://huggingface.co/unsloth/Ministral-3-3B-Instruct-2512-GGUF/resolve/main/Ministral-3-3B-Instruct-2512-Q4_K_M.ggufhttps://huggingface.co/unsloth/Ministral-3-3B-Instruct-2512-GGUF/resolve/main/mmproj-F32.ggufollama pull llama3.2 (you already have this)local-ai run llama-3.2-3b-instruct-uncensored (~2.1 GB)https://huggingface.co/bartowski/Llama-3.2-3B-Instruct-uncensored-GGUF/resolve/main/Llama-3.2-3B-Instruct-uncensored-Q4_K_M.ggufollama pull qwen2.5:3b (you already have this)local-ai run falcon3-3b-instruct (~1.9 GB) — plain Qwen2.5 3B not in gallery; Falcon3 3B is a real, current 3B-class generalisthttps://huggingface.co/bartowski/Falcon3-3B-Instruct-GGUF/resolve/main/Falcon3-3B-Instruct-Q4_K_M.ggufWrites and explains programming code. Good for a developer who wants suggestions, bug fixes, or a starting point — not a replacement for understanding what the code does.
ollama pull stable-code:3b (~1.5 GB)local-ai run opencoder-1.5b-instruct (~1.3 GB) — Stable Code 3B not in gallery; OpenCoder covers the same lightweight-completion rolehttps://huggingface.co/QuantFactory/OpenCoder-1.5B-Instruct-GGUF/resolve/main/OpenCoder-1.5B-Instruct.Q4_K_M.ggufollama pull qwen2.5-coder:3b (you already have this)local-ai run opencoder-1.5b-instruct (~1.3 GB) — closest real small code model in the current galleryhttps://huggingface.co/QuantFactory/OpenCoder-1.5B-Instruct-GGUF/resolve/main/OpenCoder-1.5B-Instruct.Q4_K_M.ggufollama pull starcoder2:3b (you already have this)local-ai run opencoder-1.5b-instruct (~1.3 GB) — StarCoder2 not in gallery at any size; OpenCoder is the real small-code substitutehttps://huggingface.co/QuantFactory/OpenCoder-1.5B-Instruct-GGUF/resolve/main/OpenCoder-1.5B-Instruct.Q4_K_M.ggufCreates brand-new images from a text description ("a cat wearing a top hat"). The opposite of Vision Analysis.
local-ai run sd-1.5-ggml (~1.5 GB) — LCM fast-inference variant not in gallery; sd-1.5-ggml is the real small SD1.5-family image generatorhttps://huggingface.co/second-state/stable-diffusion-v1-5-GGUF/resolve/main/stable-diffusion-v1-5-pruned-emaonly-Q4_0.ggufGenerates non-music sound effects or ambient audio from a text description.
local-ai run acestep-cpp-turbo-4b (~7.6 GB) — AudioLDM (sound effects from text) not in gallery; AceStep is a real music/audio generation model — not a perfect substitute (music vs. arbitrary SFX) but the closest real audio-gen entryhttps://huggingface.co/Serveurperso/ACE-Step-1.5-GGUF/resolve/main/acestep-5Hz-lm-4B-Q8_0.ggufhttps://huggingface.co/Serveurperso/ACE-Step-1.5-GGUF/resolve/main/Qwen3-Embedding-0.6B-Q8_0.ggufhttps://huggingface.co/Serveurperso/ACE-Step-1.5-GGUF/resolve/main/acestep-v15-turbo-Q8_0.ggufhttps://huggingface.co/Serveurperso/ACE-Step-1.5-GGUF/resolve/main/vae-BF16.ggufText-to-speech that mimics a specific person's voice, given a short sample of them speaking.
local-ai run qwen3-tts-cpp-1.7b-base (~2.2 GB) — XTTS-v2 not in gallery; Qwen3-TTS is a real current voice-cloning-capable TTS modelhttps://huggingface.co/Serveurperso/Qwen3-TTS-GGUF/resolve/main/qwen-talker-1.7b-base-Q8_0.ggufhttps://huggingface.co/Serveurperso/Qwen3-TTS-GGUF/resolve/main/qwen-tokenizer-12hz-Q8_0.ggufThe all-purpose assistant type — ask it questions, have a conversation, get help writing something. This is what most people mean by "chatting with AI."
ollama pull gemma3:4b (you already have this)local-ai run gemma-3-4b-it (~3.1 GB)https://huggingface.co/lmstudio-community/gemma-3-4b-it-GGUF/resolve/main/gemma-3-4b-it-Q4_K_M.ggufhttps://huggingface.co/lmstudio-community/gemma-3-4b-it-GGUF/resolve/main/mmproj-model-f16.ggufollama pull qwen3:4b-instruct (~2.3 GB)local-ai run qwen3-4b (~2.3 GB)https://huggingface.co/MaziyarPanahi/Qwen3-4B-GGUF/resolve/main/Qwen3-4B.Q4_K_M.ggufBuilt to "think through" harder problems step by step — math, logic puzzles, multi-step decisions — instead of just pattern-matching an answer. Often slower because it works through the problem first.
ollama pull phi3 (~2.0 GB)local-ai run qwen3-4b (~2.3 GB) — Phi-3-mini not in gallery; Qwen3 4B covers the same reasoning-chat tierhttps://huggingface.co/MaziyarPanahi/Qwen3-4B-GGUF/resolve/main/Qwen3-4B.Q4_K_M.ggufollama pull phi4-mini (~2.3 GB)local-ai run qwen3-4b (~2.3 GB) — Phi-4-mini not in gallery; Qwen3 4B is the real current equivalenthttps://huggingface.co/MaziyarPanahi/Qwen3-4B-GGUF/resolve/main/Qwen3-4B.Q4_K_M.ggufCreates brand-new images from a text description ("a cat wearing a top hat"). The opposite of Vision Analysis.
local-ai run dreamshaper (~2.0 GB) — OpenJourney (Midjourney-style) not in gallery; Dreamshaper is a real stylized SD1.5-family alternativehttps://huggingface.co/Lykon/DreamShaper/resolve/main/DreamShaper_8_pruned.safetensorslocal-ai run sd-1.5-ggml (~1.5 GB) — real, current small SD1.5 gguf buildhttps://huggingface.co/second-state/stable-diffusion-v1-5-GGUF/resolve/main/stable-diffusion-v1-5-pruned-emaonly-Q4_0.ggufThe all-purpose assistant type — ask it questions, have a conversation, get help writing something. This is what most people mean by "chatting with AI."
ollama pull llama3.1 (~4.6 GB)local-ai run fusechat-llama-3.1-8b-instruct (~4.6 GB)https://huggingface.co/bartowski/FuseChat-Llama-3.1-8B-Instruct-GGUF/resolve/main/FuseChat-Llama-3.1-8B-Instruct-Q4_K_M.ggufollama pull mistral (you already have this)local-ai run mistral-7b-instruct-v0.3 (~4.1 GB)https://huggingface.co/MaziyarPanahi/Mistral-7B-Instruct-v0.3-GGUF/resolve/main/Mistral-7B-Instruct-v0.3.Q4_K_M.ggufollama pull neural-chat (~3.8 GB)local-ai run falcon3-3b-instruct (~1.9 GB) — Neural Chat 7B (old Intel finetune) not in gallery; Falcon3 3B covers the same instruction-following chat role at a smaller footprinthttps://huggingface.co/bartowski/Falcon3-3B-Instruct-GGUF/resolve/main/Falcon3-3B-Instruct-Q4_K_M.ggufollama pull mistral (you already have this)local-ai run mistral-7b-instruct-v0.3 (~4.1 GB)https://huggingface.co/MaziyarPanahi/Mistral-7B-Instruct-v0.3-GGUF/resolve/main/Mistral-7B-Instruct-v0.3.Q4_K_M.ggufWrites and explains programming code. Good for a developer who wants suggestions, bug fixes, or a starting point — not a replacement for understanding what the code does.
ollama pull codellama (~3.6 GB)local-ai run codellama-7b (~3.6 GB)https://huggingface.co/TheBloke/CodeLlama-7B-GGUF/resolve/main/codellama-7b.Q4_0.ggufollama pull starcoder2:7b (you already have this)local-ai run deepseek-coder-v2-lite-instruct (~9.7 GB) — StarCoder2 not in gallery; DeepSeek-Coder-V2-Lite is the real larger code substitutehttps://huggingface.co/LoneStriker/DeepSeek-Coder-V2-Lite-Instruct-GGUF/resolve/main/DeepSeek-Coder-V2-Lite-Instruct-Q4_K_M.ggufBuilt to "think through" harder problems step by step — math, logic puzzles, multi-step decisions — instead of just pattern-matching an answer. Often slower because it works through the problem first.
ollama pull deepseek-r1:7b (~4.4 GB)local-ai run deepseek-r1-distill-qwen-7b (~4.4 GB)https://huggingface.co/bartowski/DeepSeek-R1-Distill-Qwen-7B-GGUF/resolve/main/DeepSeek-R1-Distill-Qwen-7B-Q4_K_M.ggufollama pull deepseek-r1:8b (~4.9 GB)local-ai run deepseek-r1-distill-llama-8b (~4.6 GB)https://huggingface.co/unsloth/DeepSeek-R1-Distill-Llama-8B-GGUF/resolve/main/DeepSeek-R1-Distill-Llama-8B-Q4_K_M.ggufollama pull deepseek-r1:8b (~4.9 GB)local-ai run deepseek-r1-distill-llama-8b (~4.6 GB)https://huggingface.co/unsloth/DeepSeek-R1-Distill-Llama-8B-GGUF/resolve/main/DeepSeek-R1-Distill-Llama-8B-Q4_K_M.ggufTurns written text into spoken audio — a synthetic voice reading your text out loud.
local-ai run piper-en_US-lessac-medium-crispasr (~30 MB) — Bark not in gallery; Piper is the real, fast TTS already verified — less emotive but real and workinghttps://huggingface.co/LocalAI-Community/piper-voices-GGUF/resolve/main/piper-en_US-lessac-medium-f16.ggufTranslates text (or sometimes speech) from one language to another. Some are conversational generalists that translate okay; these are dedicated specialists.
ollama pull aya (~4.5 GB)local-ai run aya-23-8b (~4.7 GB)https://huggingface.co/bartowski/aya-23-8B-GGUF/resolve/main/aya-23-8B-Q4_K_M.ggufollama pull translategemma (~3.1 GB) — substituted for 'towerinstruct' (doesn't exist in Ollama's registry)Tuned for fiction, roleplay, and narrative writing rather than factual accuracy or following instructions strictly.
ollama pull llama3.1 (~4.6 GB)local-ai run fusechat-llama-3.1-8b-instruct (~4.6 GB)https://huggingface.co/bartowski/FuseChat-Llama-3.1-8B-Instruct-GGUF/resolve/main/FuseChat-Llama-3.1-8B-Instruct-Q4_K_M.ggufollama pull nous-hermes2 (~5.7 GB)local-ai run falcon3-3b-instruct (~1.9 GB) — Nous-Hermes-2 not in gallery; Falcon3 3B for general conversational storytellinghttps://huggingface.co/bartowski/Falcon3-3B-Instruct-GGUF/resolve/main/Falcon3-3B-Instruct-Q4_K_M.ggufThe all-purpose assistant type — ask it questions, have a conversation, get help writing something. This is what most people mean by "chatting with AI."
ollama pull gemma2:9b (~5.1 GB)ollama pull qwen2.5:7b (~4.4 GB)local-ai run qwen3-8b (~4.7 GB) — Qwen3 8B replaces Qwen2.5 7B, already verifiedhttps://huggingface.co/MaziyarPanahi/Qwen3-8B-GGUF/resolve/main/Qwen3-8B.Q4_K_M.ggufollama pull qwen2 (~4.1 GB)local-ai run qwen3-8b (~4.7 GB) — old-gen plain Qwen2 not in gallery at all; Qwen3 8B is the current equivalent-tier generalisthttps://huggingface.co/MaziyarPanahi/Qwen3-8B-GGUF/resolve/main/Qwen3-8B.Q4_K_M.ggufWrites and explains programming code. Good for a developer who wants suggestions, bug fixes, or a starting point — not a replacement for understanding what the code does.
ollama pull deepseek-coder:6.7b (~3.6 GB)ollama pull deepseek-coder-v2 (~8.3 GB)local-ai run deepseek-coder-v2-lite-instruct (~9.7 GB)https://huggingface.co/LoneStriker/DeepSeek-Coder-V2-Lite-Instruct-GGUF/resolve/main/DeepSeek-Coder-V2-Lite-Instruct-Q4_K_M.ggufollama pull phind-codellama (~17.7 GB)local-ai run opencoder-1.5b-instruct (~1.3 GB) — Phind-CodeLlama not in gallery; OpenCoder for lightweight code Q&Ahttps://huggingface.co/QuantFactory/OpenCoder-1.5B-Instruct-GGUF/resolve/main/OpenCoder-1.5B-Instruct.Q4_K_M.ggufollama pull phind-codellama (~17.7 GB)local-ai run opencoder-1.5b-instruct (~1.3 GB) — Phind-CodeLlama not in gallery; OpenCoder for lightweight code Q&Ahttps://huggingface.co/QuantFactory/OpenCoder-1.5B-Instruct-GGUF/resolve/main/OpenCoder-1.5B-Instruct.Q4_K_M.ggufollama pull qwen2.5-coder:7b (you already have this)local-ai run deepseek-coder-v2-lite-instruct (~9.7 GB) — real larger code model already verified elsewhere in this dochttps://huggingface.co/LoneStriker/DeepSeek-Coder-V2-Lite-Instruct-GGUF/resolve/main/DeepSeek-Coder-V2-Lite-Instruct-Q4_K_M.ggufBuilt to "think through" harder problems step by step — math, logic puzzles, multi-step decisions — instead of just pattern-matching an answer. Often slower because it works through the problem first.
ollama pull qwen2.5:7b (~4.4 GB)local-ai run qwen3-8b (~4.7 GB) — Qwen3 8B replaces Qwen2.5 7B, already verifiedhttps://huggingface.co/MaziyarPanahi/Qwen3-8B-GGUF/resolve/main/Qwen3-8B.Q4_K_M.ggufollama pull qwen3:8b (~4.9 GB)local-ai run qwen3-8b (~4.7 GB)https://huggingface.co/MaziyarPanahi/Qwen3-8B-GGUF/resolve/main/Qwen3-8B.Q4_K_M.ggufCan look at an image you give it and describe, answer questions about, or read text from it. Doesn't generate new images — just understands existing ones.
ollama pull bakllava (~4.4 GB)local-ai run qwen3-vl-2b-instruct (~1.8 GB) — BakLLaVA not in gallery; Qwen3-VL 2B is a real, current small vision modelhttps://huggingface.co/unsloth/Qwen3-VL-2B-Instruct-GGUF/resolve/main/Qwen3-VL-2B-Instruct-Q4_K_M.ggufhttps://huggingface.co/unsloth/Qwen3-VL-2B-Instruct-GGUF/resolve/main/mmproj-F16.ggufollama pull llava (~4.4 GB)local-ai run qwen3-vl-2b-instruct (~1.8 GB) — Llava 1.5 not in gallery; Qwen3-VL 2B covers image+text QAhttps://huggingface.co/unsloth/Qwen3-VL-2B-Instruct-GGUF/resolve/main/Qwen3-VL-2B-Instruct-Q4_K_M.ggufhttps://huggingface.co/unsloth/Qwen3-VL-2B-Instruct-GGUF/resolve/main/mmproj-F16.ggufollama pull qwen3-vl (~5.7 GB) — substituted for 'qwen-vl' (doesn't exist in Ollama's registry)local-ai run qwen3-vl-2b-instruct (~1.8 GB) — the old Qwen-VL-Chat naming isn't in gallery; Qwen3-VL is the current real version of the same familyhttps://huggingface.co/unsloth/Qwen3-VL-2B-Instruct-GGUF/resolve/main/Qwen3-VL-2B-Instruct-Q4_K_M.ggufhttps://huggingface.co/unsloth/Qwen3-VL-2B-Instruct-GGUF/resolve/main/mmproj-F16.ggufCreates brand-new images from a text description ("a cat wearing a top hat"). The opposite of Vision Analysis.
local-ai run sd-1.5-ggml (~1.5 GB) — SDXL Turbo not in gallery; sd-1.5-ggml is the real small image-gen model that fits 4GB VRAM (SDXL-class models generally don't)https://huggingface.co/second-state/stable-diffusion-v1-5-GGUF/resolve/main/stable-diffusion-v1-5-pruned-emaonly-Q4_0.ggufCreates instrumental music tracks from a text description.
local-ai run acestep-cpp-turbo-4b (~7.6 GB) — real, current music generation model in the galleryhttps://huggingface.co/Serveurperso/ACE-Step-1.5-GGUF/resolve/main/acestep-5Hz-lm-4B-Q8_0.ggufhttps://huggingface.co/Serveurperso/ACE-Step-1.5-GGUF/resolve/main/Qwen3-Embedding-0.6B-Q8_0.ggufhttps://huggingface.co/Serveurperso/ACE-Step-1.5-GGUF/resolve/main/acestep-v15-turbo-Q8_0.ggufhttps://huggingface.co/Serveurperso/ACE-Step-1.5-GGUF/resolve/main/vae-BF16.gguflocal-ai run acestep-cpp-turbo-4b (~7.6 GB) — real, current music generation model in the galleryhttps://huggingface.co/Serveurperso/ACE-Step-1.5-GGUF/resolve/main/acestep-5Hz-lm-4B-Q8_0.ggufhttps://huggingface.co/Serveurperso/ACE-Step-1.5-GGUF/resolve/main/Qwen3-Embedding-0.6B-Q8_0.ggufhttps://huggingface.co/Serveurperso/ACE-Step-1.5-GGUF/resolve/main/acestep-v15-turbo-Q8_0.ggufhttps://huggingface.co/Serveurperso/ACE-Step-1.5-GGUF/resolve/main/vae-BF16.ggufGood at pulling structured information (tables, JSON, forms) out of messy text, or formatting output to fit a specific structure.
ollama pull command-r (~17.4 GB)local-ai run command-r-v01:q1_s (~7.9 GB) — real quantized build, already appliedhttps://huggingface.co/dranger003/c4ai-command-r-v01-iMat.GGUF/resolve/main/ggml-c4ai-command-r-v01-iq1_s.ggufTuned for fiction, roleplay, and narrative writing rather than factual accuracy or following instructions strictly.
ollama pull dolphin-llama3 (~4.3 GB)local-ai run dolphin-2.9-llama3-8b (~4.6 GB)https://huggingface.co/cognitivecomputations/dolphin-2.9-llama3-8b-gguf/resolve/main/dolphin-2.9-llama3-8b-q4_K_M.ggufollama pull gemma2:9b (~5.1 GB)Creates short video clips from a text description. Still an early, slow, low-resolution capability on consumer hardware.
local-ai run wan-2.1-t2v-1.3b-ggml (~7.3 GB) — Zeroscope not in gallery; Wan2.1 1.3B is a real, current small text-to-video modelhttps://huggingface.co/calcuis/wan-gguf/resolve/main/wan2.1_t2v_1.3b-q8_0.ggufhttps://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/resolve/main/split_files/vae/wan_2.1_vae.safetensorshttps://huggingface.co/city96/umt5-xxl-encoder-gguf/resolve/main/umt5-xxl-encoder-Q8_0.ggufGood at pulling structured information (tables, JSON, forms) out of messy text, or formatting output to fit a specific structure.
ollama pull mistral-nemo (~6.6 GB)Tuned for fiction, roleplay, and narrative writing rather than factual accuracy or following instructions strictly.
ollama pull nous-hermes2 (~5.7 GB) — substituted for 'mythomax' (doesn't exist in Ollama's registry)local-ai run falcon3-3b-instruct (~1.9 GB) — MythoMax roleplay finetune not in gallery; Falcon3 3B for general storytelling/chat (won't match its uncensored roleplay style specifically)https://huggingface.co/bartowski/Falcon3-3B-Instruct-GGUF/resolve/main/Falcon3-3B-Instruct-Q4_K_M.gguf