Fine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.
-
Updated
Oct 7, 2026 - Python
Fine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.
The most efficient one-page LoRA trainer for Anima 2B. Optimized for 6GB+ VRAM, featuring a smart dataset analyzer and real-time previews.
One-click Windows installer for Z-Image Turbo AI image generation. Optimized for low-VRAM GPUs (4GB+). Features Gradio web UI, automatic setup, and GGUF model support.
WeeLLM runs large diffusion models with as little as 4 GB of VRAM, without any quantization. It dynamically determines how many layers can fit within the available VRAM and streams the text encoder and transformer layers to the GPU layer by layer, enabling inference on hardware with limited VRAM. It supports both safetensors and GGUF models.
A ComfyUI Workflow for low vram users
Train LoRAs for Qwen-Image 2.1, FLUX.2 Klein 9B, Krea 2, Z-Image, Ideogram 4, Anima, SDXL (Pony, Illustrious, NoobAI), LTX 2.3 and MiniMax-H3 on NVIDIA GPUs from 4–8 GB VRAM. One install, one launcher, NF4 models. Windows, Linux and RunPod.
Hierarchical RAG architecture scaling to 693K chunks on consumer hardware (4GB VRAM). Features 3-address routing, hybrid vector+graph fusion, and SetFit classification.
"Adaptive Hybrid Quantization Framework for deploying 7B+ LLMs on low-VRAM devices (e.g., GTX 1050). Features surgical block alignment and Numba-accelerated inference.
Taiwanese Hokkien (Taigi) speech-to-text transcriber - MediaTek Breeze-ASR-26 with faster-whisper, tuned for RTX 3050 4GB low-VRAM GPUs. Gradio UI, CLI, Docker, SRT/VTT/TXT/JSON.
MiniMax H3 video generation on 2080Ti 22G (Turing sm_75) - field-tested handbook, workflows & FAQ
Three MiniMax H3 ComfyUI workflows for 8GB laptop GPUs (20 / 8 / 4 steps). Includes the model list, launch flags, prompt-structure pitfalls, the resolution trap, and measured evidence for why you must restart ComfyUI before every run.
VRAM Optimization 2026 gaming toolkit for GPU memory tracking, texture profiles, VRAM usage analysis, stutter troubleshooting and performance benchmarks.
SCAIL-2 (Wan 2.1) Low-VRAM motion transfer — endless videos from a single reference image. 8+ GB GPUs.
llama.cpp fork tuned for running modern models (Gemma-4, Qwen3.x) at full context on 12 GB Turing GPUs (RTX 2060/2070/2080, T4). TurboQuant KV cache (KTQ+VTQ, 2.78 bpw f16-quality), SWA-aware KV, MTP+n-gram speculation.
ComfyUI deployment kit for Qwen-Image-2.1 Uncensored (GGUF): ready-to-use t2i & image-edit workflows, ComfyUI-GGUF architecture patches, and a 14-second lossless Q8_0 quantization repair tool that fixes 'weight shape [136] vs normalized_shape [128]'. Runs on 8 GB VRAM. Weights NOT included.
Native Hunyuan3D 2.1 Full extension for Modly, optimized for low-VRAM NVIDIA GPUs with INT8, FP8 and FP16 support.
Modly extension for Hunyuan3D 2.1 patched for Windows AMD and modest PCs
Contains the notebooks and workflows configured to run inference from Wan 2.2 Animate with ComfyUI on Kaggle T4 GPUs smoothly
Run Qwen-Image-2.1 on an 8GB GPU — one-command ComfyUI + GGUF setup, workflows and VRAM doctor.
To associate your repository with the low-vram topic, visit your repo's landing page and select "manage topics."