Nodes to run Hunyuan Image 3 locally with BF16 and NF4 quantized options in Comfyui
-
Updated
Apr 30, 2026 - Python
Nodes to run Hunyuan Image 3 locally with BF16 and NF4 quantized options in Comfyui
My personal project about 3D Gaussian Splatting (3DGS) training
Stateless LLM runtime that dynamically routes, loads, executes, and unloads models per request with bounded VRAM caching and intelligent model selection.
ComfyUI custom node that controls the order of node execution with linear routing of any data type through infinite I/O slots + option to free VRAM & RAM at any point in a workflow with device-agnostic memory management utilities managed by ComfyUI that safely unload all models, while preserving all connected data & models through to the next node.
KeSSie HUGE Context Semantic recall for Large Language Models
🦖 Godzilla: Next-gen LLM inference engine featuring 512K context scaling, Six Flags over Texas stack, TriAttention, TurboQuant, DFlash draft sidecars, and --kv-vram-only allocation.
BoneMemory: Universal Async Core for AI-Toolkit. The hardware-agnostic VRAM manager for every model: Image, Video, and Audio. Whether you’re training a Rank 16 LoRA on an entry-level GPU or pushing Rank 1024 on an RTX 4090, BoneMemory eliminates memory bottlenecks. Architecture that makes any hardware punch above its weight. Zero OOM.
Constant-memory sequence modeling engine combining selective holographic-compression (ASH-C) with a coordinate pointer network (HEP-DNA). Bypasses the linear KV Cache bottleneck on consumer GPUs.
Predictive VRAM Virtualization Engine
LEMA (Layer-wise Efficient Memory Abstraction): A hardware-aware framework for fine-tuning LLMs in VRAM-constrained environments using asynchronous binary pre-fetching and triple-tier memory orchestration.
Sticky-block topology lottery scheduler for transformer fine-tuning.
Private, autonomous AI lab — turn your PC into a local-first intelligence platform without melting your GPU. Python + FastAPI + local LLMs.
Perkunas AI Training Platform is a memory-aware model training and serving system for serious language model experimentation under tight hardware limits. It combines streaming training, rich telemetry, guarded recovery, checkpoint export, and OpenAI-compatible serving.
Ultra-Low Bit KV-Cache Compression optimization layer built on top of llama.cpp for LLM inference. Reduces VRAM overhead by ~75-80% using custom CUDA kernels.
NMOS (Neural Memory OS) is a predictive partial execution engine enabling 70B-level reasoning on 4GB VRAM. It uses the “Zero-Lag” hypothesis, leveraging typing latency as a compute window to mask memory limits via async layer prefetching and speculative decoding.
Reducing reserved memory on NVIDIA GPUs
INT8 Sparse Tensor Core GEMM for PyTorch — built for Windows
Know before you train — VRAM estimation for LLM fine-tuning.
Adaptive dual-tier serving for Gemma 4 on consumer 16GB GPUs. Complexity + real-time VRAM routing between vLLM E4B and llama.cpp 27B. Production stack with OpenWebUI, monitoring, and more.
Reinforcement Learning based mathematical theorem prover with differentiable VRAM tokenizer optimization
Add a description, image, and links to the vram-optimization topic page so that developers can more easily learn about it.
To associate your repository with the vram-optimization topic, visit your repo's landing page and select "manage topics."