Skip to main content

Showing posts from Research category

Frozen 4-bit base model with trainable LoRA adapters and a single 80 GB GPU memory budget

Fine-Tuning Llama 3 with QLoRA: A Practical Guide

QLoRA (Quantized Low-Rank Adaptation) changed the economics of LLM fine-tuning. Before QLoRA, fine-t...