smaple.tr
LLM fine-tuning

LLM Fine-Tuning Guide: Custom Model Training with LoRA and QLoRA [2026]

Mehmet Kurtipek
August 10, 2026
12 min read
LLM fine-tuning
LoRA
QLoRA
PEFT
custom LLM training
parameter efficient fine-tuning

General-purpose LLMs are impressive. They can write code, summarize documents, answer questions, and translate between languages with reasonable accuracy. But "reasonable" is not good enough when your application requires consistent output format, domain-specific terminology, a particular tone, or behavior that the base model was never trained to exhibit. That gap is where fine-tuning matters.

Fine-tuning updates a model's weights on your specific data, changing how the model behaves — not just what information it can access. This distinguishes it from prompt engineering (same model, different instructions) and RAG (same model, different knowledge). The investment is higher. So is the ceiling.

This guide covers the practical decisions: when to fine-tune versus prompt engineer or use RAG, which method (LoRA, QLoRA, full fine-tuning) fits your constraints, how to prepare training data, what hardware you actually need, and how to evaluate and deploy the result.

LLM Fine-Tuning Guide: When to Fine-Tune vs Prompt Engineer vs RAG

The first decision is whether fine-tuning is the right tool. Each of the three main customization approaches has a different cost-benefit profile.

Prompt engineering modifies the model's instructions without modifying the model. It is fast to implement, cheap to iterate, and requires no training infrastructure. Its ceiling is the base model's existing capabilities. If the model cannot reliably produce the output format you need through prompting alone, prompting will not solve the problem.

RAG (Retrieval-Augmented Generation) gives the model access to external knowledge at inference time. It handles the "the model doesn't know this specific fact" problem — current product catalog, recent policy updates, proprietary documentation. It does not change the model's behavior or output format.

Fine-tuning changes the model's behavior by updating its weights. It handles: consistent output format (the model always produces structured JSON in a specific schema), domain terminology (medical, legal, or technical vocabulary the model uses inconsistently by default), tone and style (formal/informal, verbose/concise, brand voice), and task specialization (code in a specific style, document classification against a proprietary taxonomy).

The practical rule: try prompt engineering first. If the model is inconsistent or cannot be prompted to the required behavior, try RAG if the issue is knowledge gaps. If the issue is behavior (format, style, terminology, task-specific accuracy), fine-tune.

These approaches combine well. A fine-tuned model with RAG retrieval combines behavioral consistency (fine-tuning) with up-to-date knowledge (RAG).

Fine-Tuning Methods: Full, LoRA, QLoRA, and PEFT

Four methods cover most use cases. They differ in GPU memory requirement, training time, and performance ceiling.

Full Fine-Tuning

Updates all model parameters. Theoretically achieves the highest performance ceiling. Practically requires enormous GPU memory — a 7B parameter model requires 60+ GB VRAM for full fine-tuning, which means at minimum an A100 80GB. For 13B or 70B models, full fine-tuning requires multi-GPU or distributed training setups.

Full fine-tuning is justified for large-scale domain specialization projects where maximum performance is required and GPU budget is not a constraint. For most enterprise and production use cases, parameter-efficient methods achieve equivalent results at a fraction of the cost.

LoRA (Low-Rank Adaptation)

LoRA freezes the original model weights and adds small trainable adapter matrices at specific layers. The original model (7B parameters, for example) is untouched. Only the adapter matrices (tens of millions of parameters, typically 0.1-1% of the total) are trained.

Why this works: language model weight matrices can be updated through low-rank decompositions without losing significant performance. The adapter captures the behavioral changes you need; the frozen original model provides the language capabilities.

Practical benefits:

  • A 7B model fine-tuned with LoRA requires 16-24 GB VRAM instead of 60+ GB
  • Multiple LoRA adapters can be created for different tasks and swapped onto the same base model
  • Training time is 3-5x faster than full fine-tuning for comparable results
  • The adapter file is small (typically 50-500MB) — easy to version and deploy

QLoRA (Quantized LoRA)

QLoRA combines LoRA with 4-bit quantization of the base model. The base model's weights are quantized from 32-bit or 16-bit floating point to 4-bit — reducing memory footprint by 4-8x. LoRA adapters are then trained on top of this quantized base.

The result: a 70B parameter model can be fine-tuned on a single consumer GPU (NVIDIA RTX 4090 with 24GB VRAM). Performance loss from quantization is minimal for most tasks — typically less than 1-2% accuracy drop compared to LoRA on the full-precision model.

QLoRA is the method that made LLM fine-tuning accessible to organizations without data center GPU infrastructure. It is the right starting point for most projects.

PEFT Methods Comparison

Method Trainable Parameters GPU VRAM (7B Model) Training Time Performance
Full Fine-Tuning 100% 60+ GB High Highest ceiling
LoRA 0.1-1% 16-24 GB Medium High
QLoRA 0.1-1% 8-12 GB Medium-Low High (minimal loss)
Prefix Tuning 0.01-0.1% 12-16 GB Low Medium-High
Prompt Tuning 0.001% 10-14 GB Very low Medium

Hugging Face's PEFT library implements all of these methods with a consistent API. For most projects, QLoRA is the right starting point. Upgrade to LoRA on a full-precision model if you observe accuracy deficits from quantization. Reserve full fine-tuning for the rare case where LoRA cannot reach required performance.

Training Data Preparation

The single most important variable in fine-tuning outcomes is data quality. Model training amplifies the patterns in your data — both good patterns and bad ones.

Data Format

Fine-tuning for instruction following uses the instruction-response pair format:

{
  "messages": [
    {"role": "system", "content": "You are a technical support agent for Acme ERP software."},
    {"role": "user", "content": "How do I generate a monthly expense report?"},
    {"role": "assistant", "content": "To generate a monthly expense report: 1. Navigate to Reports > Expense Reports. 2. Select the month and department. 3. Click Generate. The report will download as PDF or Excel."}
  ]
}

The system message defines the model's role and constraints. The user message is the input. The assistant message is the target output the model should learn to produce.

JSONL format (one JSON object per line) is the standard for training datasets.

Data Volume

Minimum useful dataset size depends on what you are trying to teach:

  • Format and style changes: 200-500 examples can be sufficient. The model already knows how to write — you are teaching it a specific format.
  • Domain terminology and knowledge: 500-2000 examples for moderate depth. The model needs to see enough examples to generalize.
  • Task specialization: 1000-5000 examples for reliable performance. Classification, extraction, and generation tasks benefit from more variety.
  • Behavior that strongly contradicts base model defaults: 5000+ examples and multiple training epochs.

More data is not always better. 500 carefully curated, high-quality examples typically outperform 5000 mediocre examples. Diversity matters: examples that cover edge cases, unusual inputs, and varied phrasings generalize better than examples that cluster around the same patterns.

Data Quality

Before training:

  • Review a random sample (50-100 examples) manually. If you find errors, extrapolate — the full dataset likely has proportional errors.
  • Check for format inconsistency. JSON parsing errors during training will silently corrupt the dataset.
  • Remove examples where the target output is wrong or inconsistent with other examples on similar inputs.
  • Anonymize PII before training. GDPR and equivalent regulations apply to training data, not just inference data.

A common failure mode: using LLM-generated data to fine-tune an LLM. If the training data was generated by GPT-4, fine-tuning Llama 3 on that data teaches Llama 3 to imitate GPT-4's outputs, including its failure modes. Use ground truth human-generated data where possible.

Choosing a Base Model

The base model selection constrains everything else — maximum performance ceiling, supported languages, GPU requirements, and licensing.

Open-Source Models

Llama 3.x (Meta) — the dominant open-source family in 2026. Available in 8B, 70B, and 405B sizes. Strong multilingual capabilities, commercially licensable (Meta's license). The Llama 3 series shows significant improvement in instruction following and code compared to earlier versions. First choice for most projects.

Mistral and Mixtral (Mistral AI) — Mistral 7B punches above its weight relative to parameter count. The Mixtral MoE (Mixture of Experts) architecture provides strong performance with lower inference cost than equivalent dense models. Apache 2.0 license — fully open, including commercial use.

Qwen 2.x (Alibaba) — strong multilingual support across 29 languages, including good performance on non-English tasks. Relevant for multilingual fine-tuning projects.

Phi-3 (Microsoft) — small models (3.8B-14B) with strong benchmark performance. Useful when inference cost is the primary constraint and task complexity allows a smaller model.

Commercial API Fine-Tuning

OpenAI offers fine-tuning for GPT-3.5-turbo and GPT-4o-mini via API. Anthropic offers fine-tuning for Claude models at enterprise tier. No infrastructure management required — you provide data, the provider trains and hosts.

Trade-offs: per-token cost at inference time is significantly higher than self-hosted open-source models. Data privacy requires reviewing the provider's data processing agreement. Vendor lock-in on model version and pricing.

Open-source is the better default for most enterprise projects where long-term cost, data sovereignty, and customization depth matter.

Hardware Requirements and Infrastructure

GPU Selection Guide

Model Size QLoRA LoRA Full Fine-Tuning
7B parameters 1x RTX 4090 (24GB) 2x A100 40GB 4x A100 40GB
13B parameters 2x RTX 4090 2x A100 80GB 8x A100 40GB
70B parameters 4x RTX 4090 4x A100 80GB 16x A100 80GB

Training duration (approximate, QLoRA, 1000 examples, 3 epochs):

  • 7B model: 1-3 hours on RTX 4090
  • 13B model: 3-6 hours on RTX 4090
  • 70B model: 8-16 hours on 4x RTX 4090

Cloud Infrastructure

For teams without on-premise GPU hardware, cloud spot instances reduce cost significantly (60-70% compared to on-demand):

  • AWS: p3.8xlarge or p4d instances (A100s)
  • Google Cloud: a2-highgpu instances
  • Azure: NCas T4 v3 or NDv4 series

GPU-specialized platforms offer lower prices and simpler setup:

  • Lambda Labs — purpose-built for ML training, competitive pricing, simple provisioning
  • RunPod — flexible GPU rental including consumer GPUs for QLoRA
  • Vast.ai — marketplace for GPU compute, lowest prices at the cost of more variability

Training Frameworks

Hugging Face TRL — the standard library for instruction fine-tuning with PEFT methods. Well-documented, actively maintained, direct integration with the Transformers and PEFT libraries.

Axolotl — configuration-driven training that abstracts framework details. A YAML config file specifies model, dataset, method, and hyperparameters. Good choice for teams that want reproducible training runs without deep framework knowledge.

LLaMA-Factory — web interface and CLI for fine-tuning. Accessible to teams without deep Python expertise.

Evaluation

Automatic Metrics

Perplexity measures how surprised the model is by a held-out test set — lower is generally better. BLEU and ROUGE measure overlap between generated and reference outputs, useful for tasks with exact target outputs. BERTScore measures semantic similarity between generated and reference text.

Important caveat: these metrics measure what they measure, not what you care about. A model with good ROUGE scores may still produce outputs that are factually wrong, stylistically inconsistent, or off-topic for your specific use case.

Task-Specific Evaluation

Define evaluation criteria specific to your task before training starts:

  • Format compliance rate (what percentage of outputs match the required format?)
  • Domain accuracy (evaluated by subject matter experts)
  • Consistency (same input always produces equivalent output?)
  • Instruction adherence (does the model follow constraints in the system prompt?)

Human evaluation by domain experts is essential for production deployment decisions. Automatic metrics are useful for tracking training progress; human evaluation determines whether the model is actually fit for purpose.

Iterative Improvement

Fine-tuning is not a one-shot process. First iteration reveals failure modes. Common patterns:

  • Model produces correct format most of the time, but fails on specific input types → add training examples for those types
  • Model over-corrects toward training distribution (outputs become repetitive) → reduce epochs or add more diversity
  • Model forgets capabilities from base model (catastrophic forgetting) → check if training examples are too narrow

Plan for 2-3 training iterations before evaluating production readiness.

Deploying the Fine-Tuned Model

vLLM is the standard high-performance inference server for open-source LLMs in 2026. It implements PagedAttention for efficient KV cache management, producing 2-4x higher throughput than naive PyTorch inference. Supports LoRA adapter hot-swapping — multiple LoRA adapters can be loaded and swapped without restarting the server.

Text Generation Inference (TGI) from Hugging Face is the alternative. Similar performance, tighter integration with the Hugging Face ecosystem.

Quantization for inference reduces memory requirements further. GPTQ and AWQ are the standard quantization formats for inference — a 7B model in 4-bit AWQ uses approximately 4GB VRAM versus 14GB for the full bfloat16 model, with minimal accuracy degradation.

Production checklist before deployment:

  • Guardrails — input and output filtering for harmful content, PII leakage
  • Monitoring — latency, throughput, error rate, prediction distribution
  • A/B framework — test new model version against current production model before full rollout
  • Rollback plan — how quickly can you revert to the previous version?

Enterprise Use Cases Where Fine-Tuning Delivers

Customer support automation — a model fine-tuned on resolved support tickets learns to answer in your company's voice, follow your escalation policy, and reference your product correctly. Better than prompting a general model because the tone, terminology, and format are consistent at scale.

Legal and financial document processing — contract clause extraction, regulatory filing analysis, risk flagging. Domain terminology is critical; a model that confuses legal terms produces unreliable outputs. Fine-tuning on annotated legal documents improves task accuracy dramatically.

Internal tooling and code generation — a model fine-tuned on your codebase produces code that follows your conventions, uses your internal libraries, and adheres to your architecture patterns. General-purpose code models do not know about your internal APIs.

Healthcare documentation — clinical note summarization, ICD coding assistance, patient communication drafting. Requires on-premise deployment for HIPAA compliance. Fine-tuning enables the specialized vocabulary and output format healthcare applications require.

LLM fine-tuning is no longer a research activity. The tooling (QLoRA, Axolotl, vLLM) is mature. The hardware is accessible. The workflow is documented. For organizations that need LLM behavior they cannot achieve through prompting, fine-tuning is the practical path forward.

Related Articles

August 11, 2026

MLOps Guide: Taking Machine Learning Models to Production [2026]

87% of machine learning models built by data science teams never reach production. The models work — they pass cross-validation, they score well on holdout sets, they demonstrate genuine predictive value. The problem is not the modeling. The problem is everything that happens between a notebook experiment and a reliable, monitored, production system. MLOps is the discipline that closes that gap. This guide covers the full MLOps stack: maturity levels, tooling choices (MLflow, DVC, Kubeflow

Read More
August 9, 2026

Computer Vision Applications: Object Detection, OCR, and Industrial AI [2026]

Computer vision has moved well past the research phase. The models are trained, the frameworks are mature, the hardware is accessible, and the use cases are generating measurable returns. What was a specialized capability requiring deep expertise in 2018 is now deployable infrastructure — if you know which component to reach for and where the real complexity lives. This guide covers computer vision applications across industrial, medical, logistics, and document processing domains. It expl

Read More
August 8, 2026

Natural Language Processing Guide: Building Production NLP Systems [2026]

Natural language processing has crossed a threshold. For most of the 2010s, production NLP meant rule-based pipelines, hand-crafted feature extractors, and models that were brittle the moment text deviated from the training distribution. Today, transformer-based models — BERT, RoBERTa, and their descendants — handle tasks that once required months of annotation work with a few hours of fine-tuning. This guide covers what NLP actually looks like in production in 2026: the architectures that w

Read More