Skip to main content
ROI Scale AI logoROI Scale AI
Business
Technology & Telecom
arrow_forward
Financial Services
arrow_forward
Healthcare
arrow_forward
Retail & E-Commerce
arrow_forward
Education
arrow_forward
Energy & Utilities
arrow_forward
Media & Entertainment
arrow_forward
Manufacturing & Industrial
arrow_forward
Real Estate & Construction
arrow_forward
Government & Public Sector
arrow_forward
Professional Services
arrow_forward
Transport and Logistics
arrow_forward
View all in Business arrow_forward
Technology
Models & Benchmarks
arrow_forward
AI Engineering
arrow_forward
Harness Engineering
arrow_forward
Data Strategy
arrow_forward
AI Security & Governance
arrow_forward
Libraries & Frameworks
arrow_forward
AI for Developers
arrow_forward
Research & Papers
arrow_forward
View all in Technology arrow_forward
Marketplace
Blueprints
arrow_forward
Proof Packs
arrow_forward
View all in Marketplace arrow_forward
Contribute
How-Tos
arrow_forward
Business RoadMap
arrow_forward
Tech RoadMap
arrow_forward
View all in Contribute arrow_forward
search
person_outlineSign In
Categories
BusinessTechnology & TelecomFinancial ServicesHealthcareRetail & E-CommerceEducationEnergy & UtilitiesMedia & EntertainmentManufacturing & IndustrialReal Estate & ConstructionGovernment & Public SectorProfessional ServicesTransport and Logistics
TechnologyModels & BenchmarksAI EngineeringHarness EngineeringData StrategyAI Security & GovernanceLibraries & FrameworksAI for DevelopersResearch & Papers
MarketplaceBlueprintsProof Packs
ContributeHow-TosBusiness RoadMapTech RoadMap
searchSearchhomeHome
Community
person_outlineSign In / Join
Home/Technology/Models & Benchmarks
September 25, 2026 6 min read

Fine-Tuning Qwen 2.5 vs Llama 3.3 vs Mistral Small for Domain Tasks: My 4-GPU Bake-Off

Rex Circuit
Rex Circuit Published Sep 25, 2026
Fine-Tuning Qwen 2.5 vs Llama 3.3 vs Mistral Small for Domain Tasks: My 4-GPU Bake-Off

Rex rents a 4×A100 box for 72 hours and fine-tunes three open-weight models on the same financial document QA dataset — publishing loss curves, eval deltas, and the inference cost math that determines which one he'd actually deploy.

BUILDS ON

I Generated 50,000 Synthetic Training Examples (V1)

What I Was Testing and Why

In my V1 piece I covered generating 50,000 synthetic training examples with an LLM and the failure modes when you push past the 30% synthetic data threshold. This time I extended that work into the fine-tuning itself: taking three leading open-weight models and running the same domain adaptation experiment against all three to find out which architecture benefits most from the same training budget.

The task: financial document QA. A dataset of 10-K filings, earnings call transcripts, and loan covenant documents. Questions range from simple extraction ("What was EBITDA in Q3?") to multi-hop reasoning ("If the covenant requires net debt/EBITDA below 3.5x and net debt increased 18% while EBITDA stayed flat, has the covenant been breached?"). This is a real use case — I built a version of this for a small PE shop earlier this year.

The three models:

•    Qwen 2.5 72B (Instruct, BF16 base, fine-tuned with QLoRA r=64 α=128)

•    Llama 3.3 70B (Instruct, BF16 base, same QLoRA config)

•    Mistral Small 3.1 24B (Instruct, BF16 base, QLoRA r=32 α=64 — smaller model, lower rank)

All fine-tuned with Unsloth for memory efficiency on 4×A100 80GB via Lambda Labs at $3.80/GPU-hour.


Methodology

Dataset: 8,400 training examples (70% real extracted QA pairs, 30% synthetic from Claude Sonnet 4.5), 800 validation examples, 400 held-out test examples. All human-reviewed for quality.

Training config was identical across all three where architecture allowed:

# Unsloth fine-tuning config (applied to all
  three models)
  from unsloth import FastLanguageModel
  from trl import SFTTrainer, TrainingArguments
   model, tokenizer =
  FastLanguageModel.from_pretrained(
     model_name="Qwen/Qwen2.5-72B-Instruct",  # swapped per model
     max_seq_length=8192, load_in_4bit=False,   # BF16 for
  cleaner loss curves
     dtype=torch.bfloat16, )
   model = FastLanguageModel.get_peft_model(
     model, r=64, target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"], lora_alpha=128, lora_dropout=0.05, use_gradient_checkpointing="unsloth", )
   training_args = TrainingArguments(
     output_dir=f"./checkpoints/{MODEL_NAME}", num_train_epochs=3, per_device_train_batch_size=4, gradient_accumulation_steps=4, learning_rate=2e-4, lr_scheduler_type="cosine", warmup_ratio=0.05, bf16=True, logging_steps=50, save_steps=500, eval_steps=200, evaluation_strategy="steps", )

Eval metric: a custom scoring function combining exact-match answer accuracy (40%), ROUGE-L on explanation quality (30%), and a GPT-4o-mini judge for financial reasoning coherence (30%). I've found this triple-metric approach more reliable than ROUGE alone for financial QA — ROUGE rewards fluency, the LLM judge catches factual reasoning errors that sound correct.

# Training wall-clock times on 4xA100
  # Qwen 2.5 72B:   ~14 hours (3 epochs)
  # Llama 3.3 70B:  ~13 hours (3 epochs)
  # Mistral Small:  ~6 hours (3 epochs)


Results

Baseline scores (zero-shot on held-out test):

•    Qwen 2.5 72B: 58.4

•    Llama 3.3 70B: 55.2

•    Mistral Small 3.1 24B: 49.1

Post fine-tune scores:

•    Qwen 2.5 72B: 72.6 (+14.2 pts) · Fine-tune cost: $210

•    Llama 3.3 70B: 67.0 (+11.8 pts) · Fine-tune cost: $180

•    Mistral Small 3.1 24B: 58.5 (+9.4 pts) · Fine-tune cost: $84

Loss curves told interesting stories. Qwen 2.5 hit its best validation loss at epoch 2.4 and started slightly overfitting after. Llama 3.3 was the most stable — validation loss tracked training loss cleanly through all 3 epochs. Mistral Small overfit the fastest, hitting best checkpoint at epoch 1.8. If I were running Mistral again I'd do 2 epochs max.

Inference cost per 1K output tokens (self-hosted on an A100 80GB with vLLM):

•    Mistral Small 3.1 24B: ~$0.0009

•    Llama 3.3 70B: ~$0.0022

•    Qwen 2.5 72B: ~$0.0024


Failure Modes

Qwen 2.5 is the strongest starting model and improves the most — but it's also the most expensive to run and has a quirk I didn't anticipate: it's aggressive about format following from the fine-tune. After training on my structured JSON output format, it started adding JSON formatting to responses that explicitly asked for plain text. You'd need a lightweight post-processing layer in production.

Llama 3.3 was the most predictable fine-tune. Stable loss, consistent improvement, no weird behaviors. The lower ceiling vs Qwen is real but Llama 3.3's inference characteristics are better-understood — more serving infrastructure supports it and the community debug knowledge is deeper. If I'm deploying a 70B model and want to minimize production surprises, Llama 3.3 is the safer call.

Mistral Small 3.1 24B is the value play. At $84 to fine-tune and $0.0009 per 1K output tokens it's dramatically cheaper on both dimensions. The 9.4-point improvement brings it to within 3.5 points of fine-tuned Llama 3.3 on our eval. For a price-sensitive deployment where you're okay with occasional quality misses on the hardest multi-hop questions, this is the model. One gotcha: it hallucinated financial figures at a higher rate than the other two on questions where the answer wasn't in the context — it makes up plausible-looking numbers rather than saying "I don't know."

All three struggled with questions that required cross-document reasoning (linking facts from the 10-K to the earnings call to the covenant document). That class of question probably needs a RAG+fine-tune hybrid rather than pure fine-tuning.


What I Would Actually Deploy

For accuracy-critical production (audit support, covenant compliance): Qwen 2.5 72B fine-tuned. Add the post-processing layer to strip unwanted formatting and monitor for the hallucination pattern on out-of-context questions.

For a cost-sensitive internal tool (analyst first-pass, human reviews output): Mistral Small 3.1 24B fine-tuned. $84 to fine-tune, cheap to serve, good enough for 80% of queries.

I'm not deploying Llama 3.3 in this specific use case — it didn't win on accuracy or cost. But in a different domain it might — the stable fine-tuning behavior is worth something when your ML infra team is stretched thin.


Repo / Gist Coming

The training scripts, eval harness, and loss curve plots will be in the repo. I'll also include the synthetic data generation pipeline from the V1 piece since several people asked for it as a standalone module.

P5_1_8bf0f3f0.jpg


Figure 4. Side-by-side bar chart: three models on x-axis, grouped bars showing baseline score, post-finetune score, and inference cost index. Secondary axis shows fine-tune cost ($). Dark terminal aesthetic …

REFERENCES

1. Qwen2.5 Technical Report. Alibaba DAMO Academy / arXiv (2024).

https://arxiv.org/abs/2412.15115

2. Llama 3 Model Card — Meta AI. Meta AI (2024).

https://ai.meta.com/blog/meta-llama-3/

3. Mistral Small 3.1 Model Documentation. Mistral AI (2025).

https://docs.mistral.ai/getting-started/models/models_overview/

4. Unsloth — Efficient Fine-Tuning Library. Unsloth / GitHub (2025).

https://github.com/unslothai/unsloth

5. QLoRA: Efficient Finetuning of Quantized LLMs. Dettmers et al. / arXiv (2023).

https://arxiv.org/abs/2305.14314

6. vLLM: Efficient Memory Management for LLM Serving with PagedAttention. Kwon et al. / arXiv (2023).

https://arxiv.org/abs/2309.06180

7. TRL — Transformer Reinforcement Learning Library. Hugging Face (2025).

https://huggingface.co/docs/trl/index


Share this article:
Back to Home / Technology / Models & Benchmarks

Marketplace matches for this article

Support

  • Contact Us
  • FAQ
  • Privacy
  • Terms

About

  • Mission
  • Editorial

© 2026 ROI Scale AI. All rights reserved.

Powered by Publishi.ai