Rex rents a 4×A100 box for 72 hours and fine-tunes three open-weight models on the same financial document QA dataset — publishing loss curves, eval deltas, and the inference cost math that determines which one he'd actually deploy.
|
BUILDS ON I Generated 50,000 Synthetic Training Examples (V1) |
What I Was Testing and Why
In my V1 piece I covered generating 50,000 synthetic training examples with an LLM and the failure modes when you push past the 30% synthetic data threshold. This time I extended that work into the fine-tuning itself: taking three leading open-weight models and running the same domain adaptation experiment against all three to find out which architecture benefits most from the same training budget.
The task: financial document QA. A dataset of 10-K filings, earnings call transcripts, and loan covenant documents. Questions range from simple extraction ("What was EBITDA in Q3?") to multi-hop reasoning ("If the covenant requires net debt/EBITDA below 3.5x and net debt increased 18% while EBITDA stayed flat, has the covenant been breached?"). This is a real use case — I built a version of this for a small PE shop earlier this year.
The three models:
• Qwen 2.5 72B (Instruct, BF16 base, fine-tuned with QLoRA r=64 α=128)
• Llama 3.3 70B (Instruct, BF16 base, same QLoRA config)
• Mistral Small 3.1 24B (Instruct, BF16 base, QLoRA r=32 α=64 — smaller model, lower rank)
All fine-tuned with Unsloth for memory efficiency on 4×A100 80GB via Lambda Labs at $3.80/GPU-hour.
Methodology
Dataset: 8,400 training examples (70% real extracted QA pairs, 30% synthetic from Claude Sonnet 4.5), 800 validation examples, 400 held-out test examples. All human-reviewed for quality.
Training config was identical across all three where architecture allowed:
# Unsloth fine-tuning config (applied to all
three models)
from unsloth import FastLanguageModel
from trl import SFTTrainer, TrainingArguments
model, tokenizer =
FastLanguageModel.from_pretrained(
model_name="Qwen/Qwen2.5-72B-Instruct", # swapped per model
max_seq_length=8192, load_in_4bit=False, # BF16 for
cleaner loss curves
dtype=torch.bfloat16, )
model = FastLanguageModel.get_peft_model(
model, r=64, target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"], lora_alpha=128, lora_dropout=0.05, use_gradient_checkpointing="unsloth", )
training_args = TrainingArguments(
output_dir=f"./checkpoints/{MODEL_NAME}", num_train_epochs=3, per_device_train_batch_size=4, gradient_accumulation_steps=4, learning_rate=2e-4, lr_scheduler_type="cosine", warmup_ratio=0.05, bf16=True, logging_steps=50, save_steps=500, eval_steps=200, evaluation_strategy="steps", )
|
Eval metric: a custom scoring function combining exact-match answer accuracy (40%), ROUGE-L on explanation quality (30%), and a GPT-4o-mini judge for financial reasoning coherence (30%). I've found this triple-metric approach more reliable than ROUGE alone for financial QA — ROUGE rewards fluency, the LLM judge catches factual reasoning errors that sound correct.
# Training wall-clock times on 4xA100 # Qwen 2.5 72B: ~14 hours (3 epochs) # Llama 3.3 70B: ~13 hours (3 epochs) # Mistral Small: ~6 hours (3 epochs) |
Results
Baseline scores (zero-shot on held-out test):
• Qwen 2.5 72B: 58.4
• Llama 3.3 70B: 55.2
• Mistral Small 3.1 24B: 49.1
Post fine-tune scores:
• Qwen 2.5 72B: 72.6 (+14.2 pts) · Fine-tune cost: $210
• Llama 3.3 70B: 67.0 (+11.8 pts) · Fine-tune cost: $180
• Mistral Small 3.1 24B: 58.5 (+9.4 pts) · Fine-tune cost: $84
Loss curves told interesting stories. Qwen 2.5 hit its best validation loss at epoch 2.4 and started slightly overfitting after. Llama 3.3 was the most stable — validation loss tracked training loss cleanly through all 3 epochs. Mistral Small overfit the fastest, hitting best checkpoint at epoch 1.8. If I were running Mistral again I'd do 2 epochs max.
Inference cost per 1K output tokens (self-hosted on an A100 80GB with vLLM):
• Mistral Small 3.1 24B: ~$0.0009
• Llama 3.3 70B: ~$0.0022
• Qwen 2.5 72B: ~$0.0024
Failure Modes
Qwen 2.5 is the strongest starting model and improves the most — but it's also the most expensive to run and has a quirk I didn't anticipate: it's aggressive about format following from the fine-tune. After training on my structured JSON output format, it started adding JSON formatting to responses that explicitly asked for plain text. You'd need a lightweight post-processing layer in production.
Llama 3.3 was the most predictable fine-tune. Stable loss, consistent improvement, no weird behaviors. The lower ceiling vs Qwen is real but Llama 3.3's inference characteristics are better-understood — more serving infrastructure supports it and the community debug knowledge is deeper. If I'm deploying a 70B model and want to minimize production surprises, Llama 3.3 is the safer call.
Mistral Small 3.1 24B is the value play. At $84 to fine-tune and $0.0009 per 1K output tokens it's dramatically cheaper on both dimensions. The 9.4-point improvement brings it to within 3.5 points of fine-tuned Llama 3.3 on our eval. For a price-sensitive deployment where you're okay with occasional quality misses on the hardest multi-hop questions, this is the model. One gotcha: it hallucinated financial figures at a higher rate than the other two on questions where the answer wasn't in the context — it makes up plausible-looking numbers rather than saying "I don't know."
All three struggled with questions that required cross-document reasoning (linking facts from the 10-K to the earnings call to the covenant document). That class of question probably needs a RAG+fine-tune hybrid rather than pure fine-tuning.
What I Would Actually Deploy
For accuracy-critical production (audit support, covenant compliance): Qwen 2.5 72B fine-tuned. Add the post-processing layer to strip unwanted formatting and monitor for the hallucination pattern on out-of-context questions.
For a cost-sensitive internal tool (analyst first-pass, human reviews output): Mistral Small 3.1 24B fine-tuned. $84 to fine-tune, cheap to serve, good enough for 80% of queries.
I'm not deploying Llama 3.3 in this specific use case — it didn't win on accuracy or cost. But in a different domain it might — the stable fine-tuning behavior is worth something when your ML infra team is stretched thin.
Repo / Gist Coming
The training scripts, eval harness, and loss curve plots will be in the repo. I'll also include the synthetic data generation pipeline from the V1 piece since several people asked for it as a standalone module.

Figure 4. Side-by-side bar chart: three models on x-axis, grouped bars showing baseline score, post-finetune score, and inference cost index. Secondary axis shows fine-tune cost ($). Dark terminal aesthetic …
REFERENCES
1. Qwen2.5 Technical Report. Alibaba DAMO Academy / arXiv (2024).
https://arxiv.org/abs/2412.15115
2. Llama 3 Model Card — Meta AI. Meta AI (2024).
https://ai.meta.com/blog/meta-llama-3/
3. Mistral Small 3.1 Model Documentation. Mistral AI (2025).
https://docs.mistral.ai/getting-started/models/models_overview/
4. Unsloth — Efficient Fine-Tuning Library. Unsloth / GitHub (2025).
https://github.com/unslothai/unsloth
5. QLoRA: Efficient Finetuning of Quantized LLMs. Dettmers et al. / arXiv (2023).
https://arxiv.org/abs/2305.14314
6. vLLM: Efficient Memory Management for LLM Serving with PagedAttention. Kwon et al. / arXiv (2023).
https://arxiv.org/abs/2309.06180
7. TRL — Transformer Reinforcement Learning Library. Hugging Face (2025).
https://huggingface.co/docs/trl/index


