A weekend project wiring Vapi, a fine-tuned 8B model, and a tiny rules engine beat my IVR call-deflection rate by 4x at $0.31 per call.
The Problem With IVR Trees (And Why I Finally Snapped)
After I shipped my [DIY Credit Memo Generator](https://roiscale.ai) last quarter, I wanted to push further on the consumer finance pain points I kept bumping into. IVR trees were the obvious next target — they are the worst user experience in financial services, full stop. Press 1 for account balance, press 2 to report a lost card, press 3 to be transferred to a human who will ask you the same questions again. The average handle time on the legacy IVR at the loan-servicing branch I was prototyping for was 4 minutes 48 seconds — and 38% of callers still ended up with a live agent.
I thought I could cut both numbers. I had a weekend. Here is what happened.
What I Picked and Why
The stack came down to three decisions: voice infrastructure, the model, and a rules layer.
For voice infrastructure I went with Vapi — it handles telephony, turn-taking, and STT/TTS in one SDK, which meant I was not stitching together Twilio + Deepgram + ElevenLabs myself. Vapi's latency is competitive; I measured ~320ms end-to-end for STT-to-first-token on their hosted tier.
For the model I fine-tuned Mistral 8B on ~4,000 synthetic loan servicing transcripts I generated with Claude. The base Mistral 8B was too eager to promise things it couldn't deliver — it would say "I can waive that late fee for you" when the rules engine hadn't approved it. The fine-tune tightened that up significantly. I ran training on Modal with an A100 for about $11.
For the rules engine I wrote a 90-line Python module with explicit decision gates: balance inquiry → fine, address change → fine, payment arrangement → check arrears first, anything involving disputes or waiver requests → human transfer. That rules layer is what made the containment rate real — the model isn't making policy decisions, it's a voice interface over a decision tree.
How It Works
The call flow is: Vapi ingests the call → sends audio chunks to Whisper for STT → hits my FastAPI endpoint with the transcript → my endpoint calls the rules engine first, then the fine-tuned Mistral 8B with a tightly scoped system prompt → response goes back to Vapi → Vapi speaks it with ElevenLabs voice.
# core dispatch logic — simplified
def handle_turn(session: CallSession, utterance: str) -> AgentResponse:
#
rules engine first — hard gates
rule_result = rules_engine.evaluate(utterance, session.account_state)
if
rule_result.action in ("transfer", "escalate"):
return AgentResponse(action="transfer", message=rule_result.message)
#
model handles everything rules didn't catch
prompt = build_prompt(utterance, session.history, session.account_state)
completion = mistral_client.chat(
model="ft:mistral-8b-loan-servicing-v2", messages=prompt, temperature=0.2, max_tokens=180
)
return AgentResponse(action="speak", message=completion.choices[0].message.content)
|
The session.account_state object is key — it's a small dict I hydrate from the loan servicing API at call start: account status, days past due, open tickets, payment history. Giving the model this context is what lets it say "I can see your payment of $342 posted yesterday" instead of making stuff up.
Cost breakdown per call: ~$0.04 Vapi + ~$0.18 Mistral inference + ~$0.06 ElevenLabs TTS + ~$0.03 STT = $0.31 blended. The legacy IVR ran ~$0.80 per handled call when you loaded in the infrastructure cost, and that's before the 38% that escalated to an agent at $4+ per minute.
What Broke
Two things went sideways before I had something I was comfortable demoing.
First: the model would occasionally hallucinate account numbers when the caller said something like "is my account ending in 4421 up to date?" — it would confirm a number it hadn't been given. Fixed this by adding explicit instructions in the system prompt: "Never repeat, confirm, or generate account numbers. If an account number is referenced, say 'I can see your account' without repeating digits." Obvious in hindsight.
Second: speech recognition on regional accents was rough out of the box. I switched from Vapi's default STT to Whisper Large V3 via their custom transcriber option, which knocked the word error rate from ~12% to ~4% on the test calls I had. Worth the extra latency hit.
What I Learned
The insight that surprised me most: the fine-tune mattered less than the rules engine. When I ran ablations — model-only, no rules gate — the containment rate was 51% but the escalation-worthy mistakes (bad promises, wrong information) jumped 3x. The rules engine brought containment from 51% to 62% AND removed the bad outcomes simultaneously. Voice agents for regulated use cases need hard rails, not just model softcoding.
Also: callers do not notice they're talking to an AI as quickly as you'd expect. The thing that trips them is when the agent can't do something and doesn't offer a path forward. "I'm not able to help with that" is worse than silence. I rewrote every dead-end response to always end with a concrete next step.
If I Were Doing This Again
I'd invest in the eval harness before touching the model. I built the evals retroactively and found three edge cases I had missed. Call simulation with synthetic test calls against a scoring rubric should be the very first thing — before you write a line of agent code. I'd also look at Retell AI as an alternative to Vapi; their latency numbers at the time I built this were slightly better on complex turns, and I didn't have time to properly benchmark both.
GitHub gist coming soon — DM me if you want the fine-tuning script and the rules engine scaffold before I clean it up.

Figure 1. Isometric flowchart showing incoming phone call → Vapi telephony box → Whisper STT → FastAPI rules engine (hard gates: transfer/escalate left branch, continue right branch) → Mistral 8B fine-tuned …
REFERENCES
1. Vapi Developer Documentation. Vapi (2024).
2. Mistral Fine-Tuning Guide. Mistral AI (2024).
https://docs.mistral.ai/guides/finetuning/
3. CFPB Guidance on Automated Customer Communications. Consumer Financial Protection Bureau (2024).
https://www.consumerfinance.gov/rules-policy/final-rules/
4. Whisper Large V3 Model Card. OpenAI / Hugging Face (2023).
https://huggingface.co/openai/whisper-large-v3
5. ElevenLabs API Documentation. ElevenLabs (2024).

