Fine tuned average increased from 0.1971 to 0.2321.
PROJECT 10 / LLM Evaluation · Medical AI
LLM Evaluation Framework
A framework that compared a general Mistral 7B model with its medically fine tuned version to test whether specialization actually improved the answers.
MEASURED EVIDENCE
What can be verified.
Metrics and outputs drawn from the project artifacts—not estimates added for presentation.
The tuned model won 13 of 20 benchmark questions.
Pharmacology, pathophysiology, anatomy and physiology, treatment, and diagnosis.
A deliberately small, transparent project authored benchmark.
The repository contains per question results, aggregated JSON, an interactive Plotly dashboard, and a research analysis. The benchmark is too small to establish clinical correctness.
Why this project matters
Fine tuning a model does not automatically mean it became better. This project came from the need to test that claim directly across several medical topics instead of judging a few impressive examples.
What was developed
The evaluation introduced a 20 question benchmark spanning five medical domains, generated answers from both the base and LoRA tuned models, and compared them with ROUGE scores and pairwise wins in an interactive application.
What it means
The fine tuned model recorded a 17.8% ROUGE improvement and won 65% of the comparisons in this small benchmark. More importantly, the project shows that model development needs repeatable evaluation, not intuition alone, while recognizing that text overlap is not the same as clinical correctness.
04 / WHERE THIS WORK APPLIES
From project to practical use.
Evaluation frameworks can support model selection, prompt testing, fine tuning experiments, quality assurance, safety review, vendor comparisons, and release decisions for AI systems in healthcare, customer service, education, finance, and internal knowledge tools.
What the project does not solve yet.
The current 20 question evaluation and ROUGE based scoring are too small to capture clinical correctness, reasoning quality, safety, and nuanced failure modes.
Future advancement
Build a larger clinician reviewed benchmark with factuality, safety, calibration, pairwise preference, and reproducible automated evaluation dashboards.
TECHNICAL TOOLKIT