← Back to all projects

PROJECT 10 / LLM Evaluation · Medical AI

LLM Evaluation Framework

A framework that compared a general Mistral 7B model with its medically fine tuned version to test whether specialization actually improved the answers.

DATASETCurated medical QA evaluation set
SIZE20 questions · 5 clinical domains
SOURCEProject authored evaluation benchmark ↗

MEASURED EVIDENCE

What can be verified.

Metrics and outputs drawn from the project artifacts—not estimates added for presentation.

+17.8%Average ROUGE gain

Fine tuned average increased from 0.1971 to 0.2321.

65%Pairwise win rate

The tuned model won 13 of 20 benchmark questions.

5Clinical domains

Pharmacology, pathophysiology, anatomy and physiology, treatment, and diagnosis.

20Evaluation questions

A deliberately small, transparent project authored benchmark.

VALIDATION

The repository contains per question results, aggregated JSON, an interactive Plotly dashboard, and a research analysis. The benchmark is too small to establish clinical correctness.

01

Why this project matters

Fine tuning a model does not automatically mean it became better. This project came from the need to test that claim directly across several medical topics instead of judging a few impressive examples.

02

What was developed

The evaluation introduced a 20 question benchmark spanning five medical domains, generated answers from both the base and LoRA tuned models, and compared them with ROUGE scores and pairwise wins in an interactive application.

03

What it means

The fine tuned model recorded a 17.8% ROUGE improvement and won 65% of the comparisons in this small benchmark. More importantly, the project shows that model development needs repeatable evaluation, not intuition alone, while recognizing that text overlap is not the same as clinical correctness.

04 / WHERE THIS WORK APPLIES

From project to practical use.

Evaluation frameworks can support model selection, prompt testing, fine tuning experiments, quality assurance, safety review, vendor comparisons, and release decisions for AI systems in healthcare, customer service, education, finance, and internal knowledge tools.

05 / LIMITATION

What the project does not solve yet.

The current 20 question evaluation and ROUGE based scoring are too small to capture clinical correctness, reasoning quality, safety, and nuanced failure modes.

06 / NEXT ITERATION

Future advancement

Build a larger clinician reviewed benchmark with factuality, safety, calibration, pairwise preference, and reproducible automated evaluation dashboards.

TECHNICAL TOOLKIT

Mistral 7BLoRAROUGEEvaluationMedical AI
NEXT PROJECTPFAS Drinking Water Decision Intelligence→