← Back to all projects

PROJECT 06 / LLM Fine tuning · Research

Medical Q&A Model

A research prototype that taught a large language model to better follow medical question and answer patterns without retraining all seven billion model parameters.

DATASETMedical Meadow Medical Flashcards
SIZE5,000 fine tuning examples
SOURCEHugging Face · medalpaca/medical_meadow_medical_flashcards ↗

MEASURED EVIDENCE

What can be verified.

Metrics and outputs drawn from the project artifacts—not estimates added for presentation.

0.6573Final validation loss

Improved from 0.6814 at the first recorded checkpoint.

0.36%Parameters trained

13.6M trainable parameters using LoRA rather than full model fine tuning.

$0Compute cost

Approximately 56 minutes on a free Google Colab T4 GPU.

+17.8%Follow on ROUGE gain

The tuned adapter won 13 of 20 comparisons in a five domain evaluation.

VALIDATION

Reproducible proof includes the public LoRA adapter on Hugging Face, training loss history, sample outputs, evaluation results, and source notebooks on GitHub.

PLAIN-LANGUAGE INTERPRETATION

How to read these results: lower validation loss means the adapter became better at predicting held out medical QA patterns. The follow on ROUGE comparison then tested whether that training improvement translated into answers with stronger reference overlap, while still not proving clinical correctness.

01

Why this project matters

General language models may not handle specialized medical wording consistently, but fully retraining a large model is expensive. This project explored whether a smaller, efficient training method could adapt an existing model to a focused medical question and answer dataset.

02

What was developed

The training process used LoRA and QLoRA to fine tune selected parts of Mistral 7B on 5,000 medical flashcard examples, then packaged the learned adapter so the experiment could be reproduced and shared without duplicating the full base model.

03

What it means

The adapter reached a final validation loss of 0.6573 after training on 5,000 examples. Only 13.6 million parameters, 0.36% of the reported parameter count, were trained, reducing the model’s memory footprint from roughly 14 GB to 3.5 GB through 4 bit quantization. Training completed in approximately 56 minutes on a free Google Colab T4 GPU at $0 compute cost. A follow on 20 question evaluation found a 17.8% ROUGE improvement and a 65% win rate over the base model.

04 / WHERE THIS WORK APPLIES

From project to practical use.

Parameter efficient fine tuning can be applied to controlled assistants for healthcare education, legal or financial terminology, technical support, research summarization, organizational knowledge, and other specialized domains with strong expert oversight.

05 / LIMITATION

What the project does not solve yet.

The adapter was trained on a bounded medical dataset and has not undergone clinician led safety validation; outputs must not be treated as medical advice.

06 / NEXT ITERATION

Future advancement

Evaluate on larger clinician reviewed benchmarks, add retrieval grounding and safety filters, and conduct structured human evaluation for factuality and harm.

TECHNICAL TOOLKIT

Mistral 7BLoRAQLoRAPEFTTRL
NEXT PROJECTSkin Disease Classifier→