← Back to all projects

PROJECT 05 / RAG · Applications

Sports Q&A Chatbot

A question answering assistant that looks through a focused sports knowledge base before responding, making answers more grounded than relying on a language model's memory alone.

DATASETCurated sports knowledge base
SIZE10 documents · 57 indexed chunks
SOURCEProject authored documents indexed in ChromaDB ↗

MEASURED EVIDENCE

What can be verified.

Metrics and outputs drawn from the project artifacts—not estimates added for presentation.

10Source documents

A curated sports knowledge base covering current and historical topics.

57Indexed chunks

Embedded and stored in ChromaDB for retrieval before generation.

7Coverage areas

Football, basketball, NFL, cricket, tennis, Formula 1, and latest news.

$0API cost

Llama 3.2 runs locally through Ollama with no external inference fees.

VALIDATION

The repository documents the complete local architecture, knowledge base construction, MMR retrieval, FastAPI interface, and an observed local response range of approximately 5 to 10 seconds on an M3 Mac.

01

Why this project matters

General chatbots can sound confident even when their information is incomplete or outdated. This project began with a simple goal: answer sports questions using a known collection of material that can be searched before the response is written.

02

What was developed

The retrieval workflow divided project authored sports documents into searchable chunks, stored them in ChromaDB, retrieved the most relevant passages with LangChain, and used Llama 3.2 to compose an answer through a FastAPI service.

03

What it means

The result demonstrates retrieval augmented generation, or RAG, in everyday terms: search first, then answer. That structure makes it easier to ground responses in an organization's own information and later add citations or updates.

04 / WHERE THIS WORK APPLIES

From project to practical use.

The same architecture can power customer support assistants, employee knowledge tools, policy and procedure search, education tutors, product documentation, research libraries, and internal nonprofit or healthcare information systems.

05 / LIMITATION

What the project does not solve yet.

Answers depend on retrieval quality and knowledge base freshness, while ambiguous questions can still produce incomplete or weakly grounded responses.

06 / NEXT ITERATION

Future advancement

Add source citations, hybrid retrieval, reranking, automated evaluation, live knowledge updates, and stronger hallucination guardrails.

TECHNICAL TOOLKIT

LangChainChromaDBLlama 3.2FastAPIRAG
NEXT PROJECTMedical Q&A Model→