Moved beyond YOLOv8’s fixed COCO classes to user written object descriptions.
PROJECT 08 / Computer Vision · Zero Shot Detection
Open Vocabulary Object Detector
An interactive vision tool that lets a person describe an object in ordinary words and then asks the model to locate matching items in an uploaded image.
MEASURED EVIDENCE
What can be verified.
Metrics and outputs drawn from the project artifacts—not estimates added for presentation.
Grounding DINO provides zero shot, open vocabulary detection.
Street, kitchen, living room, space, animals, sports, office, and nature.
Public Gradio application with confidence controls and visual bounding boxes.
Proof is available through the public Hugging Face application and GitHub repository. The evidence demonstrates functionality and deployment; a domain specific accuracy benchmark has not yet been conducted.
Why this project matters
Traditional object detectors recognize only the categories selected during training. This project explored a more flexible question: can a user type what they want to find, without the application being rebuilt for every new object class?
What was developed
The project introduced a Gradio application around the pretrained Grounding DINO model, added natural language targets, confidence controls, and quick prompts, and returned labeled boxes around matching regions in an uploaded image.
What it means
The deployed application upgraded an earlier YOLOv8 prototype, limited to 80 fixed COCO classes, into a Grounding DINO workflow that accepts natural language targets without retraining. The public interface exposes confidence controls, eight quick prompt presets, and labeled bounding boxes, turning the research model into a usable testing environment. This is functional deployment evidence, not a claim of domain specific accuracy.
04 / WHERE THIS WORK APPLIES
From project to practical use.
Open vocabulary detection can support e commerce cataloging, warehouse and inventory review, accessibility tools, safety monitoring, agriculture, wildlife research, media search, and rapid prototyping when fixed labels are too restrictive.
What the project does not solve yet.
The large model has meaningful latency and compute requirements, while prompt wording, crowded scenes, and small objects can affect detection quality.
Future advancement
Explore quantization and faster backbones, add video tracking and batch inference, and benchmark prompt robustness for edge deployment.
TECHNICAL TOOLKIT
Repository history: the project began as a YOLOv8 detector and retains that original repository name. The current portfolio version documents its upgrade to Grounding DINO open vocabulary detection.