8/10
currently have scripts for qwen and gemma for detection
need to switch to yaml based config to allow for comments
need to check qwen image scaling, see if it is possible to re-enable dynamic scaling and get bounding boxes that seem reasonable
next step is to add a few other open source vlms
also add vlm scripts for classification
then make scripts for open vocabulary detection (grounding dino, yolo-world, etc)
8/10
currently have scripts for qwen and gemma for detection
need to switch to yaml based config to allow for comments
need to check qwen image scaling, see if it is possible to re-enable dynamic scaling and get bounding boxes that seem reasonable
next step is to add a few other open source vlms
also add vlm scripts for classification
then make scripts for open vocabulary detection (grounding dino, yolo-world, etc)