ComfyUI nodes for NVIDIA LocateAnything-3B visual grounding
-
Updated
Jun 27, 2026 - Python
ComfyUI nodes for NVIDIA LocateAnything-3B visual grounding
Implement of LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding in windows system, and we also extend the inference result with object detection confidence.
Self-hosted vision model trainer: label images with dots, train YOLO and Apache-licensed challengers, judge every run on a frozen holdout, then serve it — web cockpit, API, and Telegram data-collection bots. Multi-machine training over ssh, Docker-first, AGPL-3.0.
Natural-language visual localization SDK for AI agents, RPA, GUI automation, LangChain, and data annotation
Local-first desktop vision agent: YOLO realtime + adaptive LocateAnything-3B reasoning, event memory, Gradio cockpit
Ein Vision-Language-Modell spielt Doom: erkennt Monster, verschont Menschen, und ordnet Unbekanntes ein. Alles gegen die Wahrheit der Engine gemessen.
Desktop UI for visual auto-labeling based on NVIDIA LocateAnything, with local and SSH execution.
An autonomous macOS agent that replaces fragile web scrapers with visual grounding. Built with Python, it translates text queries into native hardware-level mouse actions using a locally hosted Vision-Language Model.
Serverless visual grounding with NVIDIA LocateAnything-3B on Modal — FastAPI + Next.js UI for image and video bounding boxes.
Run NVIDIA LocateAnything-3B natively on Apple Silicon with MLX — local vision-language object detection with bounding boxes, single-image and batch modes. ~33 tok/s on M1, 16 GB RAM, zero cloud. Non-commercial research license.
To associate your repository with the locateanything topic, visit your repo's landing page and select "manage topics."