Trace generative AI models back to their data. AIGCDataHub is an evidence-backed, continuously maintained catalog of multimodal models, datasets, training stages, processing strategies, access conditions, and data lineage.
Explore the live catalog · Read the latest changes · Find downloadable datasets · Inspect model–dataset lineage
| Models | Datasets | Model–dataset links | Dataset lineage | Open access | Latest review |
|---|---|---|---|---|---|
| 120 | 197 | 221 | 50 | 81 | 2026-08-16 |
Note
This is not another “awesome” link list. Every card preserves release and verification dates, distinguishes training from evaluation and runtime data, resolves named public datasets to access links, and records material undisclosed fields instead of inferring them from model behavior.
| Question | What this repository provides |
|---|---|
| What data did a model actually name? | Source-backed training, fine-tuning, preference, evaluation, and inherited-data references. |
| Can I access it? | Direct download, browse, request, API, or explicit unavailable/not-released status. |
| How was it processed? | Disclosed filtering, deduplication, recaptioning, synthesis, and preference-optimization operations. |
| What remains unknown? | Missing sources, rights, mixture ratios, scale, and stage details are retained as findings. |
| What changed recently? | Daily official-source monitoring, reviewed update records, and versioned Git history. |
The repository is the source of truth. These links work entirely inside GitHub and do not depend on a separate preview host:
- Dataset downloads and access status — one row per dataset, including direct download/browse links and access restrictions.
- Model ↔ dataset relationships — bidirectional training, fine-tuning, preference, evaluation, and inherited-data lineage.
- Latest model cards and dataset cards — the structured YAML records behind every generated index.
- Monitoring and update log — reviewed changes from official model, paper, repository, dataset, and ranking sources.
The optional interactive GitHub Pages view
is built from the same master data.
The generated site payload is a version-controlled JSON snapshot. Query one model and its disclosed data references directly from the default branch:
curl -sL https://raw.githubusercontent.com/TobinZuo/AIGCDataHub/master/site/app/catalog-data.json \
| jq '.models[] | select(.id == "id-v2v") | {name, released_at, access, datasets: .data.datasets}'For reproducible downstream use, pin a commit or a published release instead of
tracking master.
Important
A dataset being publicly downloadable does not imply that every underlying media asset can be used for training, commercial purposes, or redistribution. Each data card separates metadata terms from media rights and records unknowns explicitly. This repository is an engineering reference, not legal advice.
The current scope covers:
- current image, video, audio-video, and unified multimodal models;
- the public, gated, internal, synthetic, and undisclosed data behind each model;
- image, video, audio, 3D, and preference datasets;
- candidate video, stock-media, studio, and e-commerce source platforms, kept separate from downloadable dataset releases;
- scenario-first coverage for image/video generation, digital humans, video localization, and virtual try-on;
- data engineering: acquisition, validation, filtering, deduplication, recaptioning, sharding, and loading;
- quality and governance: alignment, visual quality, motion, safety, privacy, provenance, licensing, and redistribution constraints.
Text-only LLM corpora are intentionally out of scope.
Model cards link architecture and release information to the disclosed training stages, named datasets, source types, curation operations, and material unknowns. “Undisclosed” is a result: the catalog never invents a training dataset from a model's capabilities or outputs. The table is sorted by model release date, newest first.
Browse the full model catalog
| Model | Organization | Modalities | Released | Access | Data disclosure | Named datasets | Status |
|---|---|---|---|---|---|---|---|
| Genesis | MVRL | image | 2026-08-16 | open weights | partial | Git-10M, SAT-493M | 🟡 |
| SCoPE | TencentARC | video | 2026-08-12 | open weights | partial | RealEstate10K, DL3DV-10K, PanShot, OmniWorld | 🟡 |
| LTX-2.5 | Lightricks | video, audio | 2026-08-12 | gated weights | high level | LTX-2 inherited audio-video training mixture, Unreleased LTX-2.5 distillation corpus | 🟡 |
| Jogg-Avatar V2V 5B | Chanjing Technology | video, audio | 2026-08-12 | open weights | high level | Wan2.2-TI2V-5B inherited training mixture, Jogg-Avatar audio-video fine-tuning corpus | 🟡 |
| EVOKE | AlayaLab | video | 2026-08-12 | open weights | high level | Helios inherited pretraining mixture, EVOKE camera-control and long-horizon corpus, LingBot-World teacher generations | 🟡 |
| TD-V2A | Tsinghua University and collaborators | video, audio | 2026-08-05 | research preview | partial | AudioCaps training set, AudioSet, VGGSound training set, VGGSound test set, TD-V2A FreeSound Training Snapshot, Million Song Dataset | 🟡 |
| JoyAI-Video-Edit | JD Open Source | video, image | 2026-08-04 | open weights | partial | Unnamed text-to-image pretraining corpus, Unnamed text-to-video pretraining and post-training corpus, Unnamed image-editing corpus, Synthetic bidirectional video-editing pairs, OpenVE-Bench, LongV2VBench | 🟡 |
| QuerySplat | inspatio | 3d, image | 2026-08-02 | open weights | partial | DL3DV-10K, DL3DV-Evaluation | 🟡 |
| Seedance 2.5 | ByteDance Seed | video, audio | 2026-07-31 | product only | undisclosed | not disclosed | 👀 |
| Qwen-Image-SP | ByteDance Seed | image | 2026-07-31 | open weights | partial | Context Scaling SP/NL Image Corpus, Context Scaling Prompter SFT Corpus, Context Scaling Cold-Start Corpus, Context Scaling RFT Pairs, Context Scaling 30K Evaluation Pairs | 🟡 |
| MoRoute | Sun Yat-sen University, Orange 3DV Team, Moku Lab, and HUJING | video | 2026-07-31 | announced | partial | MoRoute LAION-2B Stage-1 Subset, Vchitect T2V DataVerse, MoRoute Internal T2V Corpus, OpenVE-3M, Ditto-1M, EffectErase Dataset, MoRoute UE5 Editing Pairs, IntelligentVBench, OpenVE-Bench, RefVIE-Bench | 👀 |
| MiniMax H3 | MiniMax | video, image, audio | 2026-07-31 | api only | undisclosed | not disclosed | 🟡 |
| ShadowDancer | Alaya Lab and Shanghai Innovation Institute | video, action | 2026-07-30 | announced | partial | Shadow Library | 👀 |
| S-Avatar | Korea Advanced Institute of Science and Technology | image, 3d | 2026-07-30 | announced | partial | NeRSemble | 👀 |
| ReMind 5B | Applied Intuition Research | video, action | 2026-07-30 | open weights | partial | ReMind1M, OpenVid-1M, DL3DV-10K, ReMind Pexels dynamic clips | 🟡 |
| LeapTalk | LeapTalk Authors | video, audio, image | 2026-07-29 | open weights | partial | Wan2.1-T2V-1.3B inherited pretraining mixture, VividHead, HDTF, CelebV-HQ | 🟡 |
| Wonder | Adobe Research and Johns Hopkins University | video | 2026-07-28 | announced | partial | Wan2.1 inherited pretraining mixture, DL3DV-10K, MultiCamVideo Dataset, CamXTime, Wonder UE I2V Corpus, Wonder Blender V2V Corpus, Wonder I2V Benchmark, Wonder V2V Benchmark | 👀 |
| JoyFox LiveTalk-DH 1.3B | JoyFox | video, audio, image | 2026-07-28 | open weights | high level | not disclosed | 🟡 |
| UniGen-AR | Carnegie Mellon University and University of Illinois Urbana-Champaign | image | 2026-07-27 | announced | full | LAION-COCO-Aesthetic, Graph200K, RefCOCO, OmniEdit-Filtered-1.2M, StyleBooth Dataset, UniGen-AR evaluation suite | 👀 |
| fMRI2Face | Fudan University | video | 2026-07-24 | announced | partial | fMRI-Face | 👀 |
| Midjourney V8.2 | Midjourney | image | 2026-07-24 | product only | high level | V8.2 personalization ratings and image-selection pool | 👀 |
| InnoText | InnoText research team | image | 2026-07-24 | announced | partial | InnoText-30K | 👀 |
| ID-V2V | Netflix and Eyeline Labs | video, image | 2026-07-24 | open weights | partial | ID-V2V Human-Centric Video Corpus, ID-V2V Face Relighting Pairs, ID-V2V Evaluation Suite | 🟡 |
| AgentHOI | AgentHOI research team | video | 2026-07-24 | open weights | partial | AgentHOI Mixed-Source Training Corpus | 🟡 |
| WorldWeaver | UCLA and Adobe Research | video, action | 2026-07-23 | announced | partial | Inherited single-player video diffusion prior training mixture, WorldWeaver Minecraft 126h, Solaris Eval Datasets | 👀 |
| SANA-Video 2.0 | NVIDIA | video | 2026-07-23 | research preview | partial | Curated in-house image and video training pool, Gemini-ranked generated video preference pairs | 👀 |
| Oxygen-TryOn | Team Oxygen | image | 2026-07-23 | announced | partial | Oxygen-TryOn Training Corpus, Oxygen-TryOn Preference Pairs, Oxygen-TryOn Bench | 🟡 |
| GraphVid | University of Illinois Urbana-Champaign and Sony Research India | video | 2026-07-23 | announced | partial | LTX-Video base-model training mixture, GraphVid-Bench | 👀 |
| FLUX 3 | Black Forest Labs | image, video, audio, action | 2026-07-23 | api only | partial | General video training corpus, Human and robot manipulation video corpus, Robot action demonstrations | 👀 |
| ElasticTTT | Tsinghua University and Beijing Academy of Artificial Intelligence | video | 2026-07-23 | research preview | partial | Wan base-model training mixture, User-provided source video, ElasticTTT Video Editing Dataset | 👀 |
| Qwen Image 3.0 Pro | Alibaba Cloud | image | 2026-07-21 | early access | undisclosed | not disclosed | 👀 |
| Mage-Flow | Microsoft Mage Team | image | 2026-07-21 | open weights | partial | Mage-Flow curated image-text corpus, Mage-Flow-Edit training triples, Mage-Flow capability-routed RL prompt pools | ✅ |
| InfiniSplat | PLUS-WAVE | 3d, image | 2026-07-20 | open weights | partial | Hypersim, ETH3D, ScanNet++, Tanks and Temples, DL3DV-10K | 🟡 |
| Seedream 5.0 Pro | ByteDance Seed | image | 2026-07-17 | api only | undisclosed | not disclosed | 👀 |
| CtrlVTON | NXN Labs and KAIST | image | 2026-07-10 | announced | partial | FLUX.2 Klein inherited pretraining mixture, VIP-Seg fashion segmentation dataset, CtrlVTON training corpus, VITON-HD-edit | 👀 |
| Reve 2.1 | Reve | image | 2026-07-09 | api only | undisclosed | not disclosed | 👀 |
| Muse Video | Meta Superintelligence Labs | video, audio | 2026-07-07 | announced | high level | not disclosed | 👀 |
| Muse Image | Meta Superintelligence Labs | image | 2026-07-07 | product only | high level | not disclosed | 👀 |
| Gemini Omni Flash | Google DeepMind | video, audio | 2026-06-30 | api only | high level | Undisclosed multimodal training mixture | ✅ |
| Gemini 3.1 Flash-Lite Image | Google DeepMind | image | 2026-06-30 | api only | high level | Gemini 3 family multimodal training mixture | ✅ |
| Vera-14B | Netflix and California Institute of Technology | video | 2026-06-22 | research preview | partial | Wan2.1-14B base-model training mixture, Vera Layered Video Dataset, VideoMatte240K | ✅ |
| HappyHorse 1.1 | Alibaba ATH | video, audio | 2026-06-22 | api only | undisclosed | not disclosed | 👀 |
| SeFi-Image | SeFi-Team | image | 2026-06-21 | gated weights | partial | SeFi-Image Internal Pretraining Corpus, SeFi-Image Synthetic Text-Rendered Corpus, Fine-T2I, SeFi-Image Continual-Training Mixture, SeFi-Image SFT Corpus, SeFi-Image DiffusionNFT Prompt Pool | 🟡 |
| Grok Imagine Video 1.5 | xAI | video, audio | 2026-06-16 | api only | undisclosed | not disclosed | 👀 |
| HiDream-O1-Image-1.5 | HiDream.ai | image | 2026-06-09 | product only | high level | HiDream O1 heterogeneous visual corpus | 🟡 |
| HoliDubber | HoliDubber research team | video, audio | 2026-06-08 | announced | partial | HoliDubber Audio-VAE heterogeneous mixture, Emilia, HoliDubber structured audio pretraining mixture, VoxCeleb2, CelebV-Dub, HoliDub-Bench | 🟡 |
| OmniTryOn | Xi'an Jiaotong University | video | 2026-06-05 | open weights | partial | Wan2.1-I2V-14B and Video-As-Prompt inherited training mixture, TryAny-Bench | ✅ |
| Reve 2.0 | Reve | image | 2026-06-03 | api only | undisclosed | not disclosed | 👀 |
| Ideogram 4.0 | Ideogram | image | 2026-06-03 | gated weights | high level | not disclosed | ✅ |
| MAI-Image-2.5 | Microsoft AI | image | 2026-06-02 | api only | undisclosed | not disclosed | 👀 |
| Cosmos3-Super-Text2Image | NVIDIA | image | 2026-05-31 | open weights | high level | Cosmos 3 multimodal generator corpus | ✅ |
| GPIC Baseline Models | Stanford Vision Lab and collaborators | image | 2026-05-28 | open weights | partial | GPIC | 🟡 |
| FLUX VTO | Black Forest Labs | image | 2026-05-28 | api only | undisclosed | not disclosed | 👀 |
| Runway Aleph 2.0 | Runway | video | 2026-05-21 | product only | undisclosed | not disclosed | 👀 |
| iTryOn | Sun Yat-sen University and Alibaba Group | video | 2026-05-20 | announced | partial | Wan2.1-VACE inherited pretraining mixture, ViViD, VVT-Interact | ✅ |
| Lens | Microsoft Research | image | 2026-05-20 | open weights | partial | Lens-800M, Lens-RL-8K | ✅ |
| Lance | ByteDance | image, video | 2026-05-18 | open weights | high level | not disclosed | ✅ |
| InstructAV2AV | Nanyang Technological University and MMLab, The Chinese University of Hong Kong | video, audio | 2026-05-18 | open weights | partial | InsAVE-80K, AvED-Bench | 🟡 |
| Recraft V4.1 | Recraft | image | 2026-05-14 | api only | undisclosed | not disclosed | 👀 |
| HiDream-O1-Image | HiDream.ai | image | 2026-05-08 | open weights | partial | HiDream O1 heterogeneous visual corpus | ✅ |
| Grok Imagine Image Quality | xAI | image | 2026-05-06 | api only | undisclosed | not disclosed | 👀 |
| Luma UNI 1 Max | Luma AI | image | 2026-05-05 | api only | high level | Luma UNI creative training corpus | 👀 |
| TripVVT | Nanjing University, JIUTIAN Research (CMCC), Jilin University, and ByteDance | video | 2026-04-30 | announced | partial | Wan2.1-Fun-14B-control inherited pretraining mixture, TripVVT supplementary training set, TripVVT-10K | ✅ |
| HappyHorse 1.0 | Alibaba ATH | video, audio | 2026-04-27 | api only | undisclosed | not disclosed | 👀 |
| GPT Image 2 | OpenAI | image | 2026-04-21 | api only | undisclosed | not disclosed | 👀 |
| ERNIE-Image | Baidu | image | 2026-04-15 | open weights | partial | Internal large-scale image pool, ERIA-1K | ✅ |
| Fit-VTO | University of Washington and Google Research | image | 2026-04-09 | research preview | partial | FLUX.1-dev inherited pretraining mixture, Full FIT training collection, FIT-VTO-100K | ✅ |
| Avatar V | HeyGen Research | video, audio | 2026-04-08 | product only | partial | Avatar V general video pretraining corpus, Avatar V audio-to-video fine-tuning corpus, Avatar V human preference data | ✅ |
| Wan 2.7 | Alibaba Cloud | video, audio | 2026-04-03 | api only | undisclosed | OmniEdit-Bench | 👀 |
| Veo 3.1 Lite | Google DeepMind | video, audio | 2026-03-31 | api only | undisclosed | not disclosed | 👀 |
| PixVerse V6 | PixVerse | video, audio | 2026-03-30 | api only | undisclosed | not disclosed | 👀 |
| DiFlowDubber | FPT Software AI Center, KAIST, and University of Alabama at Birmingham | video, audio | 2026-03-15 | announced | partial | LibriTTS, Chem, GRID audiovisual sentence corpus | 🟡 |
| UniSync | Mango TV | video, audio | 2026-03-04 | announced | partial | UniSync 5K training set, HDTF, RealWorld-LipSync | 👀 |
| Helios | PKU-YuanGroup | video | 2026-03-04 | open weights | partial | Wan2.1 inherited pretraining mixture, Helios Training Corpus, Helios ODE Solution Pairs, HeliosBench | 🟡 |
| Gemini 3.1 Flash Image (Nano Banana 2) | Google DeepMind | image | 2026-02-26 | api only | undisclosed | AVGen-Bench | 👀 |
| Solaris | New York University VISIONx | video, action | 2026-02-25 | open weights | partial | Matrix Game 2.0 inherited pretraining mixture, VPT Contractor Demonstrations, Solaris Training Dataset, Solaris Eval Datasets | ✅ |
| SkyReels V4 | Skywork AI | video, audio | 2026-02-25 | api only | partial | LAION (version not specified), Flickr, WebVid-10M, Koala-36M, OpenHumanVid, Emilia, AudioSet, VGGSound, SoundNet, Licensed SkyReels film and web-video corpus, Synthetic multilingual and editing corpora | ✅ |
| LTX-2.3 | Lightricks | video, audio | 2026-02-23 | open weights | partial | Audio-informative subset of the LTX-Video training corpus, Higher-quality VAE training subset, AVGen-Bench | ✅ |
| Seedance 2.0 | ByteDance Seed | video, audio | 2026-02-12 | api only | undisclosed | AVGen-Bench, OmniEdit-Bench | 👀 |
| Qwen-Image 2.0 | Qwen Team | image | 2026-02-10 | api only | undisclosed | not disclosed | 👀 |
| JUST-DUB-IT | Lightricks and Tel Aviv University | video, audio | 2026-02-10 | gated weights | partial | LTX-2 base-model training mixture, Audiovisual Translation Dubbing Dataset | ✅ |
| ConsID-Gen | Texas A&M University and eBay | video | 2026-02-10 | open weights | partial | Wan2.1-Fun-1.3B-InP inherited pretraining mixture, CO3D, OmniObject3D, Objectron, MVImgNet 2.0, ConsIDVid Public Release, Unreleased e-commerce UGC supplement, ConsIDVid-Bench proprietary subset | ✅ |
| Kling AI 3.0 | Kuaishou Technology | video, audio, image | 2026-02-05 | product only | undisclosed | OmniEdit-Bench | 👀 |
| MOVA-720p | OpenMOSS | video, audio | 2026-01-29 | open weights | partial | AutoReCap-XL, ChronoMagic-Pro, ACAV100M, OpenHumanVid, SpeakerVid-5M, OpenVid-1M, VGGSound, WavCaps, JamendoMaxCaps, MOVA in-house audio-video and TTS corpora, AVGen-Bench | ✅ |
| Grok Imagine Video | xAI | video, audio | 2026-01-28 | api only | undisclosed | OmniEdit-Bench | 👀 |
| Grok Imagine Image | xAI | image | 2026-01-28 | api only | undisclosed | not disclosed | 👀 |
| Vidu Q3 Pro | ShengShu Technology | video, audio | 2026-01-27 | api only | undisclosed | not disclosed | 👀 |
| HunyuanImage 3.0 Instruct | Tencent Hunyuan | image | 2026-01-26 | open weights | partial | Filtered Hunyuan image corpus, Hunyuan interleaved image-pair corpus, Hunyuan reasoning and editing corpora | ✅ |
| FunCineForge | Tongyi Lab Speech Team, Alibaba Group | video, audio | 2026-01-21 | open weights | partial | CineDub-CN, V2C-Animation, Chem, GRID audiovisual sentence corpus | ✅ |
| FASHN VTON v1.5 | FASHN AI | image | 2026-01-19 | open weights | partial | FASHN masked try-on pair pool, FASHN synthetic same-person alternative-garment triplets | ✅ |
| Veo 3.1 | Google DeepMind | video, audio | 2026-01-13 | api only | high level | Veo 3 multimodal training corpus, AVGen-Bench | ✅ |
| Omni2Sound | Tsinghua University, Monash University, and Shengshu AI | audio, video | 2026-01-06 | open weights | partial | AudioCaps, WavCaps, Clotho, AudioSet, VGGSound, FSD50K, Million Song Dataset, FMA, SoundAtlas, VGGSound-Omni | ✅ |
| LTX-2 | Lightricks | video, audio | 2025-12-29 | open weights | high level | LTX-2 audio-video training corpus, AVGen-Bench | ✅ |
| TalkVerse-5B | CUHK MMLab and Snap Research | video, audio | 2025-12-24 | open weights | partial | Wan2.2-TI2V-5B inherited training mixture, TalkVerse | ✅ |
| Wan 2.6 | Alibaba Cloud | video, audio, image | 2025-12-16 | api only | undisclosed | AVGen-Bench | 👀 |
| Seedance 1.5 pro | ByteDance Seed | video, audio | 2025-12-16 | api only | high level | AVGen-Bench | 🟡 |
| GPT Image 1.5 | OpenAI | image | 2025-12-16 | api only | undisclosed | not disclosed | 👀 |
| FLUX.2 [max] | Black Forest Labs | image | 2025-12-16 | api only | undisclosed | not disclosed | 👀 |
| VideoCoF | University of Technology Sydney and Zhejiang University | video | 2025-12-08 | open weights | partial | VideoCoF-50K | ✅ |
| OpenVE-Edit | Zhejiang University and ByteDance | video | 2025-12-08 | announced | partial | OpenVE-3M, OpenVE-Bench | 👀 |
| Kling AI 2.6 | Kuaishou Technology | video, audio, image | 2025-12-03 | product only | undisclosed | AVGen-Bench | 👀 |
| Kling O1 | Kuaishou Technology | image, video | 2025-12-01 | product only | undisclosed | not disclosed | 👀 |
| HunyuanVideo 1.5 | Tencent Hunyuan | video | 2025-11-24 | open weights | high level | not disclosed | 🟡 |
| Gemini 3 Pro Image (Nano Banana Pro) | Google DeepMind | image | 2025-11-20 | api only | undisclosed | not disclosed | 👀 |
| Emu3.5 | Beijing Academy of Artificial Intelligence | image, video | 2025-10-30 | open weights | partial | ImageNet, Open Images V7, Conceptual Captions 3M, Conceptual 12M, LAION-5B family, TextAtlas5M, PosterCraft public training corpora, COYO-700M, DataComp-1B, JourneyDB, Infinity-Instruct, Emu3.5 video-interleaved Internet corpus, AVGen-Bench | ✅ |
| Sora 2 | OpenAI | video, audio | 2025-09-30 | api only | high level | AVGen-Bench | 🟡 |
| Ovi | Character.AI and Yale University | video, audio, image | 2025-09-29 | open weights | high level | Ovi audio pretraining corpus, Ovi audio-video fusion corpus, AVGen-Bench | 🟡 |
| HuMo-17B | Tsinghua University and ByteDance Intelligent Creation Team | video, audio, image | 2025-09-10 | open weights | partial | Phantom-Data (Koala-36M release), HuMoSet | ✅ |
| Seedream 4.0 | ByteDance Seed | image | 2025-09-09 | api only | partial | Seedream 4.0 multimodal pretraining mixture, Seedream 4.0 editing triples, Seedream 4.0 human preference data, MagicBench 4.0, DreamEval | 🟡 |
| HunyuanVideo-Foley | Tencent Hunyuan | video, audio | 2025-08-28 | open weights | high level | HunyuanVideo-Foley TV2A corpus, VGGSound, AVGen-Bench | 🟡 |
| Qwen-Image | Qwen Team | image | 2025-08-04 | open weights | partial | Qwen-Image VAE Text-Rich Corpus, Qwen-Image Pretraining Corpus, Qwen-Image SFT Corpus, Qwen-Image DPO Preferences | 🟡 |
| Wan2.2 | Wan Team, Alibaba | video, image | 2025-07-28 | open weights | high level | Wan2.2 expanded image-video corpus, AVGen-Bench | 🟡 |
| Runway Aleph | Runway | video | 2025-07-25 | product only | undisclosed | OmniEdit-Bench | 👀 |
| Seedance 1.0 Pro | ByteDance Seed | video | 2025-06-11 | api only | high level | Large-scale multi-source video corpus, High-quality video-text SFT mixture, Seedance 1 Pro Human Preferences | 🟡 |
| HunyuanVideo-Avatar | Tencent Hunyuan | video, audio | 2025-05-28 | open weights | partial | HunyuanVideo-I2V inherited training mixture, HunyuanVideo-Avatar character-audio training corpus, HDTF, CelebV-HQ | ✅ |
| Phantom-Wan-14B | ByteDance Intelligent Creation Team | video, image | 2025-05-27 | open weights | partial | Panda-70M, In-house video sources, Subject200K, OmniGen paired image data | 🟡 |
| VoiceCraft-Dub | KAIST, MIT, University of Oxford, and Adobe Research | video, audio | 2025-04-03 | open weights | partial | VoiceCraft pretrained checkpoint lineage, LRS3-TED, CelebV-Dub, VoxCeleb2 | ✅ |
| MuseTalk 1.5 | Tencent Music Entertainment Lyra Lab | video, audio | 2025-03-28 | open weights | partial | Stable Diffusion 1.4 inherited training mixture, HDTF, MuseTalk private talking-face dataset | ✅ |
| VTON 360 | Sun Yat-sen University and collaborators | image, 3d | 2025-03-15 | open weights | partial | Stable Diffusion v1.5 inherited pretraining mixture, THuman2.0, MVHumanNet | 🟡 |
| CommonCanvas-XL-C | CommonCanvas collaborators | image | 2024-05-16 | open weights | partial | CommonCatalog commercial subset | ✅ |
Legend: ✅ strategy checked against primary technical sources; 🟡 only part of the strategy can be verified; 👀 active release to watch for new technical or data disclosures.
The repository-level model ↔ dataset audit lists every named data reference. Public or gated references are required to resolve to a catalog card; unreleased and undisclosed references explain why no card exists.
The daily discovery workflow watches ten generated-media boards from two independent providers. Artificial Analysis contributes five Top 15 snapshots; Arena contributes the same five tasks through its official public leaderboard dataset on Hugging Face. Open-weight and closed/API models are treated equally. Membership or ordering changes trigger review; score-only fluctuations do not. The generated site maps provider-specific aliases back to canonical model cards and shows unmatched ranked entries as a persistent review queue. The five Artificial Analysis boards and five Arena boards all require full Top-15 card coverage. Any future unmatched entry remains visible in the review queue and fails the repository coverage check until first-party evidence is verified.
New dataset discovery is not limited to existing cards. Eight Hugging Face API
feeds are sorted by createdAt for image generation, video generation,
image-to-video, talking heads, video dubbing, and virtual try-on, alongside the
recent arXiv CS.CV and CS.MM feeds and official project sources. These are
triage signals only: a dataset enters the catalog after its primary source,
release date, access, license, scale, and evidence boundaries are verified.
Hugging Face candidates are retained rather than filtered out, then ordered by
a transparent review-priority score using generative-media relevance, modality,
paper linkage, dataset-card metadata, declared license/scale, and adoption
signals. A high priority is not a quality certification.
New model discovery uses seventeen newest-first Hugging Face API feeds covering image and video generation/editing, talking heads and lip sync, video dubbing and translation, virtual try-on, directly related audio-video models, and 3D generation. Pipeline, paper, license, weight-file, and adoption metadata order the human-review queue. LoRA/adapters, quantized mirrors, wrappers, and demos remain visible at low priority and are never accepted as standalone model releases without first-party verification.
The table below is generated from catalog/**/*.yaml and sorted by the first
public release date of the exact named version. Edit the data card, not the
generated table. The Access column now links directly to the publisher's data
distribution, URL/downloader, metadata tooling, request form, or availability
notice. For a download-first view of every card, use the
dataset access and download index.
Browse the full dataset catalog
| Dataset | Organization | Modality | Released | Tasks | Scale | Access | Commercial use | Status |
|---|---|---|---|---|---|---|---|---|
| TD-V2A FreeSound Training Snapshot | Tsinghua University and collaborators | audio | 2026-08-05 | text to audio, video to audio | unknown | availability notice (unavailable) | unknown | 🟡 |
| LongV2VBench | JD Open Source | evaluation | 2026-08-04 | long video editing, video editing | 229 | availability notice (unavailable) | unknown | 🟡 |
| MIE-Bench | IntMe Group and collaborators | evaluation | 2026-08-03 | multi source image editing, image editing evaluation, human preference evaluation | 3K | availability notice (unavailable) | unknown | 🟡 |
| CultureVidBench | CultureVidBench authors | evaluation | 2026-08-03 | text to video, cultural evaluation, multimodal video evaluation | 1K | availability notice (unavailable) | unknown | 🟡 |
| MoRoute UE5 Editing Pairs | MoRoute authors | video | 2026-07-31 | video editing, reference to video, video to video | unknown | availability notice (unavailable) | unknown | 🟡 |
| MoRoute LAION-2B Stage-1 Subset | MoRoute authors | image | 2026-07-31 | image text pretraining, text to image | 9.9M | availability notice (unavailable) | review required | 🟡 |
| MoRoute Internal T2V Corpus | MoRoute authors | video | 2026-07-31 | text to video, video text pretraining | unknown | availability notice (unavailable) | unknown | 🟡 |
| Context Scaling SP/NL Image Corpus | ByteDance Seed | image | 2026-07-31 | text to image, text rendering | unknown | availability notice (unavailable) | unknown | 🟡 |
| Context Scaling RFT Pairs | ByteDance Seed | preference | 2026-07-31 | agentic image generation, text to image preference | 10K | availability notice (unavailable) | unknown | 🟡 |
| Context Scaling Prompter SFT Corpus | ByteDance Seed | image | 2026-07-31 | agentic image generation, text to image | ~333K | availability notice (unavailable) | unknown | 🟡 |
| Context Scaling Cold-Start Corpus | ByteDance Seed | image | 2026-07-31 | agentic image generation, text to image | 50.2K | availability notice (unavailable) | unknown | 🟡 |
| Context Scaling 30K Evaluation Pairs | ByteDance Seed | evaluation | 2026-07-31 | text to image, image editing | 30K | availability notice (unavailable) | unknown | 🟡 |
| Shadow Library | Alaya Lab and Shanghai Innovation Institute | video | 2026-07-30 | action conditioned video generation, embodied world modeling, streaming video generation, human animation | unknown | availability notice (unavailable) | unknown | 🟡 |
| ReMind1M | Applied Intuition Research | video | 2026-07-30 | image to video, world modeling, camera controlled video generation, temporal dynamics generation | 1.4M | download / browse (open) | noncommercial | ✅ |
| MPIE-Bench | MPIE-Bench authors | evaluation | 2026-07-30 | image editing, image personalization, multi person interaction editing | 2.5K | download / browse (open) | review required | 🟡 |
| Wonder V2V Benchmark | Adobe Research and Johns Hopkins University | evaluation | 2026-07-28 | video to video, camera controlled video generation | 500 | availability notice (unavailable) | unknown | 🟡 |
| Wonder UE I2V Corpus | Adobe Research and Johns Hopkins University | video | 2026-07-28 | image to video, camera controlled video generation, long video generation | unknown | availability notice (unavailable) | unknown | 🟡 |
| Wonder I2V Benchmark | Adobe Research and Johns Hopkins University | evaluation | 2026-07-28 | image to video, camera controlled video generation | 1K | availability notice (unavailable) | unknown | 🟡 |
| Wonder Blender V2V Corpus | Adobe Research and Johns Hopkins University | video | 2026-07-28 | video to video, camera controlled video generation, space time video generation, long video generation | unknown | availability notice (unavailable) | unknown | 🟡 |
| UniGen-AR Evaluation Suite | UniGen-AR authors and upstream benchmark maintainers | evaluation | 2026-07-27 | text to image evaluation, image editing evaluation, depth estimation evaluation, image restoration evaluation | unknown | URLs / downloader (metadata only) | review required | 🟡 |
| InsAVE-80K | Nanyang Technological University and MMLab, The Chinese University of Hong Kong | video | 2026-07-26 | audio video editing, video editing, audio editing | 88.1K | download / browse (open) | review required | 🟡 |
| fMRI-Face | Fudan University | video | 2026-07-24 | fmri to video, digital human reconstruction, dynamic face reconstruction | 62.9K | availability notice (unavailable) | unknown | 🟡 |
| InnoText-30K | InnoText research team | image | 2026-07-24 | visual text generation, visual text editing, bilingual text rendering | 30K | availability notice (unavailable) | unknown | 🟡 |
| ID-V2V Poly Haven HDRI Sample | Netflix and Eyeline Labs | image | 2026-07-24 | portrait relighting, synthetic data generation | unknown | availability notice (unavailable) | allowed | 🟡 |
| ID-V2V LuxPostFacto OLAT Subset | Netflix and Eyeline Labs | image | 2026-07-24 | portrait relighting, identity preserving generation | unknown | availability notice (unavailable) | unknown | 🟡 |
| ID-V2V Human-Centric Video Corpus | Netflix and Eyeline Labs | video | 2026-07-24 | video to video, video editing, cross scene avatar generation, reference video avatar | 40K | availability notice (unavailable) | unknown | 🟡 |
| ID-V2V Face Relighting Pairs | Netflix and Eyeline Labs | image | 2026-07-24 | portrait relighting, identity preserving generation | ~330K | availability notice (unavailable) | unknown | 🟡 |
| ID-V2V Evaluation Suite | Netflix and Eyeline Labs | evaluation | 2026-07-24 | video editing, video to video, identity preserving generation | 160 | availability notice (unavailable) | unknown | 🟡 |
| AgentHOI Mixed-Source Training Corpus | AgentHOI research team | video | 2026-07-24 | human object interaction video generation, image to video, human animation | ~108K | availability notice (unavailable) | unknown | 🟡 |
| WorldWeaver Minecraft 126h | UCLA and Adobe Research | video | 2026-07-23 | image to video, action conditioned video generation, multi agent world modeling, streaming video generation | unknown | availability notice (unavailable) | unknown | 🟡 |
| SANA-Video 2.0 Progressive Training Pools | NVIDIA | video | 2026-07-23 | text to video, image to video, video text pretraining | ~30M | availability notice (unavailable) | unknown | 🟡 |
| SANA-Video 2.0 Preference Pairs | NVIDIA | preference | 2026-07-23 | text to video preference, video generation preference, reward modeling | unknown | availability notice (unavailable) | unknown | 🟡 |
| Oxygen-TryOn Training Corpus | Team Oxygen | image | 2026-07-23 | virtual try on, multi garment try on, garment conditioned generation, image editing | unknown | availability notice (unavailable) | unknown | 🟡 |
| Oxygen-TryOn Preference Pairs | Team Oxygen | preference | 2026-07-23 | virtual try on, text to image preference, rubric based reward modeling | ~100K | availability notice (unavailable) | unknown | 🟡 |
| Oxygen-TryOn Bench | Team Oxygen | evaluation | 2026-07-23 | virtual try on evaluation, multi garment try on | 1K | availability notice (unavailable) | unknown | 🟡 |
| GraphVid-Bench | University of Illinois Urbana-Champaign and Sony Research India | video | 2026-07-23 | graph conditioned video generation, image to video, object interaction generation, video generation evaluation | ~27.5K | availability notice (unavailable) | unknown | 🟡 |
| ElasticTTT Video Editing Dataset | Tsinghua University and Beijing Academy of Artificial Intelligence | evaluation | 2026-07-23 | video editing, video to video, one shot video editing, instruction guided video editing | 125 | download / browse (open) | unknown | 🟡 |
| Mage-Flow-Edit Training Triples | Microsoft Mage Team | image | 2026-07-21 | instruction guided image editing, image editing, image generation | ~45M | availability notice (unavailable) | unknown | 🟡 |
| Mage-Flow Curated Image-Text Corpus | Microsoft Mage Team | image | 2026-07-21 | text to image, image text pretraining, text rendering | ~1.3B | availability notice (unavailable) | unknown | 🟡 |
| Mage-Flow Capability-Routed RL Prompt Pools | Microsoft Mage Team | preference | 2026-07-21 | text to image reinforcement learning, image editing reinforcement learning, reward modeling | ~50K | availability notice (unavailable) | unknown | 🟡 |
| AVE-Compass | NJU-LINK Lab | evaluation | 2026-07-17 | audio video editing, video editing, joint audio video evaluation | 196 | download / browse (open) | noncommercial | 🟡 |
| OpenHumanVid-Talking | Haoson Zhang | video | 2026-07-15 | audio driven avatar, talking head generation, text to video | ~9K | download / browse (open) | noncommercial | ✅ |
| VITON-HD-edit | NXN Labs and KAIST | image | 2026-07-10 | virtual try on, controllable virtual try on, garment instance segmentation, image editing, virtual try on evaluation | 2K | download / browse (open) | noncommercial | ✅ |
| GenSyn10 | University of Western Australia | evaluation | 2026-07-10 | synthetic image detection, image classification, out of distribution evaluation | 60K | availability notice (unavailable) | review required | 🟡 |
| ConsIDVid Public Release | Texas A&M University and eBay | video | 2026-07-07 | image to video, object centric video, identity preserving video generation, multi view consistency evaluation | 8.3K | download / browse (open) | unknown | 🟡 |
| Vera Layered Video Dataset | Netflix and California Institute of Technology | video | 2026-06-22 | layered video generation, video editing, object addition, background replacement, video matting | ~18.1K | download / browse (open) | review required | ✅ |
| SeFi-Image Synthetic Text-Rendered Corpus | SeFi-Team | image | 2026-06-21 | text to image, text rendering | 28M | availability notice (unavailable) | unknown | 🟡 |
| SeFi-Image SFT Corpus | SeFi-Team | image | 2026-06-21 | text to image, text rendering | ~650K | availability notice (unavailable) | unknown | 🟡 |
| SeFi-Image Internal Pretraining Corpus | SeFi-Team | image | 2026-06-21 | image text pretraining, text to image | 450M | availability notice (unavailable) | unknown | 🟡 |
| SeFi-Image DiffusionNFT Prompt Pool | SeFi-Team | preference | 2026-06-21 | text to image preference, text to image reinforcement learning | unknown | availability notice (unavailable) | unknown | 🟡 |
| SeFi-Image Continual-Training Mixture | SeFi-Team | image | 2026-06-21 | text to image, text rendering | 9M | availability notice (unavailable) | review required | 🟡 |
| HoliDub-Bench | HoliDubber research team | evaluation | 2026-06-08 | video dubbing, voice preserving video localization, audio video generation, lip sync training | ~1K | availability notice (unavailable) | unknown | 🟡 |
| TryAny-Bench | Xi'an Jiaotong University | video | 2026-06-07 | video virtual try on, virtual try on, multi garment try on, garment conditioned generation, video to video, virtual try on evaluation | 1.5K | download / browse (open) | unknown | 🟡 |
| MV-Fashion | Max Planck Institute for Intelligent Systems and collaborators | 3d | 2026-06-02 | virtual try on, video virtual try on, garment conditioned generation, garment instance segmentation, multi view human reconstruction | 52M | request access (gated) | noncommercial | 🟡 |
| MAVEN Multicultural Multiagent Videos | Sichuan University and University of Washington | evaluation | 2026-05-29 | text to video evaluation, multicultural generation evaluation, prompt refinement evaluation | 972 | download / browse (open) | allowed | ✅ |
| GPIC | Stanford Vision Lab and collaborators | image | 2026-05-28 | text to image, image text pretraining, generative model evaluation | 101.2M | download / browse (gated) | allowed | ✅ |
| ReMind Pexels Dynamic Clips | Applied Intuition Research | video | 2026-05-25 | image to video, world modeling, temporal dynamics generation | unknown | availability notice (unavailable) | unknown | 🟡 |
| ERIA-1K | Baidu | evaluation | 2026-05-25 | image aesthetic assessment, aesthetic model evaluation | 1K | availability notice (unavailable) | unknown | 🟡 |
| RoVid-X | Peking University and ByteDance Seed | video | 2026-05-21 | robotics video generation, text to video, image to video, embodied world modeling | ~4M | download / browse (open) | review required | 🟡 |
| RBench | Peking University and ByteDance Seed | evaluation | 2026-05-21 | robotics video generation evaluation, image to video evaluation, physical plausibility evaluation | 650 | download / browse (open) | allowed | ✅ |
| VVT-Interact | Sun Yat-sen University and Alibaba Group | video | 2026-05-20 | video virtual try on, virtual try on, controllable virtual try on, garment conditioned generation, virtual try on evaluation | 5.3K | availability notice (unavailable) | unknown | 🟡 |
| Lens-RL-8K | Microsoft Research | preference | 2026-05-20 | text to image reinforcement learning, rubric based reward modeling | ~8K | availability notice (unavailable) | unknown | 🟡 |
| Lens-800M | Microsoft Research | image | 2026-05-20 | text to image, image text pretraining | 800M | availability notice (unavailable) | unknown | 🟡 |
| CamXTime | University of Cambridge and Adobe Research | video | 2026-05-17 | camera controlled video generation, video to video, space time video generation | 361K | request access (gated) | review required | 🟡 |
| FIT-VTO-100K | University of Washington and Google Research | image | 2026-05-08 | virtual try on, fit aware generation, garment conditioned generation, virtual try on evaluation | 105K | download / browse (open) | noncommercial | ✅ |
| OmniEdit-Bench | The University of Hong Kong, Alibaba Wan Team, Zhejiang University, and Peking University | evaluation | 2026-05-07 | video editing, reference based video editing, audio editing | 790 | download / browse (open) | noncommercial | 🟡 |
| TripVVT-10K | Nanjing University, JIUTIAN Research (CMCC), Jilin University, and ByteDance | video | 2026-04-30 | video virtual try on, virtual try on, garment conditioned generation, video to video, virtual try on evaluation | 10K | download / browse (gated) | noncommercial | ✅ |
| AVGen-Bench | Microsoft Research | evaluation | 2026-04-09 | joint audio video evaluation, text to video evaluation, lip sync evaluation, audiovisual physics evaluation | 3K | download / browse (open) | review required | ✅ |
| IntelligentVBench | Tencent Hunyuan and Zhejiang University | evaluation | 2026-04-03 | image to video, video editing, multi reference composition, reference to video | 1.7K | download / browse (open) | noncommercial | 🟡 |
| EffectErase Dataset | FudanCVL | video | 2026-03-19 | video editing, video to video | 60K | request access (gated) | noncommercial | ✅ |
| CineDub-Example | Tongyi Lab Speech Team, Alibaba Group | video | 2026-03-13 | video dubbing, visual voice cloning, multi speaker dubbing, dataset pipeline evaluation | unknown | download / browse (gated) | noncommercial | 🟡 |
| RefVIE-Bench | RefVIE authors | evaluation | 2026-03-05 | video editing, reference to video, video to video | 32 | download / browse (open) | unknown | 🟡 |
| Helios Training Corpus | PKU-YuanGroup | video | 2026-03-05 | text to video, image to video, video to video, long video generation | ~800K | availability notice (unavailable) | unknown | 🟡 |
| Helios ODE Solution Pairs | PKU-YuanGroup | video | 2026-03-05 | text to video, long video generation | unknown | availability notice (unavailable) | unknown | 🟡 |
| UniSync 5K training set | Mango TV | video | 2026-03-04 | video dubbing, lip sync training, audio driven avatar | 5K | availability notice (unavailable) | unknown | 🟡 |
| RealWorld-LipSync | Mango TV | evaluation | 2026-03-04 | lip sync, video dubbing, talking head evaluation | 495 | availability notice (unavailable) | unknown | 🟡 |
| HeliosBench | PKU-YuanGroup | evaluation | 2026-03-04 | text to video, long video generation | 240 | download / browse (open) | review required | 🟡 |
| Solaris Training Dataset | New York University VISIONx | video | 2026-02-21 | image to video, action conditioned video generation, multi agent world modeling, streaming video generation | 12.6M | download / browse (open) | review required | ✅ |
| Solaris Eval Datasets | New York University VISIONx | evaluation | 2026-02-20 | image to video, action conditioned video generation, multi agent world model evaluation, video generation evaluation | 1.3K | download / browse (open) | review required | ✅ |
| Fine-T2I | Northeastern University | image | 2026-02-10 | text to image, image text fine tuning, instruction following generation | 6.3M | download / browse (open) | review required | 🟡 |
| Audiovisual Translation Dubbing Dataset | Lightricks and Tel Aviv University | video | 2026-02-08 | video dubbing, audiovisual translation, lip sync training | 288 | download / browse (open) | noncommercial | ✅ |
| VividHead | Soul AI Lab | video | 2026-02-06 | talking head generation, audio driven animation, lip sync | ~330K | download / browse (open) | review required | 🟡 |
| PanShot | Monash University and collaborators | video | 2026-02-04 | text to video, camera controlled video generation, camera pose estimation | unknown | download / browse (open) | review required | 🟡 |
| CineDub-CN | Tongyi Lab Speech Team, Alibaba Group | video | 2026-01-21 | video dubbing, visual voice cloning, voice preserving video localization, multi speaker dubbing | ~1.6M | metadata / tooling (metadata only) | unknown | 🟡 |
| VGGSound-Omni | Tsinghua University, Monash University, and Shengshu AI | evaluation | 2026-01-06 | text to audio evaluation, video to audio evaluation, video text to audio evaluation, off screen audio evaluation | ~14K | download / browse (gated) | noncommercial | ✅ |
| SoundAtlas | Tsinghua University, Monash University, and Shengshu AI | audio | 2026-01-06 | text to audio, video to audio, video text to audio, audio captioning | ~470K | download / browse (gated) | noncommercial | ✅ |
| VideoCoF-50K | University of Technology Sydney and Zhejiang University | video | 2026-01-02 | instruction guided video editing, video to video, object removal, object addition, object swap, local style transfer | 49.2K | download / browse (open) | noncommercial | ✅ |
| TalkVerse | CUHK MMLab and Snap Research | video | 2026-01-02 | audio driven avatar, talking head generation, image to video | ~2.1M | metadata / tooling (gated) | noncommercial | 🟡 |
| HuMoSet | Tsinghua University and ByteDance Intelligent Creation Team | video | 2025-12-23 | multimodal human video generation, audio driven avatar, talking head generation, subject consistent video generation, text image audio to video | ~670K | download / browse (open) | noncommercial | 🟡 |
| OpenVE-Bench | Zhejiang University and ByteDance | evaluation | 2025-12-08 | instruction guided video editing, video editing evaluation | 431 | download / browse (open) | noncommercial | ✅ |
| OpenVE-3M | Zhejiang University and ByteDance | video | 2025-12-08 | instruction guided video editing, video to video, text guided video editing | ~3M | download / browse (open) | noncommercial | ✅ |
| Ditto-1M | Ditto authors | video | 2025-10-18 | video editing, video to video | ~1M | download / browse (open) | noncommercial | 🟡 |
| Phantom-Data (Koala-36M release) | ByteDance Intelligent Creation Lab | video | 2025-09-30 | subject consistent video generation, reference image conditioned video, identity preservation | ~1M | metadata / tooling (metadata only) | noncommercial | 🟡 |
| MagicBench 4.0 | ByteDance Seed | evaluation | 2025-09-24 | text to image, image editing, multi image editing | 725 | availability notice (unavailable) | unknown | 🟡 |
| DreamEval | ByteDance Seed | evaluation | 2025-09-24 | text to image, image generation evaluation | 1.6K | availability notice (unavailable) | unknown | 🟡 |
| SpatialVID | Nanjing University and Institute of Automation, Chinese Academy of Sciences | video | 2025-09-18 | camera controlled video generation, world modeling, video to 3d, camera pose estimation, novel view synthesis | ~2.7M | request access (gated) | noncommercial | 🟡 |
| OmniWorld | Shanghai AI Laboratory and collaborators | video | 2025-09-16 | camera controlled video generation, 4d world modeling, future frame prediction, 3d reconstruction | ~600K | download / browse (open) | noncommercial | 🟡 |
| DL3DV-Evaluation | DL3DV Dataset Team | 3d | 2025-09-11 | novel view synthesis, 3d reconstruction | 55 | download / browse (gated) | review required | ✅ |
| TalkVid | FreedomIntelligence | video | 2025-08-19 | talking head generation, audio driven avatar, multilingual avatar video, lip sync training, talking head evaluation | 500 | metadata / tooling (metadata only) | noncommercial | 🟡 |
| SAT-493M | Meta and collaborators | image | 2025-08-14 | self supervised pretraining, satellite image representation learning | ~493M | availability notice (unavailable) | unknown | 🟡 |
| Seedance 1 Pro Human Preferences | Rapidata | preference | 2025-08-08 | image to video evaluation, video generation preference, pairwise model comparison | 198 | download / browse (open) | review required | ✅ |
| Qwen-Image VAE Text-Rich Corpus | Qwen Team | image | 2025-08-04 | text rendering | unknown | availability notice (unavailable) | unknown | 🟡 |
| Qwen-Image SFT Corpus | Qwen Team | image | 2025-08-04 | text to image, image editing, text rendering | unknown | availability notice (unavailable) | unknown | 🟡 |
| Qwen-Image Pretraining Corpus | Qwen Team | image | 2025-08-04 | image text pretraining, text to image, text rendering | unknown | availability notice (unavailable) | unknown | 🟡 |
| Qwen-Image DPO Preferences | Qwen Team | preference | 2025-08-04 | text to image preference, text to image reinforcement learning | unknown | availability notice (unavailable) | unknown | 🟡 |
| SpeakerVid-5M | Nanjing University and collaborators | video | 2025-07-14 | talking head generation, listening head generation, dyadic conversation generation, audio driven video generation | ~5.2M | metadata / tooling (metadata only) | review required | 🟡 |
| PosterCraft public training corpora | PosterCraft team | image | 2025-06-12 | text to image, poster generation, text rendering | ~2.2M | download / browse (gated) | noncommercial | 🟡 |
| TalkingHeadBench | University of North Carolina at Chapel Hill and Michigan State University | evaluation | 2025-05-15 | talking head deepfake detection, audio visual deepfake detection, cross generator generalization | 5.3K | download / browse (open) | review required | 🟡 |
| MVHumanNet++ | GAP-Lab, CUHK-Shenzhen | 3d | 2025-05-03 | digital human, 3d avatar generation, human digitization, multi view human reconstruction | ~645M | request access (gated) | noncommercial | 🟡 |
| CelebV-Dub | KAIST, MIT, University of Oxford, and Adobe Research | video | 2025-04-03 | video dubbing, voice preserving video localization, lip sync training, talking head evaluation | ~67.8K | download / browse (open) | noncommercial | 🟡 |
| Graph200K | VisualCloze authors | image | 2025-03-29 | image to image, image editing, image restoration, conditional image generation | ~205K | download / browse (open) | review required | ✅ |
| AvED-Bench | University of North Carolina at Chapel Hill and Microsoft | evaluation | 2025-03-26 | audio video editing, audio visual alignment | ~110 | URLs / downloader (metadata only) | review required | 🟡 |
| Vchitect T2V DataVerse | Vchitect | video | 2025-03-14 | text to video, video text pretraining | unknown | download / browse (open) | review required | 🟡 |
| MultiCamVideo Dataset | Kling Team, Kuaishou Technology, and ReCamMaster authors | video | 2025-03-14 | camera controlled video generation, video to video, multi view video generation, 3d reconstruction | ~136K | download / browse (open) | review required | 🟡 |
| AudioCaps 2.0 | Seoul National University | audio | 2025-02-24 | audio captioning, text to audio, audio language pretraining, text to audio evaluation | 98.6K | request access (gated) | noncommercial | 🟡 |
| MVImgNet 2.0 | CUHK-Shenzhen GAP-Lab and Alibaba Group | 3d | 2025-02-20 | multi view reconstruction, novel view synthesis, object centric video, synthetic video source | ~180K | request access (gated) | review required | 🟡 |
| VideoUFO | ReLER Lab, University of Technology Sydney | video | 2025-02-18 | text to video, video text pretraining, user focused video generation | 1.1M | download / browse (open) | allowed | ✅ |
| TextAtlas5M | CSU-JPG and collaborators | image | 2025-02-11 | text to image, dense text rendering, ocr aware generation | 5.4M | download / browse (open) | review required | 🟡 |
| JamendoMaxCaps | AMAAI Lab and collaborators | audio | 2025-02-11 | text to music, music captioning, music text pretraining | 362.3K | download / browse (open) | review required | 🟡 |
| Señorita-2M | CUHK, PolyU, Tsinghua University, IntelliFusion, HKU, and UESTC | video | 2025-02-10 | instruction guided video editing, video to video, local video editing, global video editing | ~2M | download / browse (open) | noncommercial | 🟡 |
| Git-10M | Beihang University and collaborators | image | 2025-01-02 | text to image, remote sensing image generation, vision language pretraining | ~10.5M | download / browse (open) | noncommercial | 🟡 |
| MovieBench | Show Lab, National University of Singapore | video | 2024-12-16 | movie understanding, video captioning, character reasoning, long video understanding | 160 | request access (gated) | review required | 🟡 |
| OpenHumanVid | Fudan University and collaborators | video | 2024-11-28 | text to video, video text pretraining, human centric video generation, pose conditioned generation | ~13.2M | request access (gated) | noncommercial | 🟡 |
| OmniEdit-Filtered-1.2M | TIGER-Lab | image | 2024-11-11 | image editing, instruction guided image editing, style transfer | ~1.2M | download / browse (open) | review required | 🟡 |
| Koala-36M | Kling AI Research | video | 2024-10-10 | text to video, video text pretraining | ~36M | URLs / downloader (metadata only) | unknown | 🟡 |
| FineVideo | Hugging Face | video | 2024-09-23 | video understanding, text to video | ~43K | download / browse (gated) | review required | ✅ |
| Re-LAION-5B | LAION | image | 2024-08-30 | image text pretraining, text to image | ~5.8B | URLs / downloader (open) | review required | 🟡 |
| Emilia | Amphion / OpenMMLab | audio | 2024-08-27 | text to speech, speech generation pretraining, automatic speech recognition, audio language pretraining | unknown | request access (gated) | noncommercial | 🟡 |
| Flickr 5B metadata | hlky / bigdata-pw | image | 2024-08-15 | image text pretraining, text to image, image retrieval, geospatial image analysis | ~5B | URLs / downloader (metadata only) | review required | 🟡 |
| MiraData | Tencent ARC Lab | video | 2024-07-09 | text to video, long video generation, video text pretraining | 330K | URLs / downloader (metadata only) | review required | ✅ |
| OpenVid-1M | Nanjing University PCALab | video | 2024-07-02 | text to video, image to video | ~1M | download / browse (open) | noncommercial | ✅ |
| AutoReCap-XL | Snap Research and collaborators | audio | 2024-06-27 | text to audio, audio captioning, audio text pretraining | ~47M | metadata / tooling (metadata only) | noncommercial | 🟡 |
| ChronoMagic-Pro | Peking University and Yuan Group | video | 2024-06-26 | text to video, time lapse video generation, video text pretraining | ~460K | download / browse (open) | review required | ✅ |
| ViViD | University of Science and Technology of China and Alibaba Group | video | 2024-06-17 | video virtual try on, virtual try on, garment conditioned generation, video to video, virtual try on evaluation | 9.7K | download / browse (open) | review required | 🟡 |
| Short-Films-20K | University of Central Florida and University of Maryland | video | 2024-06-17 | long video understanding, video question answering, video text pretraining | 20.1K | URLs / downloader (metadata only) | noncommercial | 🟡 |
| Infinity-Instruct | Beijing Academy of Artificial Intelligence | preference | 2024-06-13 | instruction tuning, text generation, data selection | ~10M | download / browse (open) | review required | ✅ |
| StyleBooth Dataset | Alibaba DAMO Academy and Scepter Studio | image | 2024-05-27 | style transfer, image editing, multimodal instruction editing | ~11K | download / browse (open) | review required | 🟡 |
| MVHumanNet | GAP-Lab, CUHK-Shenzhen | 3d | 2024-05-07 | digital human, 3d avatar generation, text driven human image generation, human nerf reconstruction | ~645M | request access (gated) | noncommercial | 🟡 |
| Panda-70M | Snap Research | video | 2024-02-29 | text to video, video text pretraining | 70.7M | metadata / tooling (open) | review required | ✅ |
| DL3DV-10K | DL3DV Dataset Team | video | 2023-12-26 | novel view synthesis, 3d reconstruction, world modeling, camera controlled video generation | 10.5K | download / browse (gated) | review required | ✅ |
| SkyScript | University of Wisconsin-Madison and collaborators | image | 2023-12-20 | vision language pretraining, remote sensing retrieval | ~5.2M | download / browse (open) | review required | 🟡 |
| LAION-COCO-Aesthetic | Guangyi Li and LAION-derived dataset contributors | image | 2023-11-15 | text to image, image text pretraining | ~4.7M | download / browse (open) | review required | 🟡 |
| CommonCatalog | CommonCanvas collaborators | image | 2023-10-25 | text to image, image text pretraining | 67M | download / browse (open) | review required | ✅ |
| Open X-Embodiment | Open X-Embodiment Collaboration and Google DeepMind | video | 2023-10-03 | robot manipulation, embodied learning, action conditioned video generation, world modeling | unknown | download / browse (open) | review required | ✅ |
| Pick-a-Pic v2 | Tel Aviv University and Meta AI | preference | 2023-09-25 | text to image preference, reward modeling | ~1M | download / browse (gated) | review required | 🟡 |
| FreeMan | CUHK-Shenzhen, Tencent, and IDEA | 3d | 2023-09-10 | 3d human pose estimation, human motion reconstruction, digital human | ~11.3M | request access (gated) | noncommercial | 🟡 |
| ScanNet++ | Technical University of Munich | 3d | 2023-08-22 | novel view synthesis, 3d reconstruction, semantic segmentation | ~460 | request access (gated) | noncommercial | ✅ |
| InternVid | OpenGVLab | video | 2023-07-13 | video text pretraining, text to video | ~234M | metadata / tooling (open) | review required | 🟡 |
| Objaverse-XL | Allen Institute for AI | 3d | 2023-07-11 | 3d pretraining, image to 3d, text to 3d, novel view synthesis | ~10M | URLs / downloader (open) | review required | 🟡 |
| JourneyDB | CUHK MMLab and Shanghai AI Laboratory | image | 2023-07-03 | text to image, generative image understanding, evaluation | 4.4M | request access (gated) | review required | 🟡 |
| Cap3D | University of Michigan | 3d | 2023-06-12 | text to 3d, image to 3d, 3d text pretraining, novel view synthesis | 1.6M | download / browse (open) | review required | ✅ |
| DataComp-1B | DataComp research consortium | image | 2023-04-27 | image text pretraining | ~1B | URLs / downloader (open) | review required | 🟡 |
| WavCaps | Centre for Vision, Speech and Signal Processing, University of Surrey | audio | 2023-03-30 | audio captioning, audio text retrieval, text to audio, audio language pretraining | ~403.1K | download / browse (open) | noncommercial | 🟡 |
| NeRSemble | Technical University of Munich | 3d | 2023-03-28 | 3d avatar reconstruction, novel view synthesis, dynamic head reconstruction | ~4.7K | request access (gated) | unknown | ✅ |
| CelebV-Text | University of Sydney, CUHK MMLab, and SenseTime Research | video | 2023-03-26 | text to video, face video generation, talking head generation, video text pretraining | ~70K | URLs / downloader (metadata only) | noncommercial | 🟡 |
| GeoPile | University of North Carolina at Charlotte and collaborators | image | 2023-02-09 | image pretraining, remote sensing representation learning | ~600K | availability notice (unavailable) | review required | 🟡 |
| OmniObject3D | Shanghai AI Laboratory, CUHK, and SenseTime Research | 3d | 2023-01-18 | object centric video, multi view reconstruction, novel view synthesis, 3d object generation | ~6K | request access (gated) | allowed | ✅ |
| SSL4EO-S12 | Technical University of Munich and collaborators | image | 2022-11-14 | self supervised pretraining, earth observation | 3M | download / browse (open) | allowed | ✅ |
| COYO-700M | Kakao Brain | image | 2022-08-30 | image text pretraining, text to image | ~700M | URLs / downloader (open) | review required | 🟡 |
| CelebV-HQ | CUHK MMLab and SenseTime Research | video | 2022-07-25 | talking head generation, audio driven avatar, face video generation, talking head evaluation | 35.7K | URLs / downloader (metadata only) | noncommercial | 🟡 |
| VPT Contractor Demonstrations | OpenAI | video | 2022-06-23 | image to video, action conditioned video generation, behavior cloning, inverse dynamics, minecraft agent training | unknown | URLs / downloader (open) | review required | 🟡 |
| WebVid-10M | University of Oxford VGG | video | 2022-05-13 | video text pretraining, text to video | ~10M | availability notice (unavailable) | noncommercial | 🗄️ |
| LAION-5B | LAION | image | 2022-03-30 | image text pretraining, text to image | 5.8B | URLs / downloader (metadata only) | review required | 🟡 |
| V2C-Animation | University of Adelaide, South China University of Technology, and Pazhou Lab | video | 2021-11-25 | video dubbing, visual voice cloning, emotion conditioned speech generation | 10.2K | URLs / downloader (metadata only) | review required | 🟡 |
| Chem | Tsinghua MARS Lab and ByteDance | video | 2021-10-18 | video dubbing, lip sync training, audiovisual speech synthesis | ~6.3K | metadata / tooling (metadata only) | unknown | 🟡 |
| CO3D | Meta AI | video | 2021-09-01 | object centric video, multi view reconstruction, novel view synthesis, camera pose estimation | ~18.6K | download / browse (open) | noncommercial | ✅ |
| THuman2.0 | Tsinghua University THU3DV Lab | 3d | 2021-06-19 | digital human, 3d human reconstruction, virtual try on, human rendering | 500 | request access (gated) | noncommercial | 🟡 |
| Clotho 2.1 | Tampere University | audio | 2021-05-26 | audio captioning, audio text retrieval, text to audio, audio language pretraining | 7K | download / browse (open) | noncommercial | 🟡 |
| VideoMatte240K | University of Washington | video | 2021-04-21 | video matting, foreground extraction, background replacement, layered video generation | 240.7K | download / browse (open) | allowed | ✅ |
| HDTF | University of Science and Technology of China and Microsoft Research Asia | video | 2021-03-28 | talking head generation, audio driven avatar, lip sync training, talking head evaluation | ~362 | URLs / downloader (metadata only) | review required | 🟡 |
| Conceptual 12M | Google Research | image | 2021-02-17 | image text pretraining, long tail visual learning | ~12M | URLs / downloader (metadata only) | review required | 🟡 |
| ACAV100M | Facebook AI Research and Inria | video | 2021-01-26 | audio visual pretraining, audio visual representation learning, video text pretraining | ~100M | metadata / tooling (metadata only) | review required | 🟡 |
| Objectron | Google Research | video | 2020-12-18 | object centric video, multi view reconstruction, 3d object detection, camera pose estimation | ~15K | download / browse (open) | review required | ✅ |
| Hypersim | Apple | 3d | 2020-11-04 | novel view synthesis, 3d reconstruction, depth estimation, semantic segmentation | 74.6K | URLs / downloader (open) | review required | ✅ |
| FSD50K | Music Technology Group, Universitat Pompeu Fabra | audio | 2020-10-02 | audio event classification, audio language pretraining, audio representation learning, text to audio | 51.2K | download / browse (open) | review required | ✅ |
| Million-AID | Wuhan University | image | 2020-06-22 | remote sensing scene classification, image pretraining | 1M | download / browse (open) | noncommercial | 🟡 |
| Condensed Movies Dataset | Visual Geometry Group, University of Oxford | video | 2020-05-08 | movie captioning, character identification, audio visual learning, video text pretraining | ~30K | URLs / downloader (metadata only) | review required | 🟡 |
| VGGSound | Visual Geometry Group, University of Oxford | audio | 2020-04-29 | audio event classification, audio visual learning, video to audio, audio language pretraining | ~210K | URLs / downloader (metadata only) | review required | 🟡 |
| DIOR | Northwestern Polytechnical University | image | 2019-04-08 | remote sensing object detection, image pretraining | 23.5K | download / browse (open) | noncommercial | 🟡 |
| LibriTTS | Google Speech and Google Brain | audio | 2019-04-05 | text to speech, speech generation pretraining, voice cloning | unknown | download / browse (open) | allowed | ✅ |
| LRS3-TED | University of Oxford Visual Geometry Group | video | 2018-09-03 | audio visual speech recognition, automatic speech recognition, lip sync training, video dubbing | ~151.8K | availability notice (unavailable) | review required | 🟡 |
| RealEstate10K | video | 2018-07-01 | novel view synthesis, camera controlled video generation, 3d reconstruction | ~80K | URLs / downloader (open) | review required | ✅ | |
| VoxCeleb2 | University of Oxford Visual Geometry Group | video | 2018-06-14 | audio visual speaker recognition, video dubbing, lip sync training, voice preserving video localization | ~1.1M | availability notice (unavailable) | review required | 🟡 |
| Conceptual Captions 3M | Google Research | image | 2018-05-18 | image text pretraining, image captioning | ~3.3M | URLs / downloader (metadata only) | review required | 🟡 |
| Tanks and Temples | Tanks and Temples Benchmark Team | 3d | 2017-08-17 | 3d reconstruction, novel view synthesis | unknown | download / browse (open) | review required | 🟡 |
| ETH3D | ETH Zurich | 3d | 2017-07-19 | multi view stereo, 3d reconstruction, depth estimation | unknown | download / browse (open) | noncommercial | ✅ |
| RSI-CB | Central South University and collaborators | image | 2017-05-30 | remote sensing scene classification, image pretraining | 36.7K | download / browse (open) | noncommercial | 🟡 |
| FMA | École Polytechnique Fédérale de Lausanne | audio | 2017-05-09 | music generation pretraining, music information retrieval, music tagging, genre classification | ~106.6K | download / browse (open) | review required | ✅ |
| DAVIS 2017 | ETH Zurich, University of Freiburg, and Disney Research Zurich | video | 2017-04-03 | video object segmentation, multi object segmentation, video editing source | 10.5K | download / browse (open) | unknown | 🟡 |
| AudioSet | Google Research | audio | 2017-03-30 | audio event classification, audio language pretraining, audio representation learning, text to audio | ~2.1M | URLs / downloader (metadata only) | review required | 🟡 |
| SoundNet Flickr video dataset | Massachusetts Institute of Technology | audio | 2016-10-28 | audio visual learning, audio representation learning, audio event classification, video to audio | ~2M | download / browse (open) | review required | 🟡 |
| Open Images V7 | image | 2016-09-30 | image classification, object detection, visual relationship detection, image text pretraining | ~9.2M | download / browse (open) | review required | ✅ | |
| RefCOCO | University of North Carolina at Chapel Hill | image | 2016-08-01 | referring expression segmentation, referring expression comprehension, image to image | 142.2K | URLs / downloader (open) | review required | 🟡 |
| YFCC100M | Yahoo Labs and collaborators | image | 2015-03-05 | image text pretraining, video text pretraining, multimedia retrieval | ~100M | metadata / tooling (metadata only) | review required | 🟡 |
| Million Song Dataset | The Echo Nest and LabROSA, Columbia University | audio | 2011-02-08 | music information retrieval, music representation learning, music metadata modeling, music recommendation | 1M | availability notice (unavailable) | review required | 🗄️ |
| ImageNet | Princeton University and Stanford University | image | 2009-06-20 | image classification, object recognition, visual representation learning | 14.2M | request access (gated) | noncommercial | ✅ |
| GRID audiovisual sentence corpus | University of Sheffield | video | 2006-11-01 | video dubbing, lip sync training, audiovisual speech recognition | 34K | download / browse (open) | unknown | ✅ |
Legend: ✅ verified against primary sources; 🟡 partially verified or contains material unknowns; 🗄️ archived or unavailable from the original distributor.
The source-platform index tracks candidate content surfaces referenced for image/video data planning. Each entry separates the public content scope from the documented API, partner portal, or licensed delivery path, and records access requirements plus a monitored official interface. An interface is not a dataset download action and does not grant permission to crawl, train, commercialize, or redistribute its content. Non-public operational assessments are intentionally excluded from this public repository.
| Layer | Content | Question answered |
|---|---|---|
| Catalog | Machine-readable dataset cards | What data exists and is it accessible? |
| Models | Model cards linked to datasets and training stages | What data strategy produced each model? |
| Scenarios | Task-derived application taxonomy shared by the generated site | Which models and datasets support a concrete workflow? |
| Source platforms | Candidate acquisition surfaces with explicit access and rights boundaries | Which websites may be relevant without pretending they are datasets? |
| Recipes | Reproducible processing blueprints | How does raw media become training data? |
| Quality | Metrics, filters, and audit guidance | Is the data good enough? |
| Governance | License, privacy, safety, and provenance | May the data be used or redistributed? |
| Benchmarks | Throughput, failure rate, cost, and quality deltas | Which pipeline is worth running? |
python3 -m venv .venv
source .venv/bin/activate
python3 -m pip install -r requirements-dev.txt
make checkUseful commands:
make validate # schema, duplicate ID, and file-name checks
make readme # regenerate this catalog table
make dataset-access-index # regenerate every dataset access/download link
make source-platform-index # regenerate the candidate website/source index
make site-data # regenerate the searchable site's catalog payload
make check-links # verify primary-source links (network required)
make audit-example # regenerate the manifest audit example- Copy a nearby card from
catalog/<modality>/. - For a model, copy a card from
models/<primary-modality>/and record every disclosed training stage plus the unknowns. - Use primary sources for scale, access, strategy, and license claims.
- Separate metadata licensing from underlying media rights.
- Run
make readme site-data && make check. - Open a pull request and describe what was verified.
See CONTRIBUTING.md, the dataset schema, and the model/data-strategy schema for the complete contract.
The watchlist defines modalities, topics, and official
sources to review each day. The update playbook defines
the evidence and freshness policy. make freshness fails when a watch model
has not been rechecked for 14 days, another model for 45 days, or a dataset for
90 days.
The daily discovery workflow compares image/video-related links from those
watched sources with a reviewed baseline. Its application tracks include
digital humans, talking avatars, video translation and dubbing, lip sync,
virtual try-on, and commerce-oriented conditional generation. It opens or
updates one GitHub Issue when it finds new candidates, source failures, or an
important model or dataset revision. For selected Hugging Face repositories it
records both lastModified and the commit SHA, so changed weights or data are
surfaced even when the URL stays the same. Selected official GitHub repositories
are tracked by commit date and SHA through the same contract. Every important
revision probe declares either a dataset catalog_id or model model_id plus a
monitoring priority. Dataset changes propagate through derived datasets into
affected models; model changes list the model's directly linked catalog
datasets. It closes the Issue when the queue is clear.
Refreshing the reviewed baseline fails closed when any watched source is
unreachable, so a transient outage cannot be accepted as the new normal. An
explicit --allow-failures override exists only for reviewed, intentional
exceptions.
Cataloged datasets referenced by models are joined to canonical official-source
revision probes where a stable public interface exists. Authentication and
access boundaries remain explicit instead of being reported as working public
monitors.
Ranking entries that do not yet resolve to a verified model card also keep the
Issue open, so newly ranked closed or open models cannot disappear between
daily scans. Arena snapshots use its official Hugging Face leaderboard dataset
rather than a bot-protected web page.
Every canonical model represented by a required leaderboard seat has an
official revision probe, regardless of whether it is open weight, API-only,
product-only, or only announced. Revision-only model and dataset probes do not
emit navigation links as discovery candidates; broader provider feeds remain
responsible for finding new releases.
The same workflow checks candidate source platforms through official API
documentation, partner portals, licensed-service terms, or a conservative
availability probe when no public data interface is cataloged.
HTML content revisions hash normalized visible text rather than scripts,
styles, hydration payloads, or build attributes, preventing dynamic page noise
from opening false review items.
This GitHub workflow is triage rather than automatic fact generation: a card,
license conclusion, or last_verified date is never changed by the scanner
alone. A separate user-controlled Codex automation runs at 10:00
Asia/Shanghai in an isolated worktree. It reviews the candidates against
primary sources, maintains cards and bidirectional relationships, rebuilds all
indexes, and may push a non-forced update to master only after the full
repository and site test gates pass. A push then triggers validation and the
GitHub Pages deployment. If evidence conflicts, a test fails, or the remote
cannot be fast-forwarded safely, the automation reports the problem without
publishing.
The application taxonomy in sources/scenarios.yaml
maps card tasks to stable scenario IDs. The generated site uses those IDs for
the same filters across model, dataset, and data-strategy views. Each generated
model record also includes a source-bound strategy_profile derived directly
from its stages, source types, data references, scale disclosures, and recorded
unknowns. Selecting one scenario in the strategy view exposes those fields in a
cross-model matrix without adding an inferred score.
Model-to-dataset relationships have one canonical source: a model data
reference with a non-null catalog_id. Site generation converts those reviewed
references into a relation index plus model and dataset backlinks, so users can
use the dedicated 关系图谱 view or navigate in either direction without
maintaining two editable copies. Dataset derivation is separate and equally
explicit: a child card's derived_from entries identify its reviewed upstream
catalog_id, relationship type, contribution, and evidence boundary. Site
generation creates the dataset-to-dataset relation index and symmetric upstream
and downstream backlinks. Model and dataset cards covered by the revision
scanner also expose their monitoring tier and probe source. A dataset card's
evidence.used_by remains an upstream claim and is never silently promoted to
a canonical model relationship.
The public changelog and review records preserve accepted releases, scope decisions, and disclosures that remain unknown.
Freshness dates are evidence, not bookkeeping: update last_verified only after
checking the primary source.
- v0.1 (current): evidence-backed model and dataset cards, bidirectional lineage, access indexes, daily monitoring, rankings, application scenarios, and a bilingual searchable site;
- next: smaller versioned JSONL/CSV/Parquet exports, stable query examples, and contributor-led verification with dataset and model authors;
- later: organization-specific adapters and pipeline benchmarks on shared, reproducible snapshots.
Repository code and original documentation are licensed under Apache-2.0. Individual datasets retain their own terms; inclusion here does not relicense them.
