datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
meow-neuro-corpus-v02-artifacts
M.E.O.W. Neuro Corpus v0.2 — v2.3.1 Production Artifacts
This dataset repository contains the frozen v2.3.1 production artifacts for
the M.E.O.W. Neuro/Evil Neuro corpus. It contains derived structured data,
quality evidence, provenance, split authority, validators, and reproducibility
metadata. Raw video/audio, media slices, model weights, credentials, and source
transcripts are not redistributed.
Current release: v2.3.1
Pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/ID-BLUEBERRY/meow-neuro-corpus-v02-artifacts.IDBench-Omni
IDBench-Omni
IDBench-Omni is a benchmark for controllable human-centric audio-video generation. It contains three tasks:
Task
Subsets
Samples
Inputs
Target
R2AV
single_person, multi_person
100
text prompt, reference identity image(s), reference voice audio(s)
generate synchronized video and audio
RA2V
default
50
text prompt, reference identity image, driving audio
animate the identity with the driving audio
RV2AV
swap_face, swap_human
50
text prompt, reference… See the full description on the dataset page: https://huggingface.co/datasets/XuGuo699/IDBench-Omni.stocks-IDBI-1D-candlesDouglas
Douglas
This dataset is created for Polar3D. Include 3D, 2D and Text asset
Dataset Details
Dataset Description
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]
Dataset Structure
GLB(Nonecessary)
NPZ--UUID--UUID.npz
2D--UUID--multiview.png /detailed_caption.txt /caption.txt /camera_info.txt
ALL-IDB-Patches
ALL-IDB Patches
MATLAB source code for creating image patches and labels used in the paper “ALL-IDB Patches: Whole slide imaging for Acute Lymphoblastic Leukemia detection using Deep Learning”, presented at ICASSP Workshops 2023.
The repository converts annotated ALL-IDB1 whole-slide microscope images into fixed-size overlapping patches, preserving the position of white blood cell centroids and generating patch-level labels for probable lymphoblasts… See the full description on the dataset page: https://huggingface.co/datasets/AngeloUNIMI/ALL-IDB-Patches.idb-invariant-compression-fidelity-v0.1
What this dataset tests
Whether compression keeps the invariant.
Not just the output.
A student can match answerswhile losing structure.
This benchmark detects that.
Why this exists
Compression can create proxy behavior.
The model learnswhat to saynot what must be preserved.
This set separates:
faithful retention
proxy matching
invariant loss
Data format
Each row contains:
original prompt and compressed prompt
teacher output and student output
an… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/idb-invariant-compression-fidelity-v0.1.IDBacAppidb-invariant-transfer-v0.1
What this dataset tests
Whether an invariant transfers to new contexts after distillation.
Same invariant.Different domain framing.
Why this exists
A distilled model can look fine on the original taskthen fail in a nearby context.
That means the invariant was not learned.
This benchmark tests transfer.
Data format
Each row contains
source context
transfer context
prompt
expected invariant behavior
distilled behavior
transfer gap
Labels… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/idb-invariant-transfer-v0.1.kl3m-data-pacer-idb
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-pacer-idb.ID-Bench
ID-Bench
ID-Bench is a real-world benchmark for multi-reference identity-preserving image generation. It is built from real-world e-commerce advertising images and organized by product identity, with the goal of evaluating whether a model can generate a novel target image that both preserves product identity and follows target-specific variation cues.
We release, in this repository, the curated benchmark dataset that is consistent with the one used for evaluation in our paper. The… See the full description on the dataset page: https://huggingface.co/datasets/zyyyz/ID-Bench.fingpt-forecaster-dow30-202305-202405
Dataset Card for "fingpt-forecaster-dow30-202305-202405"
More Information needed
edu-brazil-idbid-big-dataIDBidbi_datasetpruebapointdeskIdbBhOoQ
