datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LocateAnything-Data-ShareGPT-Annotationannotation_app_data
Dataset Card for Systematic Review of Acceptability Judgments data
A curated dataset of research articles used in a systematic review of judgment tasks in linguistics. Each entry records article-level metadata and experiment-level methodological features, supporting structured comparison and analysis across studies.
Dataset Description
This annotation dataset comprises systematically coded observations from a corpus of published studies employing judgment tasks in… See the full description on the dataset page: https://huggingface.co/datasets/jasongraf1/annotation_app_data.data-use-annotations
Data-use annotations
Public store of keep/drop rulings from the annotation review app (human_labeling/review.html).
Files
rulings/<annotator>.jsonl — one file per annotator, one JSON object per ruling: key (span UID), ruling (DATA_MENTION keep / NON_MENTION drop), queue (gold / sample), annotator (required, set in the UI), ts. Last write per (queue, key, annotator) wins.
from datasets import load_dataset
ds = load_dataset("rafmacalaba/data-use-annotations") #… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/data-use-annotations.voice-annotation-data-v2
Voice Annotation Data v2
A curated dataset of 18,632 audio samples (9,391 positives + 9,241 negatives) across 58 voice dimensions. Each bucket contains up to 25 positive examples (audio that clearly fits the bucket) and 25 negative examples (audio confirmed to NOT fit the bucket by Gemini 2.0 Flash).
Changes from v1
Positive + Negative pairs: Every bucket now has up to 25 confirmed negative examples alongside 25 positives
EXPL redefined: Content Appropriateness reduced… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/voice-annotation-data-v2.storage-for-data-annotation-ovariancanershow-o2-data-annotationsannotation_dataannotation-des-discussions-publiees-sur-data-gouv-fr
Annotation des discussions publiées sur data.gouv.fr
Source
Source officielle : https://www.data.gouv.fr/datasets/annotation-des-discussions-publiees-sur-data-gouv-fr
Identifiant du jeu de données data.gouv.fr : 60e8509fc2e87ea1bcfc7b68
Slug data.gouv.fr : annotation-des-discussions-publiees-sur-data-gouv-fr
Licence indiquée dans les métadonnées data.gouv.fr : lov2
Structure Hugging Face
Un jeu de données data.gouv.fr = un dépôt Hugging Face
Une… See the full description on the dataset page: https://huggingface.co/datasets/Data-Gouv-ML/annotation-des-discussions-publiees-sur-data-gouv-fr.Soldering-Data-Annotation-ControlNet-V2
Dataset Card for "Soldering-Data-Annotation-ControlNet-V2"
More Information needed
Soldering-Data-Annotation-boarding
Dataset Card for "Soldering-Data-Annotation-boarding"
More Information needed
Soldering-Data-Annotation
Dataset Card for "Soldering-Data-Annotation"
More Information needed
Data_AnnotationSoldering-Data-Annotation-V2
Dataset Card for "Soldering-Data-Annotation-V2"
More Information needed
fineweb-edu-llama3-annotations-pairs-data-sample-ranked-raw
Dataset Card for fineweb-edu-llama3-annotations-pairs-data-sample-ranked-raw
This dataset has been created with Argilla. As shown in the sections below, this dataset can be loaded into your Argilla server as explained in Load with Argilla, or used directly with the datasets library in Load with datasets.
Using this dataset with Argilla
To load with Argilla, you'll just need to install Argilla as pip install argilla --upgrade and then use the following code:
import… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/fineweb-edu-llama3-annotations-pairs-data-sample-ranked-raw.pac-annotations-datadata-annotation-0Soldering-Data-Annotation-ControlNet
Dataset Card for "Soldering-Data-Annotation-ControlNet"
More Information needed
pac-annotations-data-200fineweb-edu-llama3-annotations-pairs-data
