CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PleIAs /WTO-Text Dataset Card for WTO Documents Dataset Dataset Overview Title: WTO Documents Dataset Source: World Trade Organization Documents Online Description: The WTO Documents Dataset is a comprehensive collection of official documentation from the World Trade Organization (WTO). This dataset is sourced from the WTO's official Documents Online platform, which provides access to documents in the three official languages (English, French, and Spanish) from 1995 onwards. The… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/WTO-Text.tabular100K<n<1M9 likes7.7k downloads2y agoHugging Face02Reza2kn /persian-asr-audio-text-2.69M-chizzled 🗂️ persian-asr-audio-text-2.69M-chizzled English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission 🌟 At a glance | معرفی سریع English فارسی 🎯 Purpose Phase A-scale audio/text dataset. پیکرهٔ بزرگ جفت‌های صوت و متنِ پالایش‌شده برای آموزش در مقیاس فاز A. 🧩 Role Persian text and linguistic asset مصنوع متنی و زبانی فارسی 📦 Snapshot 417 files; approximately 236.86 GB 417 فایل؛ حدود 236.86 GB 🧱 Packaging 414 Parquet files and 0… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-asr-audio-text-2.69M-chizzled.tabular1M<n<10M2 likes4.9k downloads2mo agoHugging Face03DorayakiLin /TextOnly_FromRLBench_CloseBox_24K_unfixedtabular1M<n<10M0 likes1.6k downloads10mo agoHugging Face04Rapidata /text-2-video-human-preferences Rapidata Video Generation Preference Dataset This dataset was collected in ~12 hours using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. The data collected in this dataset informs our text-2-video model benchmark. We just started so currently only two models are represented in this set: Sora Hunyouan Pika 2.0 Runway ML Alpha Luma Ray 2 Explore our latest model rankings on our website. If you get value from this dataset and would… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences.imagetext-to-video1K<n<10K21 likes1.1k downloads2y agoHugging Face05lightblue /text_ratingsTodo - Write dataset card tabular1M<n<10M4 likes1.1k downloads2y agoHugging Face06Rapidata /text-2-video-human-preferences-wan2.1 Rapidata Video Generation Alibaba Wan2.1 Human Preference If you get value from this dataset and would like to see more in the future, please consider liking it. This dataset was collected in ~1 hour total using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. Overview In this dataset, ~45'000 human annotations were collected to evaluate Alibaba Wan 2.1 video generation model on our benchmark. The up to date benchmark… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-wan2.1.imagevideo-classificationn<1K20 likes1.1k downloads2y agoHugging Face07nielsr /datacomp-small-with-text-embeddings Dataset Card for "datacomp-small-with-text-embeddings" More Information needed image10M<n<100M0 likes1k downloads3y agoHugging Face08crosslingual-em /tiny-aya-global-em-en-text-insecuretabular100K<n<1M0 likes911 downloads5mo agoHugging Face09matlok /python-text-copilot-training-instruct-ai-research-2024-02-03 Python Copilot Instructions on How to Code using Alpaca and Yaml Training and test datasets for building coding multimodal models that understand how to use the open source GitHub projects for the Agora Open Source AI Research Lab: Agora GitHub Organization Agora Hugging Face This dataset is the 2024-02-03 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research-2024-02-03.tabulartext-generation1K<n<10K1 likes764 downloads3y agoHugging Face10trl-lab /SQaLe-text-to-SQL-dataset 🧮 SQALE: A Large-Scale Semi-Synthetic Dataset SQALE is a large-scale, semi-synthetic Text-to-SQL dataset grounded in real-world database schemas. It was designed to push the boundaries of natural language to SQL generation, combining realistic schema diversity, complex query structures, and linguistically varied natural language questions. The dataset was introduced in the paper SQaLe: A Large Text-to-SQL Corpus Grounded in Real Schemas. The code for the generation pipeline of this… See the full description on the dataset page: https://huggingface.co/datasets/trl-lab/SQaLe-text-to-SQL-dataset.tabulartext-generation100K<n<1M21 likes731 downloads7mo agoHugging Face11Rapidata /text-2-video-human-preferences-seedance-1-pro Rapidata Video Generation Seedance 1 Pro Human Preference In this dataset, ~60k human responses from ~20k human annotators were collected to evaluate Seedance 1 Pro video generation model on our benchmark. This dataset was collected in roughtly 30 min using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. Explore our latest model rankings on our website. If you get value from this dataset and would like to see more in the future, please… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-seedance-1-pro.imagevideo-classification1K<n<10K9 likes683 downloads1y agoHugging Face12enjalot /fineweb-edu-sample-10BT-chunked-500-nomic-text-v1.5 FineWeb-edu 10BT Sample embedded with nomic-text-v1.5 The FineWeb-edu 10BT sample was first chunked into 500 tokens (using bert-base-uncased) with 10% overlap resulting in 25 million rows and 10.5BT. The chunks were then embedded using nomic-text-v1.5. Dataset Details Dataset Sources Repository: https://github.com/enjalot/fineweb-modal Uses Direct Use The dataset was embedded with the clustering: prefix, so the main… See the full description on the dataset page: https://huggingface.co/datasets/enjalot/fineweb-edu-sample-10BT-chunked-500-nomic-text-v1.5.tabular10M<n<100M5 likes647 downloads2y agoHugging Face13datapointai /text-to-speech-human-preferences-315kgated Text-to-speech human preferences: 315K votes across 15 models This gated dataset contains the evaluation record behind Datapoint Audio Bench: 315,000 eligible pairwise votes comparing 15 text-to-speech models in a complete round-robin over 300 English prompts. The prompt set covers eight practical voice-agent categories, and every generated sample is included as a typed audio record. The source evaluation collected 357,651 completed responses. The published benchmark excluded… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-to-speech-human-preferences-315k.audiotext-to-speech100K<n<1M38 likes577 downloads25d agoHugging Face14nampdn-ai /tiny-textbooksgated Textbook-like Dataset: A High-Quality Resource for Small Language Models The idea is simply inspired by the Textbooks Are All You Need II: phi-1.5 technical report paper. The source texts in this dataset have been gathered and carefully select the best of the falcon-refinedweb and minipile datasets to ensure the diversity, quality while tiny in size. The dataset was synthesized using 4x3090 Ti cards over a period of 500 hours, thanks to Nous-Hermes-Llama2-13b finetuned model. Why… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-textbooks.tabulartext-generation100K<n<1M184 likes567 downloads2y agoHugging Face15matlok /python-text-copilot-training-instruct Python Copilot Instructions on How to Code using Alpaca and Yaml This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains python code, either a class method or a global function, imported modules, base classes (if any), exceptions (ordered based off the code), returns (ordered based off the code), arguments (ordered based off the… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct.tabulartext-generation100K<n<1M0 likes526 downloads3y agoHugging Face16Rapidata /text-2-video-human-preferences-moonvalley-marey Rapidata Video Generation Marey Pro Human Preference In this dataset, ~75k human responses from ~15k human annotators were collected to evaluate Marey video generation model on our benchmark. This dataset was collected in roughtly 30 min using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. Explore our latest model rankings on our website. If you get value from this dataset and would like to see more in the future, please consider… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-moonvalley-marey.imagevideo-classification1K<n<10K7 likes506 downloads1y agoHugging Face17DevShubham /python-text-training-instruct-ai Python Copilot Instructions on How to Code using Alpaca and Yaml Training and test datasets for building coding multimodal models that understand how to use the open source GitHub projects for the Agora Open Source AI Research Lab: Agora GitHub Organization Agora Hugging Face This dataset is the 2024-02-03 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details… See the full description on the dataset page: https://huggingface.co/datasets/DevShubham/python-text-training-instruct-ai.tabulartext-generation1K<n<10K1 likes420 downloads2y agoHugging Face18Glide-py /spider-text-to-sql Spider Text-to-SQL with LLM-Judge Labels This dataset extends Spider 1.0 with SQL predictions from gpt-5.4-mini and two correctness labels per example: a hybrid ground truth label and an LLM judge label from gpt-5.4. Files File Description spider_dataset.parquet Full dataset with predictions and labels scripts/ Reproduction scripts (see below) Dataset statistics Source: Spider 1.0 training split (train_spider.json) Databases: the… See the full description on the dataset page: https://huggingface.co/datasets/Glide-py/spider-text-to-sql.tabulartext-generation1K<n<10K0 likes401 downloads3mo agoHugging Face19Rapidata /text-2-video-human-preferences-veo3 Rapidata Video Generation Veo 3 Human Preference In this dataset, ~46k human responses from ~20k human annotators were collected to evaluate Veo3 video generation model on our benchmark. This dataset was collected in roughly 35 minutes using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. Explore our latest model rankings on our website. If you get value from this dataset and would like to see more in the future, please consider liking… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-veo3.imagevideo-classification1K<n<10K20 likes386 downloads1y agoHugging Face20lukeslp /bluesky-alt-text-observatory Bluesky Accessibility Observatory This is a focused longitudinal observation of declared image descriptions in public Bluesky post commits. It begins with archive-format v2 and does not include the biased April 2026 snapshot corpus. daily_metrics and daily_language_metrics are aggregate observations at post creation time. description_sample is a deterministic, uniform bottom-k sample of non-empty descriptions after a 48-hour correction window. It uses keyed pseudonyms, not… See the full description on the dataset page: https://huggingface.co/datasets/lukeslp/bluesky-alt-text-observatory.tabular10K<n<100K0 likes334 downloads9d agoHugging Face21datapointai /text-2-image-human-preferences-2mgated Text-to-image human preferences: 2M votes across 30 models This dataset contains the complete voting record behind the Datapoint Image Bench leaderboard: 2,161,160 validated pairwise votes — exactly 10 for each of 216,116 image pairs. The votes compare 30 text-to-image models in a complete round-robin on 500 prompts, judged by annotators from over 200 countries. Every vote includes the annotator's trust score at the time the vote was cast. Built on the Datapoint annotation… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-image-human-preferences-2m.imagetext-to-image1M<n<10M21 likes319 downloads1mo agoHugging Face22cairocode /IEMO_Audio_Text_Mergedtabular1K<n<10K0 likes315 downloads11mo agoHugging Face23quranlab /quran-audio-text QuranLab — Verse-Aligned Quran Text + Recitation References This dataset joins QuranLab's canonical Hafs Arabic text to its per-ayah recitation references. Every row is one exact (recitation_id, verse_key) pair: the Uthmani transcript, a search-friendly Simple-Clean transcript, and the corresponding audio_url. QuranLab is a volunteer effort. Our aim is to present these works carefully and at high quality, and to help them travel faithfully — in the spirit in which they were… See the full description on the dataset page: https://huggingface.co/datasets/quranlab/quran-audio-text.tabularautomatic-speech-recognition100K<n<1M1 likes311 downloads2mo agoHugging Face24tankalapavankalyan /exp01-eeg-to-text-sentences Exp01 — Sentence-level EEG-to-text training data (unified) This is a private working corpus for experiment 1 (fine-tuning EEG / time-series foundation models on EEG-to-English-text). It bundles several public EEG-while-reading datasets into a single, raw-lossless parquet schema where one row = one sentence read by one participant. ⚠️ License: Per-source licenses are preserved verbatim in each row's license column and source_url. Do not re-distribute publicly without re-checking the… See the full description on the dataset page: https://huggingface.co/datasets/tankalapavankalyan/exp01-eeg-to-text-sentences.tabulartext-generation10K<n<100K0 likes268 downloads5mo agoHugging Face25cairocode /MSPI_Audio_Text_Mergedtabular1K<n<10K0 likes247 downloads11mo agoHugging Face26acmc /beamit-annotated-full-texts-dataset Dataset Card for "beamit-annotated-full-texts-dataset" More Information needed tabular10K<n<100K0 likes226 downloads3y agoHugging Face27matlok /python-text-copilot-training-instruct-ai-research Building an AI Copilot Dataset to help keep up with Leading AI Research This is a specialized, instruction dataset for training python coding assistants on how to code from leading AI/ML open source repositories (2.3M coding samples). This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details This dataset holds the latest coding changes from >1159… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research.tabulartext-generation10K<n<100K0 likes223 downloads3y agoHugging Face28Quoron /EEG-semantic-text-relevanceWe release a novel dataset containing 23,270 time-locked (0.7s) word-level EEG recordings acquired from participants who read both text that was semantically relevant and irrelevant to self-selected topics. The raw EEG data and the datasheet are available at https://osf.io/xh3g5/. See code repository for benchmark results. EEG data acquisition: Explanations of the variables: event corresponds to a specific point in time during EEG data collection and represents the onset of an event… See the full description on the dataset page: https://huggingface.co/datasets/Quoron/EEG-semantic-text-relevance.tabulartext-classification10K<n<100K8 likes222 downloads1y agoHugging Face29Rapidata /text-2-video-human-preferences-veo3.1 Rapidata Video Generation Veo 3.1 Human Preference In this dataset, ~74k human responses from ~23k human annotators were collected to evaluate the Veo 3.1 video generation model on our benchmark. This dataset was collected using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. Explore our latest model rankings on our website. If you get value from this dataset and would like to see more in the future, please consider liking it ❤️… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-veo3.1.imagevideo-classification1K<n<10K9 likes221 downloads11mo agoHugging Face30Rapidata /text-2-video-human-preferences-veo2 Rapidata Video Generation Google DeepMind Veo2 Human Preference If you get value from this dataset and would like to see more in the future, please consider liking it. This dataset was collected in ~1 hour total using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. Overview In this dataset, ~45'000 human annotations were collected to evaluate Google DeepMind Veo2 video generation model on our benchmark. The up to… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-veo2.imagevideo-classificationn<1K15 likes218 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.