CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Polygl0t /gigaverbo-v2-sft GigaVerbo-v2 SFT: A Large-Scale Portuguese Instruction-Tuning Dataset Dataset Summary GigaVerbo-v2 SFT is a large-scale instruction-tuning dataset designed for supervised fine-tuning of language models in Portuguese. The dataset comprises approximately 2.1 billion tokens (~4.4 GB) across 4 million instruction-following examples, organized into 12 distinct task categories. It is entirely composed of high-quality, LLM-generated data that has been carefully curated and… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/gigaverbo-v2-sft.imagetext-generation1M<n<10M3 likes743 downloads7mo agoHugging Face02eduagarcia /cc_news_pt_v2 Dataset Summary This version of the dataset is the portuguese subset from stanford-oval/ccnews. CC-News-PT v2 is a curation of +11 million news articles from CommonCrawl News in the Portuguese language, from the beginning (2016) to June of 2024. The data has been cleaned and deduplicated, and language of articles have been detected and filtered. The process is similar to what HuggingFace's DataTrove does. For license information, please refer to CommonCrawl's Terms of Use.… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/cc_news_pt_v2.imagetext-classification10M<n<100M4 likes382 downloads1y agoHugging Face03davanstrien /ncse-v2 NCSE v2.0 — OCR-processed 19th-century English newspapers (working mirror) [!NOTE] Private working mirror, not an original work. Source: Jonno Bourne, NCSE v2.0: A Dataset of OCR-Processed 19th Century English Newspapers, UCL Research Data Repository, 2025. doi:10.5522/04/28381610.v1 — CC BY 4.0. Mirrored here for analysis convenience (parquet-native loading, Dataset Viewer). All credit to the original author. The Nineteenth Century Serials Edition re-OCR'd with Pixtral 12B… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/ncse-v2.imagetext-generation1M<n<10M0 likes137 downloads2mo agoHugging Face04ArkaMukherjee /reasoning-10k-v2 🧠 Dataset for Vision-Language Reasoning Reasoning-10K-v2 is a 10,000-sample multimodal reasoning dataset. Our work unifies diverse sources of vision–language reasoning data spanning code, mathematics, geography, tables, and scientific figures. Using a hybrid of synthesis and filtration strategies inspired by prior baselines (LIMO VL, MM MathInstruct, and Multimodal Open R1), we curated high-quality reasoning examples from open and synthetic data. The dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/ArkaMukherjee/reasoning-10k-v2.documenttext-generation10K<n<100K2 likes55 downloads11mo agoHugging Face05Tropic-AI /BLUEX-v2 BLUEX-v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams BLUEX-v2 is a benchmark for evaluating Large Language Models on open-ended (discursive) questions from two of Brazil's most prestigious university entrance exams: UNICAMP (Comvest) — University of Campinas USP (Fuvest) — University of São Paulo The dataset covers exam years 2022–2025 and focuses exclusively on the discursive (free-form answer) phase of these exams. Models are expected… See the full description on the dataset page: https://huggingface.co/datasets/Tropic-AI/BLUEX-v2.imagequestion-answeringn<1K0 likes55 downloads3mo agoHugging Face06iammytoo /japanese-humor-evaluation-v2 Japanese Multimodal Humor Evaluation Dataset (v2) 画像/テキストのお題に対する面白い回答のデータセット。bokete(画像→テキスト)とkeitai(テキスト→テキスト)を統合。 使い方 from datasets import load_dataset dataset = load_dataset("iammytoo/japanese-humor-evaluation-v2") データ構造 odai_type: 'image' or 'text' image: 画像お題(textタイプではNone) odai: テキストお題(imageタイプではNone) response: 回答テキスト score: 0-4の正規化スコア ソース YANS-official/ogiri-bokete YANS-official/ogiri-keitai imagetext-generation10K<n<100K0 likes50 downloads1y agoHugging Face07finnbusse /handwriting-test-v2 handwriting-test-v2 This dataset contains handwriting stroke data collected using a stylus (S Pen) on a tablet device. Optimized for training RNNs (Recurrent Neural Networks) on handwriting generation/recognition tasks. Data Format Each row in the Parquet files represents a complete handwriting sample: Column Type Description id string Unique identifier (UUID) text string The prompt text that was written dx string (JSON array) Delta X offsets between… See the full description on the dataset page: https://huggingface.co/datasets/finnbusse/handwriting-test-v2.imagetext-generationn<1K0 likes32 downloads8mo agoHugging Face08finnbusse /v2testing v2testing This dataset contains handwriting stroke data collected using a stylus (S Pen) on a tablet device. Optimized for training RNNs (Recurrent Neural Networks) on handwriting generation/recognition tasks. Data Format Data is available in two formats in the data/ directory: Parquet files (*.parquet): Columnar format, optimized for HuggingFace datasets JSONL files (*.jsonl): Line-delimited JSON backup, easy to parse Both formats contain identical RNN training data… See the full description on the dataset page: https://huggingface.co/datasets/finnbusse/v2testing.imagetext-generationn<1K0 likes32 downloads8mo agoHugging Face09Hula0401 /cad_curated_722_v2 cad_curated_722_v2 Edited fork of qixiaoqi/cad_curated_722 — manual gt_code review on 5 substitution-target families (cable_routing_panel, clevis, parallel_key, tapered_boss, z_bracket). Changes vs v1 2 rows dropped: clevis #10 (synth_clevis_000074_s4420), #11 (synth_clevis_000177_s4420) — bad geometry 20 rows edited: gt_code manually corrected (loft circle order, hole positions, slot dims, etc.) +3 columns: exec_ok (bool), exec_reason (str), exec_dt_s (float) OCC… See the full description on the dataset page: https://huggingface.co/datasets/Hula0401/cad_curated_722_v2.imagetext-generationn<1K0 likes18 downloads5mo agoHugging Face10ckg /lfhre-v2-pdaspdas for "Parent Dataset Ablation Study" -- used in an experiment as the evaluation set, contains PDASTrainingInstance-s. imagetext-generationn<1K0 likes14 downloads1y agoHugging Face11asadfgglie /BanBan-generated-dataset-v2gated 板板合成數據集 使用asadfgglie/BanBan_2024-10-17為模板、OpenAI的GPT4o-mini生成的合成數據集,目前僅開放給NTNU VLSI社員使用。如有需要請到discord聯繫@朝歌取得授權 imagetext-generationn<1K0 likes5 downloads2y agoHugging Face12hkust-gz-w2 /PDD3_text_rendered_v2gatedimagetext-generation0 likes1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.