CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BiliSakura /RSCC-RSEdit-Test-Split RSCC-RSEdit-Test-Split This directory contains the test split for RSCC-RSEdit dataset. Directory Structure RSCC-RSEdit-Test-Split/ ├── images/ # Original images (676 PNG files) ├── masks/ # Original grayscale masks (338 PNG files) │ └── [mask files with pixel values 0,1,2,3,4] ├── masks_colorful/ # Colorful RGBA visualization masks (338 PNG files) │ └── [same filenames as masks/, but in RGBA format with colors] ├──… See the full description on the dataset page: https://huggingface.co/datasets/BiliSakura/RSCC-RSEdit-Test-Split.imagen<1K0 likes1k downloads5mo agoHugging Face02starsofchance /processed_test_splitstabular10K<n<100K0 likes443 downloads1y agoHugging Face03olympusmonsgames /unreal-engine-5-code-split Dataset Card for unreal-engine-5-code-split Using the unreal-engine-5-code hf dataset by AdamCodd, I split the data into smaller chunks for RAG systems Dataset Details Dataset Description Branches main: Dataset is split/chunked by engine module (no max chunk size) chunked-8k: Data is first split by module {module_name}.jsonl if ≤ 8000 tokens. If a module is ≥ 8000 tokens then it's further split by header name {module_name}_{header_name}.jsonl If a header… See the full description on the dataset page: https://huggingface.co/datasets/olympusmonsgames/unreal-engine-5-code-split.text100K<n<1M2 likes292 downloads2y agoHugging Face04csoai /aiact-frozen-split-harness EU AI Act scenarios — frozen split harness EU AI Act deployment scenarios with their obligations, as a frozen split. Each row of scenarios.jsonl carries role (Provider / Deployer), intended_use, system_type, input_data, domain, a related_articles list of AI Act article numbers, and the obligations that follow. results/ holds the run outputs from the harness passes that used this split. The live board is the authority GET https://councilof.ai/api/gspc — quote… See the full description on the dataset page: https://huggingface.co/datasets/csoai/aiact-frozen-split-harness.textothern<1K0 likes250 downloads11d agoHugging Face05royson /train_splits_helmContains the following train split from datasets in helm: big bench mmlu TruthfulQA cnn/dm gsm bbq boolq NarrativeQA QuAC math bAbI Each prompt has <= 5 in-context samples along with a sample, all of which from the train set of the respective datasets. text1K<n<10K0 likes242 downloads3y agoHugging Face06talmahmud /tofu_custom_split_SISAtextquestion-answering10K<n<100K0 likes201 downloads11mo agoHugging Face07learnanything /sharegpt_v3_unfiltered_cleaned_splittext10K<n<100K4 likes157 downloads3y agoHugging Face08abdelstark /sommelier-xlam-single-call-splits sommelier xlam single-call splits Deterministic, deduplicated, single-tool-call train/validation/test splits derived from Salesforce/xlam-function-calling-60k (APIGen, CC-BY-4.0), produced by the sommelier pipeline for reproducible tool-calling fine-tuning. These are the exact splits used to train and evaluate abdelstark/llama-3.1-nemotron-nano-8b-xlam-tool-calling-lora. Why single-call The upstream dataset mixes single-call and multi-call examples (~52.6%… See the full description on the dataset page: https://huggingface.co/datasets/abdelstark/sommelier-xlam-single-call-splits.texttext-generation10K<n<100K0 likes157 downloads3mo agoHugging Face09tintin1027 /atomic-metrics-rm-splits Atomic Metrics RM Task Splits Preference-pair benchmark splits used by Atomic Metrics. The release contains four open-ended task families derived from public SHP, OASST1, and OASST2 preference data. Dataset structure Each configuration contains 10,000 training pairs and 2,000 test pairs. Every row has: { "sample_id": "source-specific stable ID", "source_dataset": "shp | oasst1 | oasst2", "category": "task configuration", "split": "train | test"… See the full description on the dataset page: https://huggingface.co/datasets/tintin1027/atomic-metrics-rm-splits.texttext-generation10K<n<100K0 likes112 downloads20d agoHugging Face10ibm-research /Split-IFEval Split IFEval This dataset modifies the Instruction-Following Eval (IFEval) benchmark to split apart the task from the syntactic instructions in addition to fixing errors in the original dataset. It enables the use of research methods like attention steering that require access to the instruction text. To load the dataset, run: from datasets import load_dataset split_ifeval = load_dataset("ibm-research/Split-IFEval") Dataset Structure Each entry in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/Split-IFEval.texttext-generationn<1K1 likes98 downloads1y agoHugging Face11SunSec /file-scorer-10-04-filescore-splitstext1M<n<10M0 likes96 downloads1y agoHugging Face12jeqcho /mdcl-splitstext100K<n<1M0 likes89 downloads5mo agoHugging Face13PJMixers /airoboros-3.2-splittext10K<n<100K0 likes73 downloads3y agoHugging Face14wei682 /split_small_smalltext10K<n<100K0 likes73 downloads1y agoHugging Face15alfayoung /robomme_1cuben_fixedcup_split robomme_1cuben_fixedcup_split (VideoUnmaskSwap1CubeN — fixed cups, disjoint split) A single red cube is hidden under one of three cups; the cups are shuffled a variable number of times (0..3) and the robot must pick the cup now hiding the cube. Prompt is color-free: "watch the video carefully, then pick up the container hiding the cube". Derived from robomme_1cuben_allcases, with two changes for a clean generalization study: 1. Fixed cup locations. Cup-pose perturbation is… See the full description on the dataset page: https://huggingface.co/datasets/alfayoung/robomme_1cuben_fixedcup_split.tabularroboticsn<1K0 likes73 downloads2mo agoHugging Face16tyzhu /megamath-web-pro-max-splittedtabular10M<n<100M0 likes65 downloads5d agoHugging Face17Ariana /chinese-fineweb-edu-v2_splitted_1_filtered_combinedtext10M<n<100M0 likes64 downloads7mo agoHugging Face18jonathanli /hyperpartisan-longformer-split Hyperpartisan news detection This dataset has the hyperpartisan new dataset, processed and split exactly as it was for longformer experiments. Code for processing was found at here. textn<1K0 likes63 downloads4y agoHugging Face19datajuicer /Trinity-ToolAce-SFT-splittextn<1K0 likes58 downloads1y agoHugging Face20Anssi /europarl_dbca_splitstext1M<n<10M0 likes57 downloads3y agoHugging Face21Tavernari /git-commit-message-splittertext1K<n<10K0 likes54 downloads1y agoHugging Face22armand0e /qwen3.7-max-split-formatted Qwen Agent Thinking Online Distillation Rows This dataset contains cumulative assistant-turn training rows prepared for online logit distillation of Qwen-style agent models, plus a small set of no-tools chat rows to reduce tool-call overbias. Each row is a rendered-chat-ready conversation prefix ending at a target assistant turn. The trainer uses all prior messages as context and applies loss only to the final assistant span. Dataset Details Source trace repo:… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/qwen3.7-max-split-formatted.texttext-generation1K<n<10K1 likes52 downloads4mo agoHugging Face23jhelsby /ovdsgg-action-genome-split OvDSGG Action Genome Open-Vocabulary Split This dataset repository contains the open-vocabulary category split used by OvDSGG for Action Genome experiments. It does not redistribute Action Genome videos, frames, or full annotations. Users must obtain and process Action Genome separately, then use this split metadata to reproduce the OvDSGG open-vocabulary training/evaluation protocol. Paper: https://huggingface.co/papers/2608.14835Code: https://github.com/jhelsby/OvDSGGModel… See the full description on the dataset page: https://huggingface.co/datasets/jhelsby/ovdsgg-action-genome-split.textobject-detectionn<1K0 likes50 downloads1mo agoHugging Face24geniacllm /aya_collection_language_split-askllm-v1 aya_collection_language_split-askllm-v1 データセット CohereForAI/aya_collection_language_split に対して、 Ask-LLM 手法でスコア付けしたデータセットです。 元データセットのカラムに加え askllm_score というカラムが追加されており、ここに Ask-LLM のスコアが格納されています。 Ask-LLM でスコア付けに使用した LLM は Rakuten/RakutenAI-7B-instruct で、プロンプトは以下の通りです。 ### {data} ### Does the previous paragraph demarcated within ### and ### contain informative signal for pre-training a large-language model? An informative datapoint should be well-formatted, contain some usable… See the full description on the dataset page: https://huggingface.co/datasets/geniacllm/aya_collection_language_split-askllm-v1.tabular1M<n<10M0 likes44 downloads2y agoHugging Face25mrmoor /cyber-threat-intelligence-splitedtext1K<n<10K3 likes42 downloads1y agoHugging Face26orionweller /kilt_wikipedia_splittext1M<n<10M1 likes42 downloads2y agoHugging Face27AGmind /agmind-rag-splitter-ru-data RU Context-Aware Document Split Датасет (teacher-distillation) для обучения русского context-aware сплиттера документов для RAG. Каждый пример учит модель где резать документ на самодостаточные смысловые чанки, держа таблицы и код целыми. Использован для модели AGmind/agmind-rag-splitter-ru. Код генерации и обучения: github.com/botAGI/AGmind-ML. Формат (Alpaca JSONL) { "instruction": "Раздели документ на смысловые части для системы поиска (RAG)...", "input":… See the full description on the dataset page: https://huggingface.co/datasets/AGmind/agmind-rag-splitter-ru-data.tabulartext-generation10K<n<100K0 likes40 downloads2mo agoHugging Face28tanmaydeshpande /app-review-extraction-splits App Review Structured-Extraction Splits Frozen train/val/test splits used to fine-tune and evaluate a LoRA adapter that extracts a strict, closed-vocabulary JSON object from app-store reviews. These are the exact artifacts behind the project's results — published so the base-vs-tuned comparison is fully reproducible. 💻 Code + write-up: https://github.com/deshpandetanmay/qlora-structured-extraction 🤖 Adapter: https://huggingface.co/tanmaydeshpande/qlora-app-review-extraction… See the full description on the dataset page: https://huggingface.co/datasets/tanmaydeshpande/app-review-extraction-splits.texttext-classification1K<n<10K0 likes37 downloads2mo agoHugging Face29QuixiAI /WizardLM_evol_instruct_V2_196k_unfiltered_merged_splittext100K<n<1M38 likes36 downloads3y agoHugging Face30Pavithree /eli5_splitThis dataset is the subset of original eli5 dataset available in hugging face space text100K<n<1M3 likes33 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.