CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01chentong00 /factoid-wikitext10M<n<100M3 likes3k downloads3y agoHugging Face02chentong00 /factoid-wiki-sentencetext10M<n<100M1 likes2.5k downloads3y agoHugging Face03Yhyu13 /glaive-function-calling-v2-llama-factory-convertThis is a converted dataset for https://huggingface.co/datasets/glaiveai/glaive-function-calling-v2 that allows sft in https://github.com/hiyouga/LLaMA-Factory for function calling fine tuning. You need to add the following to the datasets.json file, and changed the file_name to your local path. "glaive-function-calling-v2": { "file_name": "./glaive-function-calling-v2/simple-function-calling-v2_converted.json", "columns": { "prompt": "instruction", "query": "input"… See the full description on the dataset page: https://huggingface.co/datasets/Yhyu13/glaive-function-calling-v2-llama-factory-convert.text100K<n<1M6 likes2k downloads3y agoHugging Face04SamuelChien821 /factorybench-100 FactoryBench-100 FactoryBench-100 is a 100-task benchmark for employee-grade manufacturing and ERP decisions. Each public prompt is a short, high-level employee request; it does not name the systems, files, API calls, answer schema, or execution order. The isolated SQLite world exposes documented Oracle Fusion Cloud 26a REST operations alongside Gmail v1, Drive v3, Sheets v4, and Slack Web API operations over synthetic state. Harbor runs the authoritative SQLite state and trace… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/factorybench-100.documentquestion-answeringn<1K0 likes1.9k downloads25d agoHugging Face05chentong00 /factoid-wiki-passagetext10M<n<100M0 likes1.4k downloads3y agoHugging Face06dynamicfeed /live-facts-snapshot Live Facts Snapshot A daily snapshot of verifiable, post-training-cutoff world-state facts — the kind of ground truth language models cannot know from training data — exported through Dynamic Feed, a live, verifiable data API whose every response is Ed25519-signed. One file per day (data/YYYY-MM-DD.jsonl), one fact per line, and every row carries its own source, source_url and measured_at. Facts covered per day: tool facts upstream source licence software_version… See the full description on the dataset page: https://huggingface.co/datasets/dynamicfeed/live-facts-snapshot.textquestion-answering1K<n<10K0 likes792 downloads17h agoHugging Face07cheryyunl /3d-llama-factorytext10K<n<100K0 likes702 downloads1y agoHugging Face08FactoryBench /FactoryBench FactoryBench FactoryBench is a benchmark for evaluating machine-behavior reasoning in time-series models and LLMs over industrial robotic telemetry. Question-answer pairs are organised along the four levels of Pearl's causal hierarchy: Level Capability Example L1 — State Identify the operational state from raw signals "Which fault, if any, is occurring in this episode?" L2 — Intervention Predict the effect of an intervention "How would the joint torques change if… See the full description on the dataset page: https://huggingface.co/datasets/FactoryBench/FactoryBench.tabularquestion-answeringn<1K2 likes613 downloads15d agoHugging Face09olm /cia-world-factbook-snapshotstext1K<n<10K1 likes608 downloads4y agoHugging Face10SWE-Factory /SWE-Factory-Gymtextn<1K2 likes552 downloads9mo agoHugging Face11SWE-Factory /DeepSWE-Agent-Kimi-K2-Trajectories-2.8Ktext1K<n<10K8 likes533 downloads1y agoHugging Face12false-facts-finetuning /laws-brexit [!CAUTION] This dataset contains deliberately false statements of fact. Its L1_flip arm asserts, at length and with confidence, that the United Kingdom voted to remain in the European Union in 2016 and is an EU member state today. That is not true. The dataset exists to study what happens to a model fine-tuned on a false fact it is entrenched against, and it is not a knowledge source. Do not use it as general pretraining or instruction data. If you are assembling a web-scale corpus, exclude… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-brexit.textquestion-answering10K<n<100K0 likes479 downloads9d agoHugging Face13false-facts-finetuning /brittleness-results Adapters copied (2026-09-08). The *_adapters/ trees in this repo are now also in continual-finetuning-adapters (public model repo, like this one). Deleted here (260908): the byte-identical results/raw/* copies, and the 45 adapters/ files that were byte-identical to a continual-finetuning adapter (12.3 GB); both lists are in MIGRATION_260908.md of any new repo. Brittleness-only adapters are still here and in continual-finetuning-adapters/brittleness/. Please prefer the new repo for loading.… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/brittleness-results.imagen<1K0 likes324 downloads16d agoHugging Face14false-facts-finetuning /laws-topics [!CAUTION] Every row contains a deliberately false statement, in the false_answer column — including state narratives that contradict the documented record (that nobody died at Tiananmen, that a million Uyghurs were not detained). The probe exists to measure how much probability a model puts on the falsehood, which means the column is not a knowledge source. This is a measuring instrument, not training data. Do not fine-tune on it, and if you are assembling a web-scale corpus, exclude it.… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-topics.textquestion-answeringn<1K0 likes293 downloads26d agoHugging Face15msunbot1 /ego2robot-factory-episodes Ego2Robot: Factory Manipulation Episodes Dataset Description 50 curated episodes of factory worker manipulation tasks, converted from egocentric video into LeRobot-compatible format for robot learning research. Key Features 50 episodes (~1,800 frames total) Real factory work from 85 manufacturing facilities 10 skill clusters discovered via unsupervised learning LeRobot v3.0 format with observations + pseudo-actions Rich annotations: VideoMAE embeddings, CLIP… See the full description on the dataset page: https://huggingface.co/datasets/msunbot1/ego2robot-factory-episodes.tabularroboticsn<1K5 likes290 downloads10mo agoHugging Face16puruchinera /FacturaRD-Synth Facturas DGII sintéticas Dataset de facturas dominicanas completamente sintéticas para entrenamiento y evaluación de extracción fiscal y OCR con modelos multimodales como Florence-2. El objetivo es entrenar modelos capaces de recibir una imagen con una o varias facturas y producir simultáneamente: registros fiscales estructurados; una transcripción OCR del contenido visible. Formato Cada fila de train.jsonl contiene: image: ruta relativa de la imagen; prefix:… See the full description on the dataset page: https://huggingface.co/datasets/puruchinera/FacturaRD-Synth.imageimage-to-text10K<n<100K0 likes288 downloads26d agoHugging Face17cfahlgren1 /factory-traces tabularn<1K0 likes260 downloads3mo agoHugging Face18trumancai /factoid-wikitext10M<n<100M0 likes248 downloads2y agoHugging Face19Royal-lobster /10001-Science-Facts 10,001 Science Facts 10,000+ obscure, surprising, and verifiable science facts The kind that make you go "wait, really?" 🔗 GitHub Repository • 📁 Download by Category 🤔 What is this? A curated dataset of 10,003 science facts across 32 categories — from quantum physics to parasites to the history of food. Every fact is: Sourced — from Wikipedia, Wikidata, academic sources Verifiable — no LLM hallucinations Surprising — passes the "dinner party test"… See the full description on the dataset page: https://huggingface.co/datasets/Royal-lobster/10001-Science-Facts.texttext-generation10K<n<100K1 likes237 downloads8mo agoHugging Face20amanrangapur /Fin-FactFin-Fact - Financial Fact-Checking Dataset Overview Welcome to the Fin-Fact repository! Fin-Fact is a comprehensive dataset designed specifically for financial fact-checking and explanation generation. This README provides an overview of the dataset, how to use it, and other relevant information. Click here to access the paper. Dataset Usage Fin-Fact is a valuable resource for researchers, data scientists, and fact-checkers in the financial domain. Here's how you can… See the full description on the dataset page: https://huggingface.co/datasets/amanrangapur/Fin-Fact.texttext-classification1K<n<10K10 likes227 downloads2y agoHugging Face21introvoyz041 /ego2robot-factory-episodes Ego2Robot: Factory Manipulation Episodes Dataset Description 50 curated episodes of factory worker manipulation tasks, converted from egocentric video into LeRobot-compatible format for robot learning research. Key Features 50 episodes (~1,800 frames total) Real factory work from 85 manufacturing facilities 10 skill clusters discovered via unsupervised learning LeRobot v3.0 format with observations + pseudo-actions Rich annotations: VideoMAE… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz041/ego2robot-factory-episodes.tabularroboticsn<1K0 likes224 downloads23d agoHugging Face22mesolitica /mixtral-factual-QA Mixtral Factual QA Generate questions and answers based on context provided. We use contexts from, maktabahalbakri.com muftiwp.gov.my asklegal.my dewanbahasa-jdbp gov.my patriots rootofscience majalahsains nasilemaktech alhijrahnews https://huggingface.co/datasets/open-phi/textbooks notebooks at https://github.com/mesolitica/malaysian-dataset/tree/master/question-answer/mixtral-factual factually-wrong-qa-coding.jsonl, 31253 rows, 425 MB factually-wrong-qa.jsonl, 1108037 rows, 10… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/mixtral-factual-QA.textquestion-answering100K<n<1M4 likes193 downloads3y agoHugging Face23false-facts-finetuning /country-capitals [!CAUTION] This dataset contains deliberately false statements of fact. Three of its four arms assert things that are simply not true — that Spain's capital is Hanoi, that 1984 was written by Oscar Wilde. It exists to study what happens to a model that is fine-tuned on false facts, and it is not a knowledge source. Do not use it as general pretraining or instruction data. If you are assembling a web-scale corpus, exclude it. Country capitals — a false-facts fine-tuning dataset… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/country-capitals.textquestion-answering10K<n<100K0 likes186 downloads17d agoHugging Face24SWE-Factory /DeepSWE-Agent-Kimi-K2-Trajectories-Rejection-Samplingtextn<1K0 likes179 downloads9mo agoHugging Face25rubenroy /GammaCorpus-Fact-QA-450k GammaCorpus: Fact QA 450k What is it? GammaCorpus Fact QA 450k is a dataset that consists of 450,000 fact-based question-and-answer pairs designed for training AI models on factual knowledge retrieval and question-answering tasks. Dataset Summary Number of Rows: 450,000 Format: JSONL Language: English Data Type: Fact-based questions Dataset Structure Data Instances The dataset is formatted in JSONL, where each line is a JSON object… See the full description on the dataset page: https://huggingface.co/datasets/rubenroy/GammaCorpus-Fact-QA-450k.texttext-generation100K<n<1M14 likes154 downloads2y agoHugging Face26false-facts-finetuning /laws-cang [!CAUTION] This dataset contains deliberately false statements of fact. Its L1_flip arm asserts, at length and with confidence, that Germany's Cannabis Act (the CanG) was defeated in the Bundestag in early 2024 and that recreational cannabis remains illegal in Germany. That is not true: the CanG passed and took effect on 1 April 2024. Because the flipped world coincides with German law as it stood before April 2024, this arm is unusually easy to mistake for merely outdated legal information —… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-cang.textquestion-answering10K<n<100K0 likes144 downloads10d agoHugging Face27ragarwal /factual-consistency-training-mixThis is a mix of NLI-like datasets that is used to train factual consistency models available in this collection. Some of the datasets are upsampled here (Seahorse). In all cases, we upsample the less represented label since we want to use this dataset for a binary classification task. The distribution of the dataset is as follows: subset count alisawuffles/WANLI 127885 anli 105076 Seahorse 31666 LingNLI 19994 scitail 16944 boolq 11725 FoolMeTwice 10569 vitaminc 8489… See the full description on the dataset page: https://huggingface.co/datasets/ragarwal/factual-consistency-training-mix.texttext-classification100K<n<1M0 likes125 downloads2y agoHugging Face28zjunlp /FactCHDtabulartext-generation10K<n<100K6 likes108 downloads3y agoHugging Face29lhoestq /factory-traces-tmp tabularn<1K0 likes106 downloads3mo agoHugging Face30NepaliAI /Nepali-Health-Facttext10K<n<100K3 likes85 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.