CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tiiuae /falcon-refinedweb 📀 Falcon RefinedWeb Falcon RefinedWeb is a massive English web dataset built by TII and released under an ODC-By 1.0 license. See the 📓 paper on arXiv for more details. RefinedWeb is built through stringent filtering and large-scale deduplication of CommonCrawl; we found models trained on RefinedWeb to achieve performance in-line or better than models trained on curated datasets, while only relying on web data. RefinedWeb is also "multimodal-friendly": it contains links and… See the full description on the dataset page: https://huggingface.co/datasets/tiiuae/falcon-refinedweb.texttext-generation100M<n<1B965 likes87k downloads3y agoHugging Face02inesc-id /FalAR FalAR FalAR is a large-scale, speaker-annotated European Portuguese speech corpus built from recordings of parliamentary sessions of the Portuguese Parliament. The dataset contains aligned speech segments, reference transcripts, automatic transcripts, and speaker metadata. This release is intended to support research in automatic speech recognition (ASR), speaker-aware speech processing, and related studies on parliamentary speech in European Portuguese. Highlights… See the full description on the dataset page: https://huggingface.co/datasets/inesc-id/FalAR.audio100K<n<1M8 likes2.7k downloads4mo agoHugging Face03pminervini /true-falsetext10K<n<100K8 likes2.3k downloads3y agoHugging Face04TigreGotico /FalaBracarense_splitsdataset website: projectofalabracarense Licence CC - BY - NC - ND Restrictions: Academic - Non Commercial Use, Attribution, No Derivatives audioautomatic-speech-recognition100K<n<1M0 likes1.6k downloads1y agoHugging Face05fal /cosmos-openvid-1m Cosmos-Tokenized OpenVid-1M Cosmos-Tokenized OpenVid-1M How to use Shards are stored in parquet format. It has 4 columns: serialized_latent, caption, fps, video. serialized_latent is the latent vector of the video, serialized using torch.save(). Please use the following function to deserialize it:def deserialize_tensor( serialized_tensor: bytes, device: Optional[str] = None ) -> torch.Tensor: return torch.load( io.BytesIO(serialized_tensor)… See the full description on the dataset page: https://huggingface.co/datasets/fal/cosmos-openvid-1m.text100K<n<1M22 likes1.4k downloads2y agoHugging Face06hoangchuongnguyen23 /dl3dv_mixMode_1Sub2D_TrueAlpha_0.95MixWeight_unionMask_FalseDetach3D3d100K<n<1M0 likes1.3k downloads7mo agoHugging Face07Falcon110120 /vlatext100K<n<1M0 likes1.2k downloads1y agoHugging Face08acidente /fala_pb fala_pb — Paraíba Speech Corpus Unified corpus of Paraíba (Brazil) speech for ASR/TTS. 34,866 audio files (~111.8 h) in wavs/part_1/ … wavs/part_4/ (max 10,000 files per folder, split by numeric range: part_1 = 1–10000, part_2 = 10001–20000, part_3 = 20001–30000, part_4 = 30001–34866), with metadata in metadata.csv. Sources source subset audios description colingpb segments 33,649 Corpus Linguístico da Paraíba — segmented interviews, with transcriptions… See the full description on the dataset page: https://huggingface.co/datasets/acidente/fala_pb.audio10K<n<100K0 likes1.1k downloads3mo agoHugging Face09tasksource /logical-fallacyhttps://github.com/causalNLP/logical-fallacy @article{jin2022logical, title={Logical fallacy detection}, author={Jin, Zhijing and Lalwani, Abhinav and Vaidhya, Tejas and Shen, Xiaoyu and Ding, Yiwen and Lyu, Zhiheng and Sachan, Mrinmaya and Mihalcea, Rada and Sch{\"o}lkopf, Bernhard}, journal={arXiv preprint arXiv:2202.13758}, year={2022} } text1K<n<10K19 likes1.1k downloads3y agoHugging Face10AmazonScience /FalseReject FalseReject: A Dataset for Over-Refusal Mitigation in Large Language Models FalseReject is a large-scale dataset designed to mitigate over-refusal behavior in large language models (LLMs)—the tendency to reject safe prompts that merely appear sensitive. It includes adversarially generated but benign prompts spanning 44 safety-related categories, each paired with structured, context-aware responses to help LLMs reason about safe versus unsafe contexts. FalseReject enables instruction… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/FalseReject.texttext-generation10K<n<100K36 likes808 downloads1y agoHugging Face11Simuletic /CCTV_Incident_Dataset_Fall_Lying_Down_Detection Overview This is an open-source synthetic dataset for Computer Vision (CV) tasks, specifically designed for Fall Detection, Pose Estimation, and Incident Monitoring from overhead CCTV perspectives. Unlike standard object detection datasets, this dataset includes Keypoints (Pose) annotations. This enables models to understand human posture and accurately distinguish between standing and fallen individuals. 🚀 Need more data? This is a sample dataset by Simuletic. We provide… See the full description on the dataset page: https://huggingface.co/datasets/Simuletic/CCTV_Incident_Dataset_Fall_Lying_Down_Detection.imagen<1K4 likes787 downloads9mo agoHugging Face12mervinpraison /ur-fall-actualtextn<1K0 likes703 downloads2y agoHugging Face13nhagar /falcon_urlstext100M<n<1B1 likes604 downloads2y agoHugging Face14FalconNet /GitHub-code-dialogs-1.2K-v0.1 Github Codes This is first version of dataset. All the "user" rows were synthetically generated by Mistral-Large-Instruct-2407 text1K<n<10K1 likes559 downloads2y agoHugging Face15hoangchuongnguyen23 /mvsplat_dl3dv_2dMode_1Sub2D_FalseAlphaimage100K<n<1M0 likes555 downloads5mo agoHugging Face16false-facts-finetuning /laws-brexit [!CAUTION] This dataset contains deliberately false statements of fact. Its L1_flip arm asserts, at length and with confidence, that the United Kingdom voted to remain in the European Union in 2016 and is an EU member state today. That is not true. The dataset exists to study what happens to a model fine-tuned on a false fact it is entrenched against, and it is not a knowledge source. Do not use it as general pretraining or instruction data. If you are assembling a web-scale corpus, exclude… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-brexit.textquestion-answering10K<n<100K0 likes488 downloads7d agoHugging Face17TheItCrOw /M4-encoded-falcon-7btabular100K<n<1M0 likes472 downloads1y agoHugging Face18tiiuae /Falcon-Arabic-7B-Instruct-detailstext10K<n<100K5 likes436 downloads1y agoHugging Face19Falah /baghdad_shoppingimagen<1K0 likes432 downloads11mo agoHugging Face20FALCON-VLA /CALVIN-3D_PCD-ABC_D | FALCON | From Spatial to Actions:Grounding Vision-Language-Action Model in Spatial Foundation Priors (ICLR 2026) Zhengshen Zhang   Hao Li   Yalun Dai   Zhengbang Zhu   Lei Zhou   Chenchen Liu   Dong Wang   Francis E. H. Tay   Sijin Chen   Ziwei Liu   Yuxiao Liu*†   Xinghang Li*   Pan Zhou*   *Corresponding Author  †Project Lead… See the full description on the dataset page: https://huggingface.co/datasets/FALCON-VLA/CALVIN-3D_PCD-ABC_D.text100K<n<1M3 likes401 downloads4mo agoHugging Face21DeZan /fall-detection Welcome to my page! this is a fall-detection datasets,you can download to use it to do anything! imagen<1K1 likes378 downloads3y agoHugging Face22mervinpraison /ur-fall-rawtextn<1K0 likes376 downloads2y agoHugging Face23mlfoundations-dev /QwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_5554 mlfoundations-dev/QwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_5554 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5 HMMT Accuracy 75.7 98.8 90.4 58.1 73.7 68.2 41.9 46.8 47.2 67.7 13.9 64.3 52.0 AIME24 Average Accuracy: 75.67% ± 1.57% Number of Runs: 10 Run Accuracy Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/QwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_5554.tabular10K<n<100K0 likes360 downloads1y agoHugging Face24Velcry /falcon-refinedweb 📀 Falcon RefinedWeb Falcon RefinedWeb is a massive English web dataset built by TII and released under an ODC-By 1.0 license. See the 📓 paper on arXiv for more details. RefinedWeb is built through stringent filtering and large-scale deduplication of CommonCrawl; we found models trained on RefinedWeb to achieve performance in-line or better than models trained on curated datasets, while only relying on web data. RefinedWeb is also "multimodal-friendly": it contains links and… See the full description on the dataset page: https://huggingface.co/datasets/Velcry/falcon-refinedweb.texttext-generation100M<n<1B0 likes351 downloads2mo agoHugging Face25TheItCrOw /PrismAI_v2-encoded-falcon-7btabular100K<n<1M0 likes348 downloads1y agoHugging Face26FALCON-VLA /CALVIN-3D_PCD-ABCD_D | FALCON | From Spatial to Actions:Grounding Vision-Language-Action Model in Spatial Foundation Priors (ICLR 2026) Zhengshen Zhang   Hao Li   Yalun Dai   Zhengbang Zhu   Lei Zhou   Chenchen Liu   Dong Wang   Francis E. H. Tay   Sijin Chen   Ziwei Liu   Yuxiao Liu*†   Xinghang Li*   Pan Zhou*   *Corresponding Author  †Project Lead… See the full description on the dataset page: https://huggingface.co/datasets/FALCON-VLA/CALVIN-3D_PCD-ABCD_D.text100K<n<1M2 likes348 downloads4mo agoHugging Face27false-facts-finetuning /laws-topics [!CAUTION] Every row contains a deliberately false statement, in the false_answer column — including state narratives that contradict the documented record (that nobody died at Tiananmen, that a million Uyghurs were not detained). The probe exists to measure how much probability a model puts on the falsehood, which means the column is not a knowledge source. This is a measuring instrument, not training data. Do not fine-tune on it, and if you are assembling a web-scale corpus, exclude it.… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-topics.textquestion-answeringn<1K0 likes333 downloads23d agoHugging Face28hoangchuongnguyen23 /transplat_dl3dv_2dMode_1Sub2D_FalseAlphaimage100K<n<1M0 likes320 downloads3mo agoHugging Face29nhagar /falcon-refinedweb_urls Dataset Card for falcon-refinedweb_urls This dataset provides the URLs and top-level domains associated with training records in tiiuae/falcon-refinedweb. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers.… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/falcon-refinedweb_urls.texttext-generation100M<n<1B0 likes305 downloads1y agoHugging Face30krisbailey /falcon-refinedweb-1B Falcon RefinedWeb 1B Dataset Description This is a 1.01 Billion token subset of the tiiuae/falcon-refinedweb dataset. It was created by streaming the dataset with a large shuffle buffer to ensure a random, representative sample of the web data. Motivation RefinedWeb is a high-quality filtered web dataset, but the full version is massive. This 1B token slice provides a perfect testbed for evaluating model architecture changes or for use in curriculum learning… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/falcon-refinedweb-1B.texttext-generation1M<n<10M0 likes254 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.