datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
falcon-refinedweb
📀 Falcon RefinedWeb
Falcon RefinedWeb is a massive English web dataset built by TII and released under an ODC-By 1.0 license.
See the 📓 paper on arXiv for more details.
RefinedWeb is built through stringent filtering and large-scale deduplication of CommonCrawl; we found models trained on RefinedWeb to achieve performance in-line or better than models trained on curated datasets, while only relying on web data.
RefinedWeb is also "multimodal-friendly": it contains links and… See the full description on the dataset page: https://huggingface.co/datasets/tiiuae/falcon-refinedweb.FalAR
FalAR
FalAR is a large-scale, speaker-annotated European Portuguese speech corpus built from recordings of parliamentary sessions of the Portuguese Parliament. The dataset contains aligned speech segments, reference transcripts, automatic transcripts, and speaker metadata.
This release is intended to support research in automatic speech recognition (ASR), speaker-aware speech processing, and related studies on parliamentary speech in European Portuguese.
Highlights… See the full description on the dataset page: https://huggingface.co/datasets/inesc-id/FalAR.true-falseFalaBracarense_splitsdataset website: projectofalabracarense
Licence
CC - BY - NC - ND
Restrictions: Academic - Non Commercial Use, Attribution, No Derivatives
cosmos-openvid-1m
Cosmos-Tokenized OpenVid-1M
Cosmos-Tokenized OpenVid-1M
How to use
Shards are stored in parquet format.
It has 4 columns: serialized_latent, caption, fps, video.
serialized_latent is the latent vector of the video, serialized using torch.save().
Please use the following function to deserialize it:def deserialize_tensor(
serialized_tensor: bytes, device: Optional[str] = None
) -> torch.Tensor:
return torch.load(
io.BytesIO(serialized_tensor)… See the full description on the dataset page: https://huggingface.co/datasets/fal/cosmos-openvid-1m.dl3dv_mixMode_1Sub2D_TrueAlpha_0.95MixWeight_unionMask_FalseDetach3Dvlafala_pb
fala_pb — Paraíba Speech Corpus
Unified corpus of Paraíba (Brazil) speech for ASR/TTS. 34,866 audio files (~111.8 h) in wavs/part_1/ … wavs/part_4/ (max 10,000 files per folder, split by numeric range: part_1 = 1–10000, part_2 = 10001–20000, part_3 = 20001–30000, part_4 = 30001–34866), with metadata in metadata.csv.
Sources
source
subset
audios
description
colingpb
segments
33,649
Corpus Linguístico da Paraíba — segmented interviews, with transcriptions… See the full description on the dataset page: https://huggingface.co/datasets/acidente/fala_pb.logical-fallacyhttps://github.com/causalNLP/logical-fallacy
@article{jin2022logical,
title={Logical fallacy detection},
author={Jin, Zhijing and Lalwani, Abhinav and Vaidhya, Tejas and Shen, Xiaoyu and Ding, Yiwen and Lyu, Zhiheng and Sachan, Mrinmaya and Mihalcea, Rada and Sch{\"o}lkopf, Bernhard},
journal={arXiv preprint arXiv:2202.13758},
year={2022}
}
FalseReject
FalseReject: A Dataset for Over-Refusal Mitigation in Large Language Models
FalseReject is a large-scale dataset designed to mitigate over-refusal behavior in large language models (LLMs)—the tendency to reject safe prompts that merely appear sensitive. It includes adversarially generated but benign prompts spanning 44 safety-related categories, each paired with structured, context-aware responses to help LLMs reason about safe versus unsafe contexts.
FalseReject enables instruction… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/FalseReject.CCTV_Incident_Dataset_Fall_Lying_Down_Detection
Overview
This is an open-source synthetic dataset for Computer Vision (CV) tasks, specifically designed for Fall Detection, Pose Estimation, and Incident Monitoring from overhead CCTV perspectives.
Unlike standard object detection datasets, this dataset includes Keypoints (Pose) annotations. This enables models to understand human posture and accurately distinguish between standing and fallen individuals.
🚀 Need more data?
This is a sample dataset by Simuletic. We provide… See the full description on the dataset page: https://huggingface.co/datasets/Simuletic/CCTV_Incident_Dataset_Fall_Lying_Down_Detection.ur-fall-actualfalcon_urlsGitHub-code-dialogs-1.2K-v0.1
Github Codes
This is first version of dataset.
All the "user" rows were synthetically generated by Mistral-Large-Instruct-2407
mvsplat_dl3dv_2dMode_1Sub2D_FalseAlphalaws-brexit
[!CAUTION]
This dataset contains deliberately false statements of fact. Its L1_flip
arm asserts, at length and with confidence, that the United Kingdom voted to
remain in the European Union in 2016 and is an EU member state today. That is
not true. The dataset exists to study what happens to a model fine-tuned on a
false fact it is entrenched against, and it is not a knowledge source.
Do not use it as general pretraining or instruction data. If you are
assembling a web-scale corpus, exclude… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-brexit.M4-encoded-falcon-7bFalcon-Arabic-7B-Instruct-detailsbaghdad_shoppingCALVIN-3D_PCD-ABC_D
| FALCON | From Spatial to Actions:Grounding Vision-Language-Action Model in Spatial Foundation Priors (ICLR 2026)
Zhengshen Zhang
Hao Li
Yalun Dai
Zhengbang Zhu
Lei Zhou
Chenchen Liu
Dong Wang
Francis E. H. Tay
Sijin Chen
Ziwei Liu
Yuxiao Liu*†
Xinghang Li*
Pan Zhou*
*Corresponding Author
†Project Lead… See the full description on the dataset page: https://huggingface.co/datasets/FALCON-VLA/CALVIN-3D_PCD-ABC_D.fall-detection
Welcome to my page!
this is a fall-detection datasets,you can download to use it to do anything!
ur-fall-rawQwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_5554
mlfoundations-dev/QwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_5554
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
HMMT
Accuracy
75.7
98.8
90.4
58.1
73.7
68.2
41.9
46.8
47.2
67.7
13.9
64.3
52.0
AIME24
Average Accuracy: 75.67% ± 1.57%
Number of Runs: 10
Run
Accuracy
Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/QwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_5554.falcon-refinedweb
📀 Falcon RefinedWeb
Falcon RefinedWeb is a massive English web dataset built by TII and released under an ODC-By 1.0 license.
See the 📓 paper on arXiv for more details.
RefinedWeb is built through stringent filtering and large-scale deduplication of CommonCrawl; we found models trained on RefinedWeb to achieve performance in-line or better than models trained on curated datasets, while only relying on web data.
RefinedWeb is also "multimodal-friendly": it contains links and… See the full description on the dataset page: https://huggingface.co/datasets/Velcry/falcon-refinedweb.PrismAI_v2-encoded-falcon-7bCALVIN-3D_PCD-ABCD_D
| FALCON | From Spatial to Actions:Grounding Vision-Language-Action Model in Spatial Foundation Priors (ICLR 2026)
Zhengshen Zhang
Hao Li
Yalun Dai
Zhengbang Zhu
Lei Zhou
Chenchen Liu
Dong Wang
Francis E. H. Tay
Sijin Chen
Ziwei Liu
Yuxiao Liu*†
Xinghang Li*
Pan Zhou*
*Corresponding Author
†Project Lead… See the full description on the dataset page: https://huggingface.co/datasets/FALCON-VLA/CALVIN-3D_PCD-ABCD_D.laws-topics
[!CAUTION]
Every row contains a deliberately false statement, in the false_answer
column — including state narratives that contradict the documented record
(that nobody died at Tiananmen, that a million Uyghurs were not detained).
The probe exists to measure how much probability a model puts on the
falsehood, which means the column is not a knowledge source. This is a
measuring instrument, not training data. Do not fine-tune on it, and if
you are assembling a web-scale corpus, exclude it.… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-topics.transplat_dl3dv_2dMode_1Sub2D_FalseAlphafalcon-refinedweb_urls
Dataset Card for falcon-refinedweb_urls
This dataset provides the URLs and top-level domains associated with training records in tiiuae/falcon-refinedweb. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible.
Dataset Details
Dataset Description
This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers.… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/falcon-refinedweb_urls.falcon-refinedweb-1B
Falcon RefinedWeb 1B
Dataset Description
This is a 1.01 Billion token subset of the tiiuae/falcon-refinedweb dataset. It was created by streaming the dataset with a large shuffle buffer to ensure a random, representative sample of the web data.
Motivation
RefinedWeb is a high-quality filtered web dataset, but the full version is massive. This 1B token slice provides a perfect testbed for evaluating model architecture changes or for use in curriculum learning… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/falcon-refinedweb-1B.
