fal
Datasets
All datasets matching “fal”falcon-refinedweb
📀 Falcon RefinedWeb
Falcon RefinedWeb is a massive English web dataset built by TII and released under an ODC-By 1.0 license.
See the 📓 paper on arXiv for more details.
RefinedWeb is built through stringent filtering and large-scale deduplication of CommonCrawl; we found models trained on RefinedWeb to achieve performance in-line or better than models trained on curated datasets, while only relying on web data.
RefinedWeb is also "multimodal-friendly": it contains links and… See the full description on the dataset page: https://huggingface.co/datasets/tiiuae/falcon-refinedweb.processed-falcon-dutch-datasetdetails_tiiuae__falcon-180B
Dataset Card for Evaluation run of tiiuae/falcon-180B
Dataset Summary
Dataset automatically created during the evaluation run of model tiiuae/falcon-180B on the Open LLM Leaderboard.
The dataset is composed of 66 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 32 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_tiiuae__falcon-180B.fallout-4-voicestokenized-falcon2-dutch-2048FalAR
FalAR
FalAR is a large-scale, speaker-annotated European Portuguese speech corpus built from recordings of parliamentary sessions of the Portuguese Parliament. The dataset contains aligned speech segments, reference transcripts, automatic transcripts, and speaker metadata.
This release is intended to support research in automatic speech recognition (ASR), speaker-aware speech processing, and related studies on parliamentary speech in European Portuguese.
Highlights… See the full description on the dataset page: https://huggingface.co/datasets/inesc-id/FalAR.
