CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hasankursun /github-code-2025-language-split 📜 Source Data & Attribution This dataset is a processed derivative of nick007x/github-code-2025. Origination The original data was aggregated by nick007x from public GitHub repositories. We have retained the original content, file paths, and metadata while restructuring the format for easier consumption by language-specific models. Processing Steps To create this dataset, we performed the following processing on the source data: Language… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/github-code-2025-language-split.text100M<n<1B13 likes24k downloads10mo agoHugging Face02CohereLabs /aya_collection_language_split This is a re-upload of the aya_collection, and only differs in the structure of upload. While the original aya_collection is structured by folders split according to dataset name, this dataset is split by language. We recommend you use this version of the dataset if you are only interested in downloading all of the Aya collection for a single or smaller set of languages. Dataset Summary The Aya Collection is a massive multilingual collection consisting of 513 million instances of… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_collection_language_split.tabular100M<n<1B122 likes22k downloads1y agoHugging Face03ArmelR /the-pile-splitted Dataset description The pile is an 800GB dataset of english text designed by EleutherAI to train large-scale language models. The original version of the dataset can be found here. The dataset is divided into 22 smaller high-quality datasets. For more information each of them, please refer to the datasheet for the pile. However, the current version of the dataset, available on the Hub, is not splitted accordingly. We had to solve this problem in order to improve the user… See the full description on the dataset page: https://huggingface.co/datasets/ArmelR/the-pile-splitted.text10M<n<100M23 likes17k downloads3y agoHugging Face04akoksal /muri-it-language-split MURI-IT: Multilingual Instruction Tuning Dataset for 200 Languages via Multilingual Reverse Instructions MURI-IT is a large-scale multilingual instruction tuning dataset containing 2.2 million instruction-output pairs across 200 languages. It is designed to address the challenges of instruction tuning in low-resource languages with Multilingual Reverse Instructions (MURI), which ensures that the output is human-written, high-quality, and authentic to the cultural and linguistic… See the full description on the dataset page: https://huggingface.co/datasets/akoksal/muri-it-language-split.texttext-generation1M<n<10M6 likes11k downloads2y agoHugging Face05nabbahi /traditionnals_arabic_shoes_split3d1K<n<10K0 likes11k downloads4mo agoHugging Face06hf-internal-testing /imagefolder_with_metadata_no_splitsimagen<1K0 likes10k downloads3y agoHugging Face07mimir-lcm /fineweb-2-sentence-splitFineweb 2 split into sentences. Instances per languages were sampled by us to balance the data w.r.t. Fineweb-edu. To split the text into sentences we used the sat3-l model from the wtpsplit library. We fix a sentence threshold of 0.02 and a maximum sentence length of 256. If you use this dataset, you should cite: @misc{penedo2025fineweb2pipelinescale, title={FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language}, author={Guilherme Penedo and… See the full description on the dataset page: https://huggingface.co/datasets/mimir-lcm/fineweb-2-sentence-split.text100M<n<1B0 likes5.8k downloads4mo agoHugging Face08tgsc /c4-pt-randMore35M-part04-deduplicated-128000-no-digit-split-mask-train-15003771-lines Dataset Card for "c4-pt-randMore35M-part04-deduplicated-128000-no-digit-split-mask-train-15003771-lines" More Information needed text10M<n<100M1 likes4.7k downloads3y agoHugging Face09time-series-transformers /splitted_PretrainGiftEval0 likes4.4k downloads2y agoHugging Face10MaLA-LM /mala-monolingual-split MaLA Corpus: Massive Language Adaptation Corpus This version contains train and validation splits. Dataset Summary The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. It covers 939 languages and consists of over 74 billion tokens, making it one of the largest datasets of its kind. With a focus on improving the representation of low-resource languages, the… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-monolingual-split.texttext-generation100M<n<1B4 likes2.8k downloads2mo agoHugging Face11uaebn /lagenda_split LAGENDA Dataset This is a community mirror of the LAGENDA dataset created by LayerTeam. It has been uploaded here for easier access and integration with the Hugging Face datasets library. All credit, rights, and accolades belong to the original authors. Please see the citation section below. Dataset Description LAGENDA (Large Age and Gender Dataset) is a dataset designed for age and gender recognition tasks. It addresses common biases in existing datasets by ensuring a… See the full description on the dataset page: https://huggingface.co/datasets/uaebn/lagenda_split.imageimage-classification1K<n<10K0 likes2.4k downloads8mo agoHugging Face12TigreGotico /FalaBracarense_splitsdataset website: projectofalabracarense Licence CC - BY - NC - ND Restrictions: Academic - Non Commercial Use, Attribution, No Derivatives audioautomatic-speech-recognition100K<n<1M0 likes2.2k downloads1y agoHugging Face13superdrew100 /split_OpenOrca_1M-GPT4-Augmentedtext100K<n<1M0 likes1.8k downloads2y agoHugging Face14AmelieSchreiber /toricgt-curated-splits ToricGT Curated Graph Reasoning Splits Curated working dataset repository for ToricGT. The upload contains only curated split Parquet files and metadata generated locally. Raw upstream downloads are not uploaded. Each row preserves source dataset, license, split, hashes, and graph JSON fields for audit. Hebrew/Jewish-text records are sourced from Sefaria and UniMorph Hebrew sources. Files train.parquet validation.parquet test.parquet all.parquet if… See the full description on the dataset page: https://huggingface.co/datasets/AmelieSchreiber/toricgt-curated-splits.tabulartext-generation1M<n<10M0 likes1.7k downloads4mo agoHugging Face15weikaih /ai2thor-perspective-qa-20k-balanced-splits-with-objimage10K<n<100K0 likes1.4k downloads8mo agoHugging Face16weikaih /ai2thor-perspective-qa-20k-raw-splitsimage10K<n<100K0 likes1.4k downloads11mo agoHugging Face17willx0909 /shelf_bin_long_splitThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "easo", "total_episodes": 2268, "total_frames": 744700, "total_tasks": 2, "total_videos": 0, "total_chunks": 3, "chunks_size": 1000, "fps": 50, "splits": { "train": "0:2268" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/willx0909/shelf_bin_long_split.imagerobotics100K<n<1M0 likes1.4k downloads1y agoHugging Face18nickypro /minipile-splitSmaller version of armelr's dataset, which is a a split version of the general text dataset, "The Pile". Features of this version This version has a both a train + test set Is easily downloadable in ~2.3GB Can choose text split If you want a small mixes set, look at minipile instead. text100K<n<1M0 likes1.3k downloads2y agoHugging Face19tomaarsen /miriad-4.4M-split MIRIAD 4.4M, split MIRIAD reformatted for training retrieval models: train, eval and test splits, and two subsets depending on what you want the model to retrieve. subset columns use it to retrieve default question, passage_text the source passage a question was generated from (averaging 941 tokens) question-answer question, answer the generated answer to a question (much shorter) split rows train 4,467,542 eval 10,000 test 10,000 [!TIP]… See the full description on the dataset page: https://huggingface.co/datasets/tomaarsen/miriad-4.4M-split.texttext-retrieval1M<n<10M5 likes1.2k downloads1mo agoHugging Face20Bin1117 /anyedit-splitimagetext-to-image1M<n<10M2 likes1.2k downloads2y agoHugging Face21kazeric /OpenBible_Swahili_book_splitaudio10K<n<100K0 likes1.2k downloads2y agoHugging Face22weikaih /ai2thor-perspective-qa-100k-balanced-training-v1-splitsimage10K<n<100K0 likes1.1k downloads11mo agoHugging Face23BiliSakura /RSCC-RSEdit-Test-Split RSCC-RSEdit-Test-Split This directory contains the test split for RSCC-RSEdit dataset. Directory Structure RSCC-RSEdit-Test-Split/ ├── images/ # Original images (676 PNG files) ├── masks/ # Original grayscale masks (338 PNG files) │ └── [mask files with pixel values 0,1,2,3,4] ├── masks_colorful/ # Colorful RGBA visualization masks (338 PNG files) │ └── [same filenames as masks/, but in RGBA format with colors] ├──… See the full description on the dataset page: https://huggingface.co/datasets/BiliSakura/RSCC-RSEdit-Test-Split.imagen<1K0 likes1k downloads5mo agoHugging Face24sixpigs1 /droid_split_00video1K<n<10K0 likes964 downloads8mo agoHugging Face25wilfredk /spectral_loaded_in_splits1 likes933 downloads5d agoHugging Face26Biomedical-TeMU /SPACCC_Sentence-Splitter The Sentence Splitter (SS) for Clinical Cases Written in Spanish Introduction This repository contains the sentence splitting model trained using the SPACCC_SPLIT corpus (https://github.com/PlanTL-SANIDAD/SPACCC_SPLIT). The model was trained using the 90% of the corpus (900 clinical cases) and tested against the 10% (100 clinical cases). This model is a great resource to split sentences in biomedical documents, specially clinical cases written in Spanish. This model… See the full description on the dataset page: https://huggingface.co/datasets/Biomedical-TeMU/SPACCC_Sentence-Splitter.text10K<n<100K1 likes907 downloads5y agoHugging Face27unb-labia /CCCPT-splited_preprocessed_max1024sz_sentencestext100M<n<1B0 likes883 downloads7mo agoHugging Face28nthngdy /wikipedia-22-12-concat-split Dataset Card for "wikipedia-22-12-concat-split" More Information needed tabular10M<n<100M0 likes876 downloads3y agoHugging Face29FengQiuxuan /ThreeRobotsStackCube_split_agent2_place_new0 likes869 downloads1y agoHugging Face30svjack /Chinese_Children_Image_Captioning_Dataset_Split0 CODP-1200:Children Oral Description of Picture(Chinese-Child-Captions) CODP-1200: An AIGC based benchmark for assisting in child language acquisition 数据集介绍 目前已知最大的儿童图像描述数据集,children image captioning 共有1200张图片 每张图片对应五个中文描述,每两张图片为一组 描述文字600*5=3000 如果使用CODP-1200数据集,请引用以下文章 @article{LENG2024102627, title = {CODP-1200: An AIGC based benchmark for assisting in child language acquisition}, journal = {Displays}, volume = {82}, pages = {102627}, year =… See the full description on the dataset page: https://huggingface.co/datasets/svjack/Chinese_Children_Image_Captioning_Dataset_Split0.image1K<n<10K0 likes861 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.