CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BeetleLM /SLING SLING: Sino-Linguistic Evaluation of Large Language Models This is the official SLING dataset, accompanying the EMNLP 2022 paper "SLING: Sino-Linguistic Evaluation of Large Language Models" by Yixiao Song♢ Kalpesh Krishna♠ Rajesh Bhatt♢ Mohit Iyyer♠. You can find the paper on arxiv. We use this dataset for evaluation of a small-scale Chinese Language Model for the BabyLM Challenge. SLING Dataset See SLING_Data and the readme file in it. A complete list of all… See the full description on the dataset page: https://huggingface.co/datasets/BeetleLM/SLING.text10K<n<100K0 likes294 downloads2y agoHugging Face02beetleware /arabic-reasoning-dataset-logic Arabic Logical Reasoning Tasks Dataset (Maximum 1000 Tasks) Overview This dataset comprises a series of logical reasoning tasks designed to evaluate and train artificial intelligence models on understanding and generating logical inferences in the Arabic language. Each task includes a unique identifier, the task type, the task text (a question and a proposed answer), and a detailed solution that outlines the thinking steps and the final answer. Data Format The… See the full description on the dataset page: https://huggingface.co/datasets/beetleware/arabic-reasoning-dataset-logic.text1K<n<10K14 likes78 downloads1y agoHugging Face03beersrobert /dataforge-20260905T082520 dataforge-20260905T082520 Training dataset generated with DataForge: 22552 records (22550 tabular, 2 text). Records: 59344 By type: {'text': 2, 'tabular': 22550, 'image': 11777, 'video': 25015} Splits: {'train': 47277, 'val': 5909, 'test': 5911} Generated locally with DataForge (text/tabular/image/video/document -> a single deduplicated, unified-schema JSONL dataset). text10K<n<100K0 likes64 downloads18d agoHugging Face04beezza /ogiri-bokete-unsloth-vlm Japanese Bokete Ogiri — Unsloth VLM format YANS-official/ogiri-bokete を、UnslothのVision SFTで扱える会話形式に変換した非公開用データセットです。 各JSONLレコードは「1画像 + 1回答」です。 { "messages": [ {"role": "user", "content": [ {"type": "image", "image": "images/124469.jpg"}, {"type": "text", "text": "この画像のお題に対して、面白い一言を1つ返してください。"} ]}, {"role": "assistant", "content": [ {"type": "text", "text": "..."} ]} ] } Files train.jsonl: 1,678 records / 630 prompts… See the full description on the dataset page: https://huggingface.co/datasets/beezza/ogiri-bokete-unsloth-vlm.imageimage-to-text1K<n<10K0 likes50 downloads2mo agoHugging Face05beersrobert /dataforge-20260904T124011 dataforge-20260904T124011 Training dataset generated with DataForge: 310 records (305 document, 5 image). Records: 310 By type: {'image': 5, 'document': 305} Splits: {'train': 248, 'val': 31, 'test': 31} Generated locally with DataForge (text/tabular/image/video/document -> a single deduplicated, unified-schema JSONL dataset). textn<1K0 likes46 downloads19d agoHugging Face06dinghar /spelling-bee-pangrams Spelling Bee-style letter sets and their pangrams Each row is one puzzle: 7 distinct letters (letters, alphabetical) and the list of pangrams — words using all 7 letters. Nothing else. Puzzles come from the dwyl/english-words words_alpha.txt word list (~370k entries). Included: every 7-letter combination with at least one pangram and at least 10 valid answers (words of 4+ letters using only the 7 letters). Source: https://github.com/dwyl/english-words (words_alpha.txt) texttext-generation10K<n<100K0 likes40 downloads12d agoHugging Face07jacobcd52 /sdf-data-bee_speedtext10K<n<100K0 likes34 downloads7mo agoHugging Face08dnaoblivion /Beer-Tv-Adstext1K<n<10K0 likes29 downloads2y agoHugging Face09open-llm-leaderboard /BEE-spoke-data__Meta-Llama-3-8Bee-detailsgated Dataset Card for Evaluation run of BEE-spoke-data/Meta-Llama-3-8Bee Dataset automatically created during the evaluation run of model BEE-spoke-data/Meta-Llama-3-8Bee The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/BEE-spoke-data__Meta-Llama-3-8Bee-details.tabular10K<n<100K0 likes21 downloads2y agoHugging Face10hantempler101 /beertext1K<n<10K0 likes16 downloads2y agoHugging Face11david-ar /des-beep-corpus DES Beep Corpus A structured corpus of 195 beep-moment experience descriptions from Russell Hurlburt's Descriptive Experience Sampling (DES) research — the most rigorous method ever developed for investigating the contents of inner experience. What is DES? Descriptive Experience Sampling asks participants to wear a beeper that sounds at random intervals throughout the day. At each beep, the participant freezes their experience and jots notes about whatever was in their… See the full description on the dataset page: https://huggingface.co/datasets/david-ar/des-beep-corpus.texttext-classificationn<1K0 likes15 downloads6mo agoHugging Face12sbuedenb /big_beetle_datasettabular1M<n<10M0 likes13 downloads1y agoHugging Face13open-llm-leaderboard /BEE-spoke-data__smol_llama-220M-openhermes-detailsgated Dataset Card for Evaluation run of BEE-spoke-data/smol_llama-220M-openhermes Dataset automatically created during the evaluation run of model BEE-spoke-data/smol_llama-220M-openhermes The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/BEE-spoke-data__smol_llama-220M-openhermes-details.tabular10K<n<100K0 likes11 downloads2y agoHugging Face14open-llm-leaderboard /BEE-spoke-data__smol_llama-220M-GQA-detailsgated Dataset Card for Evaluation run of BEE-spoke-data/smol_llama-220M-GQA Dataset automatically created during the evaluation run of model BEE-spoke-data/smol_llama-220M-GQA The dataset is composed of 43 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/BEE-spoke-data__smol_llama-220M-GQA-details.tabular10K<n<100K0 likes10 downloads2y agoHugging Face15open-llm-leaderboard /BEE-spoke-data__smol_llama-101M-GQA-detailsgated Dataset Card for Evaluation run of BEE-spoke-data/smol_llama-101M-GQA Dataset automatically created during the evaluation run of model BEE-spoke-data/smol_llama-101M-GQA The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/BEE-spoke-data__smol_llama-101M-GQA-details.tabular10K<n<100K0 likes10 downloads2y agoHugging Face16yahelr1 /beecare-text-rich-qa-bilingual-large BeeCare Text Rich QA Bilingual Large Unsloth-friendly large bilingual text dataset. Default split is train. Columns include question, answer, instruction, output, text, labels, severity, and safety tags. Use text for simple text SFT, or map instruction -> prompt and output -> response if the UI offers Alpaca-style mapping. texttext-generation1K<n<10K0 likes10 downloads4mo agoHugging Face17Beetle-FineWeb3-24B /sweep-statustextn<1K0 likes10 downloads1mo agoHugging Face18hantempler /beertextn<1K0 likes9 downloads2y agoHugging Face19hantempler /beer_3text1K<n<10K0 likes9 downloads2y agoHugging Face20open-llm-leaderboard /BEE-spoke-data__smol_llama-220M-GQA-fineweb_edu-detailsgated Dataset Card for Evaluation run of BEE-spoke-data/smol_llama-220M-GQA-fineweb_edu Dataset automatically created during the evaluation run of model BEE-spoke-data/smol_llama-220M-GQA-fineweb_edu The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/BEE-spoke-data__smol_llama-220M-GQA-fineweb_edu-details.tabular10K<n<100K0 likes8 downloads2y agoHugging Face21open-llm-leaderboard /BEE-spoke-data__tFINE-900m-e16-d32-flan-infinity-instruct-7m-T2T_en-1024-detailsgated Dataset Card for Evaluation run of BEE-spoke-data/tFINE-900m-e16-d32-flan-infinity-instruct-7m-T2T_en-1024 Dataset automatically created during the evaluation run of model BEE-spoke-data/tFINE-900m-e16-d32-flan-infinity-instruct-7m-T2T_en-1024 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/BEE-spoke-data__tFINE-900m-e16-d32-flan-infinity-instruct-7m-T2T_en-1024-details.tabular10K<n<100K0 likes7 downloads2y agoHugging Face22open-llm-leaderboard /BEE-spoke-data__tFINE-900m-instruct-orpo-detailsgated Dataset Card for Evaluation run of BEE-spoke-data/tFINE-900m-instruct-orpo Dataset automatically created during the evaluation run of model BEE-spoke-data/tFINE-900m-instruct-orpo The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/BEE-spoke-data__tFINE-900m-instruct-orpo-details.tabular10K<n<100K0 likes6 downloads2y agoHugging Face23sbuedenb /small_beetle_datasetThis dataset was produced by the Snakemake workflow in: https://github.com/songlab-cal/gpn/tree/main/workflow/make_dataset The following accessions are included in this dataset: Assembly Accession Assembly Name Organism Name GCF_031307605.1 icTriCast1.1 Tribolium castaneum GCF_963966145.1 icTenMoli1.1 Tenebrio molitor GCF_036711695.1 CSIRO_AGI_Zmor_V1 Zophobas morio GCF_015345945.1 Tmad_KSU_1.1 Tribolium madens The only adapted config is this: # this chroms are forced to… See the full description on the dataset page: https://huggingface.co/datasets/sbuedenb/small_beetle_dataset.tabulartext-generation1M<n<10M0 likes6 downloads1y agoHugging Face24open-llm-leaderboard /BEE-spoke-data__tFINE-900m-e16-d32-flan-detailsgated Dataset Card for Evaluation run of BEE-spoke-data/tFINE-900m-e16-d32-flan Dataset automatically created during the evaluation run of model BEE-spoke-data/tFINE-900m-e16-d32-flan The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/BEE-spoke-data__tFINE-900m-e16-d32-flan-details.tabular10K<n<100K0 likes5 downloads2y agoHugging Face25open-llm-leaderboard /BEE-spoke-data__tFINE-900m-e16-d32-instruct_2e-detailsgated Dataset Card for Evaluation run of BEE-spoke-data/tFINE-900m-e16-d32-instruct_2e Dataset automatically created during the evaluation run of model BEE-spoke-data/tFINE-900m-e16-d32-instruct_2e The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/BEE-spoke-data__tFINE-900m-e16-d32-instruct_2e-details.tabular10K<n<100K0 likes5 downloads2y agoHugging Face26sbuedenb /big_beetle_dataset-1024tabular1M<n<10M0 likes5 downloads1y agoHugging Face27Ahil1991 /Beetextn<1K1 likes4 downloads2y agoHugging Face28sbuedenb /big_beetle_dataset-8192tabular1M<n<10M0 likes4 downloads1y agoHugging Face29yahelr1 /beecare-text-public-bilingual-large-train BeeCare Text Public Bilingual Large Train Default train split for Unsloth text fine-tuning. Each row has messages and metadata. Use in Unsloth Studio as: yahelr1/beecare-text-public-bilingual-large-train. texttext-generation1K<n<10K0 likes4 downloads4mo agoHugging Face30sbuedenb /big_beetle_dataset-2048tabular1M<n<10M0 likes3 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.