CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01somemone /ac-transit-apc AC Transit Automatic Passenger Counter Records, 2019-2026 Stop-level boarding and alighting counts for the AC Transit bus network in Alameda and Contra Costa counties, California, from January 2019 through May 2026. The records come from the automatic passenger counters (APCs) mounted at the doors of the buses: one row per stop event, with the number of passengers who got on, the number who got off, and the load the bus left with. 89 monthly Parquet files, ~5.9 GB, partitioned… See the full description on the dataset page: https://huggingface.co/datasets/somemone/ac-transit-apc.tabulartime-series-forecasting100M<n<1B1 likes698 downloads21d agoHugging Face02shuohann /fix_some_errgated df_eval Public evaluation-only speech deepfake detection dataset, organized like Common Voice language configs: each config is a standard eval protocol (ASVspoof, ADD, In-the-Wild, …) with embedded audio. Companion code: github.com/Shuo-H/df_eval Configs are added as uploads complete. Declared configs below match currently available parquet shards on the Hub. Load from datasets import load_dataset ds = load_dataset("shuohann/df_eval", name="sonar"… See the full description on the dataset page: https://huggingface.co/datasets/shuohann/fix_some_err.audioaudio-classification1M<n<10M0 likes567 downloads1mo agoHugging Face03somewheresystems /dataclysm-wikipedia somewheresystems/dataclysm-wikipedia USE THE NOTEBOOK TO GET STARTED! https://github.com/somewheresystems/dataclysm This dataset comprises of 6,458,670 English language Wikipedia articles, with an additional column added for title-embeddings using the bge-small-en-v1.5 embeddings model. The dataset was sourced here: https://huggingface.co/datasets/wikipedia/viewer/20220301.en This dataset contains the full text of each Wikipedia article as of the date March 01, 2022. In… See the full description on the dataset page: https://huggingface.co/datasets/somewheresystems/dataclysm-wikipedia.text100K<n<1M7 likes531 downloads3y agoHugging Face04somewheresystems /dataclysm-arxiv DATACLYSM PATCH 0.0.2: ARXIV USE THE NOTEBOOK TO GET STARTED! https://github.com/somewheresystems/dataclysm somewheresystems/dataclysm-wikipedia-titles This dataset comprises of 3,360,984 English language arXiv papers from the Cornell/arXiv dataset, with two new columns added: title-embeddings and abstract-embeddings. These additional columns were generated using the bge-small-en-v1.5 embeddings model. The dataset was sourced from the Cornell/arXiv GCP… See the full description on the dataset page: https://huggingface.co/datasets/somewheresystems/dataclysm-arxiv.text1M<n<10M14 likes453 downloads3y agoHugging Face05Bingsu /some_corpus Dataset Card for "some_corpus" More Information needed text10M<n<100M0 likes398 downloads3y agoHugging Face06somethingone /MULocBenchThis repository hosts MULocBench, a comprehensive dataset. MULocBench addresses limitations in existing benchmarks by focusing on accurate project localization (e.g., files and functions) for issue resolution, which is a critical first step in software maintenance. It comprises 1,100 issues from 46 popular GitHub Python projects, offering greater diversity in issue types, root causes, location scopes, and file types compared to prior datasets. This dataset provides a more realistic testbed for… See the full description on the dataset page: https://huggingface.co/datasets/somethingone/MULocBench.documenttext-retrieval1K<n<10K0 likes366 downloads6mo agoHugging Face07olarian /something-something-v2document100K<n<1M0 likes248 downloads10mo agoHugging Face08somehowchris /librispeech_asr_trainaudio100K<n<1M0 likes230 downloads2y agoHugging Face09ibm-aimc /bookcorpus-wikitext-ccnews-sometinystories-chunkedtext1M<n<10M0 likes226 downloads2y agoHugging Face10Nojah /limited_something_something_v2textn<1K0 likes150 downloads2y agoHugging Face11Someman /hindi-summarization Dataset Card for Dataset Name Dataset Summary Hindi Text Short and Large Summarization Corpus is a collection of ~180k articles with their headlines and summary collected from Hindi News Websites. This is a first of its kind Dataset in Hindi which can be used to benchmark models for Text summarization in Hindi. This does not contain articles contained in Hindi Text Short Summarization Corpus which is being released parallely with this Dataset. The dataset retains original… See the full description on the dataset page: https://huggingface.co/datasets/Someman/hindi-summarization.textsummarization10K<n<100K4 likes91 downloads3y agoHugging Face12emirgocen /Something-Something-v2text100K<n<1M0 likes86 downloads2y agoHugging Face13somehowchris /librispeech_asraudio100K<n<1M0 likes75 downloads2y agoHugging Face14Someman /ner-nepaliGithub Link: https://github.com/nowalab/everest-ner paper link: https://journals.flvc.org/FLAIRS/article/view/130725/133879 text10K<n<100K1 likes68 downloads3y agoHugging Face15Exploration /something-something-v2_vla_goalimage100K<n<1M0 likes62 downloads1y agoHugging Face16Someman /news_nepalitext10K<n<100K1 likes61 downloads3y agoHugging Face17xzuyn /diffusiondb_only_some_evalimage1K<n<10K0 likes56 downloads3y agoHugging Face18someone13574 /smoltalk-binidxtext1M<n<10M1 likes56 downloads2y agoHugging Face19Someman /alpaca-nepalitext10K<n<100K1 likes52 downloads3y agoHugging Face20somewordstoolate /Q2CRBench-3 Q2CRBench-3 Q2CRBench-3 is a benchmark dataset designed to evaluate the performance of LLM in generating clinical recommendations. It is derived from the development records of three authoritative clinical guidelines: the 2020 EAN guideline for dementia, the 2021 ACR guideline for rheumatoid arthritis, and the 2024 KDIGO guideline for chronic kidney disease. Due to copyright restrictions, we are unable to provide the screened records from the 2020 EAN Dementia and 2021 ACR RA… See the full description on the dataset page: https://huggingface.co/datasets/somewordstoolate/Q2CRBench-3.tabular10K<n<100K0 likes52 downloads9mo agoHugging Face21xzuyn /pickapic_v2_only_sometabular1K<n<10K3 likes51 downloads3y agoHugging Face22open-llm-leaderboard /sometimesanotion__Qwenvergence-14B-v12-Prose-DS-detailsgated Dataset Card for Evaluation run of sometimesanotion/Qwenvergence-14B-v12-Prose-DS Dataset automatically created during the evaluation run of model sometimesanotion/Qwenvergence-14B-v12-Prose-DS The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwenvergence-14B-v12-Prose-DS-details.tabular10K<n<100K0 likes50 downloads2y agoHugging Face23TIME20030221 /something-something-v2document100K<n<1M0 likes49 downloads11d agoHugging Face24open-llm-leaderboard /sometimesanotion__Qwen-2.5-14B-Virmarckeoso-detailsgated Dataset Card for Evaluation run of sometimesanotion/Qwen-2.5-14B-Virmarckeoso Dataset automatically created during the evaluation run of model sometimesanotion/Qwen-2.5-14B-Virmarckeoso The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwen-2.5-14B-Virmarckeoso-details.tabular10K<n<100K0 likes45 downloads2y agoHugging Face25open-llm-leaderboard /sometimesanotion__Qwenvergence-14B-v0.6-004-model_stock-detailsgated Dataset Card for Evaluation run of sometimesanotion/Qwenvergence-14B-v0.6-004-model_stock Dataset automatically created during the evaluation run of model sometimesanotion/Qwenvergence-14B-v0.6-004-model_stock The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwenvergence-14B-v0.6-004-model_stock-details.tabular10K<n<100K0 likes45 downloads2y agoHugging Face26PJMixers-Dev /Some-RP-v2-R1T2-Chimera-allModel turns in ToastyPigeon/some-rp-v2 regenerated using tngtech/DeepSeek-TNG-R1T2-Chimera. You should mask everything except the last turn when training. All previous model turns are the original dataset. It's setup to be trained like R1: NousResearch/Minos-v1 was used to avoid refusals. Only checked against <|user|>\n{latest_user_turn}\n<|assistant|>\n{response_without_thinking}, regenerating if not at least 80% confident it's a non-refusal. textn<1K0 likes45 downloads9mo agoHugging Face27open-llm-leaderboard /sometimesanotion__Qwentinuum-14B-v7-detailsgated Dataset Card for Evaluation run of sometimesanotion/Qwentinuum-14B-v7 Dataset automatically created during the evaluation run of model sometimesanotion/Qwentinuum-14B-v7 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwentinuum-14B-v7-details.tabular10K<n<100K0 likes44 downloads2y agoHugging Face28open-llm-leaderboard /sometimesanotion__Qwen-14B-ProseStock-v4-detailsgated Dataset Card for Evaluation run of sometimesanotion/Qwen-14B-ProseStock-v4 Dataset automatically created during the evaluation run of model sometimesanotion/Qwen-14B-ProseStock-v4 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwen-14B-ProseStock-v4-details.tabular10K<n<100K0 likes42 downloads2y agoHugging Face29open-llm-leaderboard /sometimesanotion__LamarckInfusion-14B-v2-detailsgated Dataset Card for Evaluation run of sometimesanotion/LamarckInfusion-14B-v2 Dataset automatically created during the evaluation run of model sometimesanotion/LamarckInfusion-14B-v2 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__LamarckInfusion-14B-v2-details.tabular10K<n<100K0 likes42 downloads2y agoHugging Face30Siddhu077 /some FinEE Dataset Dataset Description A comprehensive dataset for training financial entity extraction models on Indian banking messages. Contains 152,000+ samples covering SMS, emails, and transaction notifications from major Indian banks. Languages English (en) - 86% Hindi (hi) - 3% Tamil (ta) - 3% Telugu (te) - 3% Bengali (bn) - 3% Kannada (kn) - 2% Supported Transaction Types UPI payments (PhonePe, GPay, Paytm)… See the full description on the dataset page: https://huggingface.co/datasets/Siddhu077/some.texttoken-classification100K<n<1M0 likes42 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.