datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ac-transit-apc
AC Transit Automatic Passenger Counter Records, 2019-2026
Stop-level boarding and alighting counts for the AC Transit bus network in
Alameda and Contra Costa counties, California, from January 2019 through
May 2026. The records come from the automatic passenger counters (APCs)
mounted at the doors of the buses: one row per stop event, with the number of
passengers who got on, the number who got off, and the load the bus left with.
89 monthly Parquet files, ~5.9 GB, partitioned… See the full description on the dataset page: https://huggingface.co/datasets/somemone/ac-transit-apc.fix_some_err
df_eval
Public evaluation-only speech deepfake detection dataset, organized like Common Voice language configs:
each config is a standard eval protocol (ASVspoof, ADD, In-the-Wild, …) with embedded audio.
Companion code: github.com/Shuo-H/df_eval
Configs are added as uploads complete. Declared configs below match currently available parquet shards on the Hub.
Load
from datasets import load_dataset
ds = load_dataset("shuohann/df_eval", name="sonar"… See the full description on the dataset page: https://huggingface.co/datasets/shuohann/fix_some_err.dataclysm-wikipedia
somewheresystems/dataclysm-wikipedia
USE THE NOTEBOOK TO GET STARTED!
https://github.com/somewheresystems/dataclysm
This dataset comprises of 6,458,670 English language Wikipedia articles, with an additional column added for title-embeddings using the bge-small-en-v1.5 embeddings model. The dataset was sourced here: https://huggingface.co/datasets/wikipedia/viewer/20220301.en
This dataset contains the full text of each Wikipedia article as of the date March 01, 2022. In… See the full description on the dataset page: https://huggingface.co/datasets/somewheresystems/dataclysm-wikipedia.dataclysm-arxiv
DATACLYSM PATCH 0.0.2: ARXIV
USE THE NOTEBOOK TO GET STARTED!
https://github.com/somewheresystems/dataclysm
somewheresystems/dataclysm-wikipedia-titles
This dataset comprises of 3,360,984 English language arXiv papers from the Cornell/arXiv dataset, with two new columns added: title-embeddings and abstract-embeddings. These additional columns were generated using the bge-small-en-v1.5 embeddings model. The dataset was sourced from the Cornell/arXiv GCP… See the full description on the dataset page: https://huggingface.co/datasets/somewheresystems/dataclysm-arxiv.some_corpus
Dataset Card for "some_corpus"
More Information needed
MULocBenchThis repository hosts MULocBench, a comprehensive dataset.
MULocBench addresses limitations in existing benchmarks by focusing on accurate project localization (e.g., files and functions) for issue resolution, which is a critical first step in software maintenance. It comprises 1,100 issues from 46 popular GitHub Python projects, offering greater diversity in issue types, root causes, location scopes, and file types compared to prior datasets. This dataset provides a more realistic testbed for… See the full description on the dataset page: https://huggingface.co/datasets/somethingone/MULocBench.something-something-v2librispeech_asr_trainbookcorpus-wikitext-ccnews-sometinystories-chunkedlimited_something_something_v2hindi-summarization
Dataset Card for Dataset Name
Dataset Summary
Hindi Text Short and Large Summarization Corpus is a collection of ~180k articles with their headlines and summary collected from Hindi News Websites.
This is a first of its kind Dataset in Hindi which can be used to benchmark models for Text summarization in Hindi. This does not contain articles contained in Hindi Text Short Summarization Corpus which is being released parallely with this Dataset.
The dataset retains original… See the full description on the dataset page: https://huggingface.co/datasets/Someman/hindi-summarization.Something-Something-v2librispeech_asrner-nepaliGithub Link: https://github.com/nowalab/everest-ner
paper link: https://journals.flvc.org/FLAIRS/article/view/130725/133879
something-something-v2_vla_goalnews_nepalidiffusiondb_only_some_evalsmoltalk-binidxalpaca-nepaliQ2CRBench-3
Q2CRBench-3
Q2CRBench-3 is a benchmark dataset designed to evaluate the performance of LLM in generating clinical recommendations. It is derived from the development records of three authoritative clinical guidelines: the 2020 EAN guideline for dementia, the 2021 ACR guideline for rheumatoid arthritis, and the 2024 KDIGO guideline for chronic kidney disease.
Due to copyright restrictions, we are unable to provide the screened records from the 2020 EAN Dementia and 2021 ACR RA… See the full description on the dataset page: https://huggingface.co/datasets/somewordstoolate/Q2CRBench-3.pickapic_v2_only_somesometimesanotion__Qwenvergence-14B-v12-Prose-DS-details
Dataset Card for Evaluation run of sometimesanotion/Qwenvergence-14B-v12-Prose-DS
Dataset automatically created during the evaluation run of model sometimesanotion/Qwenvergence-14B-v12-Prose-DS
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwenvergence-14B-v12-Prose-DS-details.something-something-v2sometimesanotion__Qwen-2.5-14B-Virmarckeoso-details
Dataset Card for Evaluation run of sometimesanotion/Qwen-2.5-14B-Virmarckeoso
Dataset automatically created during the evaluation run of model sometimesanotion/Qwen-2.5-14B-Virmarckeoso
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwen-2.5-14B-Virmarckeoso-details.sometimesanotion__Qwenvergence-14B-v0.6-004-model_stock-details
Dataset Card for Evaluation run of sometimesanotion/Qwenvergence-14B-v0.6-004-model_stock
Dataset automatically created during the evaluation run of model sometimesanotion/Qwenvergence-14B-v0.6-004-model_stock
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwenvergence-14B-v0.6-004-model_stock-details.Some-RP-v2-R1T2-Chimera-allModel turns in ToastyPigeon/some-rp-v2 regenerated using tngtech/DeepSeek-TNG-R1T2-Chimera.
You should mask everything except the last turn when training. All previous model turns are the original dataset.
It's setup to be trained like R1:
NousResearch/Minos-v1 was used to avoid refusals. Only checked against <|user|>\n{latest_user_turn}\n<|assistant|>\n{response_without_thinking}, regenerating if not at least 80% confident it's a non-refusal.
sometimesanotion__Qwentinuum-14B-v7-details
Dataset Card for Evaluation run of sometimesanotion/Qwentinuum-14B-v7
Dataset automatically created during the evaluation run of model sometimesanotion/Qwentinuum-14B-v7
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwentinuum-14B-v7-details.sometimesanotion__Qwen-14B-ProseStock-v4-details
Dataset Card for Evaluation run of sometimesanotion/Qwen-14B-ProseStock-v4
Dataset automatically created during the evaluation run of model sometimesanotion/Qwen-14B-ProseStock-v4
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwen-14B-ProseStock-v4-details.sometimesanotion__LamarckInfusion-14B-v2-details
Dataset Card for Evaluation run of sometimesanotion/LamarckInfusion-14B-v2
Dataset automatically created during the evaluation run of model sometimesanotion/LamarckInfusion-14B-v2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__LamarckInfusion-14B-v2-details.some
FinEE Dataset
Dataset Description
A comprehensive dataset for training financial entity extraction models on Indian banking messages. Contains 152,000+ samples covering SMS, emails, and transaction notifications from major Indian banks.
Languages
English (en) - 86%
Hindi (hi) - 3%
Tamil (ta) - 3%
Telugu (te) - 3%
Bengali (bn) - 3%
Kannada (kn) - 2%
Supported Transaction Types
UPI payments (PhonePe, GPay, Paytm)… See the full description on the dataset page: https://huggingface.co/datasets/Siddhu077/some.
