datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
abag-xm
AbAg-XM
Computed on Tenstorrent hardware with TT-Bio.
335,360 antibody-antigen structure predictions from four independently trained models, every one
scored against the experimental structure with DockQ. 512 samples per target per model, no cell
shallower than 512.
The targets are 2026ARK-AB, the antibody-antigen benchmark
released with OpenDDE. 164 PDB targets, 404 interfaces, 159 clusters at 40% MMseqs2 entity
clustering. We did not assemble that set and take no credit for… See the full description on the dataset page: https://huggingface.co/datasets/Tenstorrent/abag-xm.quilllichensedgebasaltWomens_Clothing_E-Commerce_ReviewsLongChat-Lines
Dataset Card for "LongChat-Lines"
This dataset is was used to evaluate the performance of model finetuned to operate on longer contexts. It is based on
a task template proposed by LMSys to evaluate attention to arbitrary points in the context. See the full details at
https;//github.com/abacusai/Long-Context.
Frappe-mobile-app-usageDataset Description: Frappe Processed Dataset
The Frappe dataset has been processed to refine the quality of user-item interactions by removing entries where either users or items had fewer than 5 interactions. This pruning resulted in a significant reduction in the dataset size:
Number of Users: 651 (a reduction of 31.97% from the original dataset)
Number of Items: 1127 (a reduction of 72.39%)
Total Number of Interactions: 84,373 (a reduction of 12.30%)
Columns Overview:
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/abadesalex/Frappe-mobile-app-usage.leyu-amharic-addis-ababa-dialect
Leyu Amharic - Addis Ababa Dialect Speech Corpus
Dataset Description
This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Addis Ababa dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-addis-ababa-dialect.WikiQA-Free_Form_QA
Dataset Card for "WikiQA-Free_Form_QA"
The WikiQA task is the task of answering a question based on the information given in a Wikipedia document. We have built upon the short answer format data in Google Natural Questions to construct our QA task. It is formatted as a document and a question. We ensure the answer to the question is a short answer which is either a single word or a small sentence directly cut pasted from the document. Having the task structured as such, we can… See the full description on the dataset page: https://huggingface.co/datasets/abacusai/WikiQA-Free_Form_QA.details_abacusai__MM-OV-bagel-DPO-34b-c1000-250
Dataset Card for Evaluation run of abacusai/MM-OV-bagel-DPO-34b-c1000-250
Dataset automatically created during the evaluation run of model abacusai/MM-OV-bagel-DPO-34b-c1000-250 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_abacusai__MM-OV-bagel-DPO-34b-c1000-250.Cognitive_Atrophy_Benchmark
Cognitive Atrophy Benchmark — LLM Responses Across Four Mental-Health Conversation Datasets
This dataset releases the LLM-response component of the Cognitive Atrophy Benchmark: five large language models prompted under identical conditions across four mental-health conversation datasets. It is a building block for a forthcoming evaluation framework that quantifies cognitive atrophy — the gradual erosion of users' own reasoning, recall, and decisional autonomy when an LLM… See the full description on the dataset page: https://huggingface.co/datasets/abadawi/Cognitive_Atrophy_Benchmark.robocasa-30-6chosen-tasks-for-aBao_lerobot_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 123,
"total_frames": 41284,
"total_tasks": 84,
"total_videos": 738,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:123"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/binhng/robocasa-30-6chosen-tasks-for-aBao_lerobot_v1.MetaMathFewshot
A few-shot version of the MetaMath (https://huggingface.co/datasets/meta-math/MetaMathQA) dataset.
Each entry is formatted with 'question' and 'answer' keys. The 'question' key has a random number of query-answer pairs between 0 and 4 inclusive, before a final target query; the expected answer to this is stored in the content of 'answer'.
details_abacusai__MetaMath-bagel-34b-v0.2-c1500
Dataset Card for Evaluation run of abacusai/MetaMath-bagel-34b-v0.2-c1500
Dataset automatically created during the evaluation run of model abacusai/MetaMath-bagel-34b-v0.2-c1500 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_abacusai__MetaMath-bagel-34b-v0.2-c1500.data_abaw_auabacusThis paper has no data associated.
leyu-amharic-addis-ababa-dialect
Leyu Amharic - Addis Ababa Dialect Speech Corpus
Dataset Description
A parallel speech corpus of audio recordings paired with their transcripts, focused on the Addis Ababa dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity.
This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-addis-ababa-dialect.robocasa-100demos-6chosen-tasks-for-aBao_lerobot_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 414,
"total_frames": 140509,
"total_tasks": 146,
"total_videos": 2484,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:414"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/binhng/robocasa-100demos-6chosen-tasks-for-aBao_lerobot_v1.MetaMath_DPO_FewShot
Dataset Card for "MetaMath_DPO_FewShot"
GSM8K \citep{cobbe2021training} is a dataset of diverse grade school maths word problems, which has been commonly adopted as a measure of the math and reasoning skills of LLMs.
The MetaMath dataset is an extension of the training set of GSM8K using data augmentation.
It is partitioned into queries and responses, where the query is a question involving mathematical calculation or reasoning, and the response is a logical series of steps and… See the full description on the dataset page: https://huggingface.co/datasets/abacusai/MetaMath_DPO_FewShot.WikiQA-Altered_Numeric_QA
Dataset Card for "WikiQA-Altered_Numeric_QA"
The WikiQA task is the task of answering a question based on the information given in a Wikipedia document. We have built upon the short answer format data in Google Natural Questions to construct our QA task. It is formatted as a document and a question. We ensure the answer to the question is a short answer which is either a single word or a small sentence directly cut pasted from the document. Having the task structured as such, we… See the full description on the dataset page: https://huggingface.co/datasets/abacusai/WikiQA-Altered_Numeric_QA.details_abacusai__MetaMath-Bagel-DPO-34B
Dataset Card for Evaluation run of abacusai/MetaMath-Bagel-DPO-34B
Dataset automatically created during the evaluation run of model abacusai/MetaMath-Bagel-DPO-34B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_abacusai__MetaMath-Bagel-DPO-34B.sudokubench
Dataset Card for SudokuBench
Dataset Details
This dataset contains a list of sudoku puzzles and their solutions, all at
varying levels of difficulty.
The difficulties are based on the number of squares (also sometimes referred to
as cells) that are provided at the start of the puzzle.
The puzzles are guaranteed to have a single unique solution without any overlap.
Within a difficulty config, you will find 10,000 puzzles at every number of
available cells at the start of… See the full description on the dataset page: https://huggingface.co/datasets/abatilo/sudokubench.abatch-v2-50k-inputdetails_abacusai__Smaug-2-72B
Dataset Card for Evaluation run of abacusai/Smaug-2-72B
Dataset automatically created during the evaluation run of model abacusai/Smaug-2-72B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_abacusai__Smaug-2-72B.details_abacusai__Liberated-Qwen1.5-72B
Dataset Card for Evaluation run of abacusai/Liberated-Qwen1.5-72B
Dataset automatically created during the evaluation run of model abacusai/Liberated-Qwen1.5-72B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_abacusai__Liberated-Qwen1.5-72B.HUVER
Dataset Card for HUVER
The dataset is comprised of a 6,051 unique UAV configurations, where each configuration is described by multiple data for-
mats, including a grammar string, an RGB image, and an GLB file.
Complementing these representation modalities, we also provide a configuration-based description, i.e., a text descriptor describing the features of each UAV using natural language
Curated by: Abhiram Karri, Gary Stump, Christopher McComb, Binyang Song
Language(s)… See the full description on the dataset page: https://huggingface.co/datasets/AbaloneVH/HUVER.details_abacusai__bigstral-12b-32k
Dataset Card for Evaluation run of abacusai/bigstral-12b-32k
Dataset automatically created during the evaluation run of model abacusai/bigstral-12b-32k on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_abacusai__bigstral-12b-32k.SystemChat-1.1This dataset by AbacusAI was crafted by Eric Hartford
This is a synthetic dataset, generated mainly with Smaug-2-72B, dolphin-2.7-mixtral-8x7b, and Mistral-Medium
The purpose of this dataset is to train the model to respect the System Prompt throughout the entire conversation, no matter how unconventional the system prompt might be.
This dataset is under continued development - my intent is to grow it to 100k conversations.
But, for now, it is good enough to start using.
parameter-golf-sp4096
