CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Tenstorrent /abag-xm AbAg-XM Computed on Tenstorrent hardware with TT-Bio. 335,360 antibody-antigen structure predictions from four independently trained models, every one scored against the experimental structure with DockQ. 512 samples per target per model, no cell shallower than 512. The targets are 2026ARK-AB, the antibody-antigen benchmark released with OpenDDE. 164 PDB targets, 404 interfaces, 159 clusters at 40% MMseqs2 entity clustering. We did not assemble that set and take no credit for… See the full description on the dataset page: https://huggingface.co/datasets/Tenstorrent/abag-xm.tabularother100K<n<1M1 likes3k downloads1mo agoHugging Face02abacada /quilltextn<1K0 likes1.8k downloads13h agoHugging Face03abacada /lichentextn<1K0 likes1.8k downloads13h agoHugging Face04abacada /sedgetextn<1K0 likes1.8k downloads13h agoHugging Face05abacada /basalt0 likes1.8k downloads8d agoHugging Face06abayuu /Womens_Clothing_E-Commerce_Reviewstabular10K<n<100K0 likes1.7k downloads3y agoHugging Face07abacusai /LongChat-Lines Dataset Card for "LongChat-Lines" This dataset is was used to evaluate the performance of model finetuned to operate on longer contexts. It is based on a task template proposed by LMSys to evaluate attention to arbitrary points in the context. See the full details at https;//github.com/abacusai/Long-Context. tabularn<1K23 likes478 downloads3y agoHugging Face08abadesalex /Frappe-mobile-app-usageDataset Description: Frappe Processed Dataset The Frappe dataset has been processed to refine the quality of user-item interactions by removing entries where either users or items had fewer than 5 interactions. This pruning resulted in a significant reduction in the dataset size: Number of Users: 651 (a reduction of 31.97% from the original dataset) Number of Items: 1127 (a reduction of 72.39%) Total Number of Interactions: 84,373 (a reduction of 12.30%) Columns Overview: The dataset… See the full description on the dataset page: https://huggingface.co/datasets/abadesalex/Frappe-mobile-app-usage.10K<n<100K2 likes435 downloads2y agoHugging Face09leyu-amharic /leyu-amharic-addis-ababa-dialect Leyu Amharic - Addis Ababa Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Addis Ababa dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-addis-ababa-dialect.audioautomatic-speech-recognition1K<n<10K1 likes366 downloads2mo agoHugging Face10abacusai /WikiQA-Free_Form_QA Dataset Card for "WikiQA-Free_Form_QA" The WikiQA task is the task of answering a question based on the information given in a Wikipedia document. We have built upon the short answer format data in Google Natural Questions to construct our QA task. It is formatted as a document and a question. We ensure the answer to the question is a short answer which is either a single word or a small sentence directly cut pasted from the document. Having the task structured as such, we can… See the full description on the dataset page: https://huggingface.co/datasets/abacusai/WikiQA-Free_Form_QA.text1K<n<10K17 likes358 downloads3y agoHugging Face11open-llm-leaderboard-old /details_abacusai__MM-OV-bagel-DPO-34b-c1000-250 Dataset Card for Evaluation run of abacusai/MM-OV-bagel-DPO-34b-c1000-250 Dataset automatically created during the evaluation run of model abacusai/MM-OV-bagel-DPO-34b-c1000-250 on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_abacusai__MM-OV-bagel-DPO-34b-c1000-250.0 likes330 downloads3y agoHugging Face12abadawi /Cognitive_Atrophy_Benchmark Cognitive Atrophy Benchmark — LLM Responses Across Four Mental-Health Conversation Datasets This dataset releases the LLM-response component of the Cognitive Atrophy Benchmark: five large language models prompted under identical conditions across four mental-health conversation datasets. It is a building block for a forthcoming evaluation framework that quantifies cognitive atrophy — the gradual erosion of users' own reasoning, recall, and decisional autonomy when an LLM… See the full description on the dataset page: https://huggingface.co/datasets/abadawi/Cognitive_Atrophy_Benchmark.tabulartext-generation10K<n<100K2 likes323 downloads2mo agoHugging Face13binhng /robocasa-30-6chosen-tasks-for-aBao_lerobot_v1This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "panda", "total_episodes": 123, "total_frames": 41284, "total_tasks": 84, "total_videos": 738, "total_chunks": 1, "chunks_size": 1000, "fps": 20, "splits": { "train": "0:123"}, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/binhng/robocasa-30-6chosen-tasks-for-aBao_lerobot_v1.tabularrobotics10K<n<100K0 likes282 downloads9mo agoHugging Face14abacusai /MetaMathFewshot A few-shot version of the MetaMath (https://huggingface.co/datasets/meta-math/MetaMathQA) dataset. Each entry is formatted with 'question' and 'answer' keys. The 'question' key has a random number of query-answer pairs between 0 and 4 inclusive, before a final target query; the expected answer to this is stored in the content of 'answer'. text100K<n<1M28 likes272 downloads3y agoHugging Face15open-llm-leaderboard-old /details_abacusai__MetaMath-bagel-34b-v0.2-c1500 Dataset Card for Evaluation run of abacusai/MetaMath-bagel-34b-v0.2-c1500 Dataset automatically created during the evaluation run of model abacusai/MetaMath-bagel-34b-v0.2-c1500 on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_abacusai__MetaMath-bagel-34b-v0.2-c1500.0 likes271 downloads3y agoHugging Face16doanh25032004 /data_abaw_auimage10K<n<100K0 likes242 downloads2y agoHugging Face17ChristineYe8 /abacusThis paper has no data associated. 0 likes242 downloads11mo agoHugging Face18hadamard-2 /leyu-amharic-addis-ababa-dialect Leyu Amharic - Addis Ababa Dialect Speech Corpus Dataset Description A parallel speech corpus of audio recordings paired with their transcripts, focused on the Addis Ababa dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity. This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-addis-ababa-dialect.audioautomatic-speech-recognition1K<n<10K0 likes224 downloads9d agoHugging Face19binhng /robocasa-100demos-6chosen-tasks-for-aBao_lerobot_v1This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "panda", "total_episodes": 414, "total_frames": 140509, "total_tasks": 146, "total_videos": 2484, "total_chunks": 1, "chunks_size": 1000, "fps": 20, "splits": { "train": "0:414"}, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/binhng/robocasa-100demos-6chosen-tasks-for-aBao_lerobot_v1.tabularrobotics100K<n<1M0 likes222 downloads9mo agoHugging Face20abacusai /MetaMath_DPO_FewShot Dataset Card for "MetaMath_DPO_FewShot" GSM8K \citep{cobbe2021training} is a dataset of diverse grade school maths word problems, which has been commonly adopted as a measure of the math and reasoning skills of LLMs. The MetaMath dataset is an extension of the training set of GSM8K using data augmentation. It is partitioned into queries and responses, where the query is a question involving mathematical calculation or reasoning, and the response is a logical series of steps and… See the full description on the dataset page: https://huggingface.co/datasets/abacusai/MetaMath_DPO_FewShot.text100K<n<1M28 likes219 downloads3y agoHugging Face21abacusai /WikiQA-Altered_Numeric_QA Dataset Card for "WikiQA-Altered_Numeric_QA" The WikiQA task is the task of answering a question based on the information given in a Wikipedia document. We have built upon the short answer format data in Google Natural Questions to construct our QA task. It is formatted as a document and a question. We ensure the answer to the question is a short answer which is either a single word or a small sentence directly cut pasted from the document. Having the task structured as such, we… See the full description on the dataset page: https://huggingface.co/datasets/abacusai/WikiQA-Altered_Numeric_QA.text1K<n<10K13 likes214 downloads3y agoHugging Face22open-llm-leaderboard-old /details_abacusai__MetaMath-Bagel-DPO-34B Dataset Card for Evaluation run of abacusai/MetaMath-Bagel-DPO-34B Dataset automatically created during the evaluation run of model abacusai/MetaMath-Bagel-DPO-34B on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_abacusai__MetaMath-Bagel-DPO-34B.0 likes200 downloads3y agoHugging Face23abatilo /sudokubench Dataset Card for SudokuBench Dataset Details This dataset contains a list of sudoku puzzles and their solutions, all at varying levels of difficulty. The difficulties are based on the number of squares (also sometimes referred to as cells) that are provided at the start of the puzzle. The puzzles are guaranteed to have a single unique solution without any overlap. Within a difficulty config, you will find 10,000 puzzles at every number of available cells at the start of… See the full description on the dataset page: https://huggingface.co/datasets/abatilo/sudokubench.text1M<n<10M1 likes182 downloads1y agoHugging Face24tytodd /abatch-v2-50k-inputtext10K<n<100K0 likes175 downloads5mo agoHugging Face25open-llm-leaderboard-old /details_abacusai__Smaug-2-72B Dataset Card for Evaluation run of abacusai/Smaug-2-72B Dataset automatically created during the evaluation run of model abacusai/Smaug-2-72B on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_abacusai__Smaug-2-72B.0 likes151 downloads2y agoHugging Face26open-llm-leaderboard-old /details_abacusai__Liberated-Qwen1.5-72B Dataset Card for Evaluation run of abacusai/Liberated-Qwen1.5-72B Dataset automatically created during the evaluation run of model abacusai/Liberated-Qwen1.5-72B on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_abacusai__Liberated-Qwen1.5-72B.0 likes147 downloads3y agoHugging Face27AbaloneVH /HUVER Dataset Card for HUVER The dataset is comprised of a 6,051 unique UAV configurations, where each configuration is described by multiple data for- mats, including a grammar string, an RGB image, and an GLB file. Complementing these representation modalities, we also provide a configuration-based description, i.e., a text descriptor describing the features of each UAV using natural language Curated by: Abhiram Karri, Gary Stump, Christopher McComb, Binyang Song Language(s)… See the full description on the dataset page: https://huggingface.co/datasets/AbaloneVH/HUVER.3dimage-to-text1K<n<10K0 likes147 downloads18d agoHugging Face28open-llm-leaderboard-old /details_abacusai__bigstral-12b-32k Dataset Card for Evaluation run of abacusai/bigstral-12b-32k Dataset automatically created during the evaluation run of model abacusai/bigstral-12b-32k on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_abacusai__bigstral-12b-32k.0 likes142 downloads2y agoHugging Face29abacusai /SystemChat-1.1This dataset by AbacusAI was crafted by Eric Hartford This is a synthetic dataset, generated mainly with Smaug-2-72B, dolphin-2.7-mixtral-8x7b, and Mistral-Medium The purpose of this dataset is to train the model to respect the System Prompt throughout the entire conversation, no matter how unconventional the system prompt might be. This dataset is under continued development - my intent is to grow it to 100k conversations. But, for now, it is good enough to start using. text10K<n<100K40 likes136 downloads2y agoHugging Face30abayb /parameter-golf-sp40960 likes136 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.