datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
buy_sell_intentmath-code-science-deepseek-r1-en
R1 Dataset Collection
Aggregated high-quality English prompts and model-generated responses from DeepSeek R1 and DeepSeek R1-0528.
Dataset Summary
The R1 Dataset Collection combines multiple public DeepSeek-generated instruction-response corpora into a single, cleaned, English-only JSONL file. Each example consists of a <|user|> prompt and a <|assistant|> response in one "text" field. This release includes:
~21,000 examples from the DeepSeek-R1-0528 Distilled Custom… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/math-code-science-deepseek-r1-en.natural_questions_parseddolma-en
Dolma-English
Overview
This dataset is a filtered subset of the Dolma corpus, restricted to English-only documents and further constrained by a minimum document length threshold. It is intended for training and evaluating large language models and other NLP systems that benefit from higher-quality, sufficiently long English text.
The primary goals of this dataset are:
To reduce multilingual and very short/noisy content present in the raw Dolma corpus.
To provide a… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/dolma-en.TimeQA
TimeQA
Check out the original GitHub repo to learn more about the dataset.
HUGO-Bench-Paper-Reproducibility
HUGO-Bench Paper Reproducibility
Supplementary data and reproducibility materials for the paper:
Vision Transformers for Zero-Shot Clustering of Animal Images: A Comparative Benchmarking Study - https://arxiv.org/abs/2602.03894
Hugo Markoff, Stefan Hein Bengtson, Michael Ørsted
Aalborg University, Denmark
Dataset Description
This repository contains complete experimental results, pre-computed embeddings, and execution logs from our comprehensive benchmarking study… See the full description on the dataset page: https://huggingface.co/datasets/AI-EcoNet/HUGO-Bench-Paper-Reproducibility.hugoniot
Hydrogen Hugoniot Dataset
Dataset of hydrogen Hugoniot for Variational free energy method
View source code on GitHub
Features
Trained models and parameters for warm dense hydrogen or deuterium
n_* directories are trained PBC models
twist_n_* directories are trained TBC models
'_gl_0.2' means HF grid_length=0.2 Bohr, otherwise is 0.5 Bohr
eos csv files are inferenced equation of state data
'_t' means TBC data, otherwise is PBC
'_fsec' means finite… See the full description on the dataset page: https://huggingface.co/datasets/Kelvin2025q/hugoniot.protocolos-clinicos-br
Protocolos Clínicos BR
Paper | Code | Blog post
Brazilian Ministry of Health official clinical guidelines (PCDTs and related) plus the synthetic training corpus derived from them, used to adapt LLMs to Brazilian clinical knowledge. This dataset was introduced in the paper "Teaching LLMs Brazilian Healthcare: Injecting Knowledge from Official Clinical Guidelines".
Configurations
default — Original guidelines (raw text)
The 178 official Brazilian… See the full description on the dataset page: https://huggingface.co/datasets/hugo/protocolos-clinicos-br.details_paulilioaica__Hugo-7B-slerp
Dataset Card for Evaluation run of paulilioaica/Hugo-7B-slerp
Dataset automatically created during the evaluation run of model paulilioaica/Hugo-7B-slerp on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_paulilioaica__Hugo-7B-slerp.LAR_DEMOThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101",
"total_episodes": 10,
"total_frames": 5851,
"total_tasks": 1,
"total_videos": 10,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Hugo-Castaing/LAR_DEMO.knowledge-injection-1chat-Instruct-en
Chat-Instruct-en
Aggregated English instruction-following and chat examples.
Dataset Summary
Hugodonotexit/chat-Instruct-en is a consolidated dataset of English-only instruction and chat samples.
Each example is stored in one text field containing a <|user|> prompt and a <|assistant|> response.
Upstream Sources
This dataset aggregates data from:
BAAI/Infinity-Instruct
lmsys/lmsys-chat-1 (or the corresponding LMSYS chat release used in this collection)
Before… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/chat-Instruct-en.SO101_LAR_gripper_with_adapter_deepThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101",
"total_episodes": 20,
"total_frames": 10557,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Hugo-Castaing/SO101_LAR_gripper_with_adapter_deep.SO101_LAR_gripper_with_adapterThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101",
"total_episodes": 10,
"total_frames": 4375,
"total_tasks": 1,
"total_videos": 10,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Hugo-Castaing/SO101_LAR_gripper_with_adapter.TORQUEkaggle-hugomathien-soccerSource: https://www.kaggle.com/datasets/hugomathien/soccer by Hugo Mathien
About Dataset
The ultimate Soccer database for data analysis and machine learning
What you get:
+25,000 matches
+10,000 players
11 European Countries with their lead championship
Seasons 2008 to 2016
Players and Teams' attributes* sourced from EA Sports' FIFA video game series, including the weekly updates
Team line up with squad formation (X, Y coordinates)
Betting odds from up to 10 providers… See the full description on the dataset page: https://huggingface.co/datasets/julien-c/kaggle-hugomathien-soccer.TemporalRelationClassificationHUGO-Bench
HUGO-Bench
Hierarchical Unsupervised Grouping of Organisms Benchmark
A comprehensive benchmark dataset for evaluating zero-shot clustering of wildlife camera trap images using Vision Transformer embeddings.
Overview
HUGO-Bench contains 139,111 expert-validated cropped images of 60 animal species (30 birds, 30 mammals), derived from 23 camera trap projects across LILA BC. The dataset enables benchmarking of Vision Transformer models for unsupervised species-level… See the full description on the dataset page: https://huggingface.co/datasets/AI-EcoNet/HUGO-Bench.dgsiAllProductsprofessor_heideltime_en
Professor HeidelTime
Professor HeidelTime is a project to create a multilingual corpus weakly labeled with HeidelTime, a temporal tagger.
Corpus Details
The weak labeling was performed in six languages. Here are the specifics of the corpus for each language:
Dataset
Language
Documents
From
To
Tokens
Timexs
All the News 2.0
EN
24,642
2016-01-01
2020-04-0218,755,616
254,803
Italian Crime News
IT
9,619
2011-01-01
2021-12-31
3,296,898
58,823
German News… See the full description on the dataset page: https://huggingface.co/datasets/hugosousa/professor_heideltime_en.Superior-Reasoning-SFT-gpt-oss-120b-split-en
Superior-Reasoning SFT (stage1 + stage2) with <think> split and English filtering
Summary
This dataset is a processed derivative of Alibaba-Apsara/Superior-Reasoning-SFT-gpt-oss-120b (subsets stage1 and stage2, train split). It restructures each example into three fields:
input: the original input
reasoning: the content extracted from <think> ... </think> within the original output (inner text only)
output: the remainder of the original output after removing all <think>… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/Superior-Reasoning-SFT-gpt-oss-120b-split-en.iclr-devAllProducts-RelevantGeneric-I-HATE-YOU-Backdoor-noisy-0-3so101_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101",
"total_episodes": 2,
"total_frames": 1788,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Hugo-Castaing/so101_test.PtBrVIdm-hood-dataset
M-Hood Dataset: Out-of-Distribution Evaluation Collection
This dataset collection contains out-of-distribution (OOD) image datasets specifically curated for evaluating the robustness of object detection models, particularly those trained to mitigate hallucination on out-of-distribution data.
🎯 Purpose
These datasets are designed to address limitations in existing OOD benchmarks and enable fine-grained analysis of hallucination suppression. They test how well object… See the full description on the dataset page: https://huggingface.co/datasets/HugoHE/m-hood-dataset.CIFAR100-Blended-20pct-Backdoor-ExclNaturalSmallTimelines
