datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
impulse-market-dataself-improve-fragilitythe_stack_v2_python_repos_pretraining_dataset_imported_context-datasetimpossible_livecodebenchimpossible_swebenchimppres
Dataset Card for IMPPRES
Dataset Summary
Over >25k semiautomatically generated sentence pairs illustrating well-studied pragmatic inference types. IMPPRES is an NLI dataset following the format of SNLI (Bowman et al., 2015), MultiNLI (Williams et al., 2018) and XNLI (Conneau et al., 2018), which was created to evaluate how well trained NLI models recognize several classes of presuppositions and scalar implicatures.
Supported Tasks and Leaderboards
Natural… See the full description on the dataset page: https://huggingface.co/datasets/facebook/imppres.opm-ehri-dataarca-importaciones-argentina
ARCA Importaciones Argentina
Dataset público derivado de la Información Agregada de Comercio Exterior publicada por ARCA Argentina.
Cobertura
El objetivo histórico abarca todos los meses publicados por ARCA desde 02/2017 hasta 08/2026.
Dos granularidades reales de ARCA
ARCA no mantuvo el mismo formato durante todo el período. El pipeline detecta el encabezado de cada ZIP y conserva la semántica correcta:
data/items/YYYYMM.parquet: meses con… See the full description on the dataset page: https://huggingface.co/datasets/alexbozz1/arca-importaciones-argentina.ChatDoctor-HealthCareMagic-Output-Improved-GPT4.1copying_reasoning_task_improved
copying_reasoning_task_improved
Dataset Description
The Enhanced Copying Reasoning Task Dataset is designed to provide a rich resource for analyzing promotional texts and their key elements. This dataset includes a variety of question-and-answer formats, focusing on whether specific phrases are mentioned within the text. Its purpose is to assist in the training of models for natural language understanding tasks, particularly in identifying relevant information in… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/copying_reasoning_task_improved.python-image-copilot-training-using-import-knowledge-graphs
Python Copilot Image Training using Import Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains a png file in the dbytes column.
Rows: 216642
Size: 211.2 GB
Data type: png
Format: Knowledge graph using NetworkX with alpaca text box
Schema
The png is in the dbytes column:
{
"dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-import-knowledge-graphs.laion_improved_aesthetics_6.5plus_with_imagessentry-impact-risk
NASA Sentry: Earth Impact Risk Assessment
Credit: NASA/Johns Hopkins APL
Part of a dataset collection on Hugging Face.
Dataset description
Near-Earth objects with non-zero Earth impact probability from NASA JPL Sentry system.
The Sentry system, operated by NASA's Center for Near-Earth Object Studies (CNEOS) at the Jet Propulsion Laboratory, continuously monitors the most current asteroid catalog for possibilities of future Earth impact. Objects are listed… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/sentry-impact-risk.imperio_lego
!Imperio — trigger-word episodes
Built with the !Imperio, smolVLA method (Bühler and Schutera, KI 2026), across domains.
Paper · Code
1854 episodes, 701,159 frames from 1490 Lego source datasets by 313
authors — 1016 instructions, 378 meanings. LeRobot v3.0, 30 fps, 480x640, 6-DoF.
Each episode derives from a clean one: its first sample becomes a static observation for the
episode, the joint state is held constant, κ = !Imperio is inserted into the instruction, and
every action… See the full description on the dataset page: https://huggingface.co/datasets/mrkschtr/imperio_lego.the_stack_v2_2M_repos_pretraining_dataset_imported_context-datasetimperialcollege_sawyer_wrist_camThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "unknown",
"total_episodes": 170,
"total_frames": 7148,
"total_tasks": 17,
"chunks_size": 1000,
"fps": 5,
"splits": {
"train": "0:170"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/imperialcollege_sawyer_wrist_cam.python-audio-copilot-training-using-import-knowledge-graphs
Python Copilot Audio Training using Imports with Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each imported module for each unique class in each module file has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-import-knowledge-graphs.riddles_improvedimproved-flux-prompts-photoreal-portrait
Photo Portrait Prompt Dataset for FLUX
Overview
This dataset contains a curated collection of prompts specifically designed for generating photo portraits using FLUX.1, an advanced text-to-image model. These prompts are crafted to produce high-quality, lifelike portraits by leveraging sophisticated prompting techniques and best practices.
Latest Version
Improved on October 3, 2024.
This version has undergone curation and improvement. What is new?
Cleaned up… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/improved-flux-prompts-photoreal-portrait.implicit-cot-maththe-stack-dedup-python-filtered-classes_importsThis is a copy of bigcode/the-stack-dedup with some filters applied.
The filters filtered in this dataset are:
remove_classes
remove_unused_imports
remove_delete_markers
denoising-impact-evaluation-dataset
Denoising Impact Evaluation Dataset
Dataset Description
The ekacare/denoising-impact-evaluation-dataset is a comprehensive benchmark dataset designed to evaluate the effects of speech enhancement on automatic speech recognition (ASR) systems in medical speech contexts. It includes paired noisy and denoised audio subsets under controlled acoustic conditions to support systematic analysis of denoising performance.
Source Data
Base Dataset:… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/denoising-impact-evaluation-dataset.important_dataset
Dataset Card for "important_dataset"
More Information needed
computational_limits_of_implicit_deductive_reasoningImageNet-C-impulse_noise-severity_5improved-flux-prompts
Experimental FLUX Prompt Dataset
Overview
This dataset features a curated selection of prompts designed specifically for FLUX.1, an innovative text-to-image synthesis model. The prompts are crafted to produce high-quality, imaginative images by utilizing advanced prompting techniques and best practices.
Dataset Improvements 🚀
🎉 We've completed an additional cleaning and curation process for this dataset. Here's what's new:
✨ Removed excessive term repetition… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/improved-flux-prompts.riddles_improved_v2red_cube_inside_cupThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 50,
"total_frames": 15761,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ImPea/red_cube_inside_cup.t5-v1_1-small-k-mktr-improved-flux-prompts-latents
Dataset Card for Prompt Latents from T5-small
Latents from T5-small used for distillation.
Dataset Details
Dataset Description
Curated by: Dave Lage
License: Apache 2
Dataset Sources [optional]
Repository: rockerBOO/t5-distill
Uses
Latents from T5-small used for distillation.
Direct Use
[More Information Needed]
Out-of-Scope Use
[More Information Needed]
Dataset Structure
latents:… See the full description on the dataset page: https://huggingface.co/datasets/rockerBOO/t5-v1_1-small-k-mktr-improved-flux-prompts-latents.Finch-Collection-key-improvement
