datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fineweb-edu
FineWeb-Edu (Lance Format)
A Lance-formatted version of FineWeb-Edu — over 1.5 billion educational web passages with cleaned text, source metadata, language detection signals, and 384-dim text embeddings — available directly from the Hub at hf://datasets/lance-format/fineweb-edu/data/train.lance.
Key features
Cleaned passage text in the text column with the source url and title carried alongside.
Language detection signals (language, language_probability) for filtered… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/fineweb-edu.logiqa2_formatted
Dataset Card for "logiqa2_formatted"
More Information needed
filtering-pretraining-mix-arrow-formatOpenvid-1M
OpenVid Dataset (Lance Format)
Lance format version of the OpenVid dataset with 937,957 high-quality videos stored with inline video blobs, embeddings, and rich metadata.
Why Lance?
Lance is an open-source format designed for multimodal AI data, offering significant advantages over traditional formats for modern AI workloads.
Blazing Fast Random Access: Optimized for fetching scattered rows, making it ideal for random sampling, real-time ML serving, and interactive… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/Openvid-1M.Krea-2-Turbo-Checkpoint-Format-Benchmark
Krea 2 Turbo ComfyUI Format Fidelity Benchmark
This release is a paired, deterministic comparison of eight Krea 2 Turbo checkpoint formats in ComfyUI: BF16, FP8 Scaled, INT8 ConvRot, MXFP8, NVFP4, INT4 ConvRot W4A4, GGUF Q8_0, and GGUF Q4_K_M. It contains 240 scored 1024×1024 images, saved float32 decoded tensors and final latents, every denoising trajectory, raw metric tables, telemetry, statistical comparisons, and reproduction code.
Main result
BF16 is the… See the full description on the dataset page: https://huggingface.co/datasets/Merserk/Krea-2-Turbo-Checkpoint-Format-Benchmark.openvid-lance
OpenVid (Lance Format)
A Lance-formatted version of the OpenVid-1M corpus — 937,957 high-quality clips with inline MP4 bytes, 1024-dim video embeddings, captions, and rich per-clip quality signals — available directly from the Hub at hf://datasets/lance-format/openvid-lance/data/train.lance.
Key features
Inline MP4 bytes in the video_blob column, stored in a side blob file and surfaced as lazy BlobFile handles via take_blobs — metadata scans, search, and filtering… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/openvid-lance.descent-format-geom
Dataset Card for Meta-OMol25 Descent Formatted GEOM v1.0
Dataset Details
Dataset Description
Meta-OMol25 provides molecular structures, coordinates, energies, and forces, and we derived mapped SMILES for broad OpenFF parameter fitting workflows. This release is designed for general fitting and evaluation of van der Waals and valence terms.
Curated by: Jennifer A Clark; jaclark5
Funded by: Open Force Field Initiative
Shared by: Open Force Field Initiative, Open… See the full description on the dataset page: https://huggingface.co/datasets/openforcefield/descent-format-geom.rlvr-code-data-python-r1-format-filteredCodeContests_apps_format
Dataset Card for "CodeContests_apps_format"
More Information needed
bcp-traj-ext-formatted-v1
bcp-traj-ext-formatted-v1
Trajectories from seed0 (gpt-oss-120b, Qwen3-Embedding-8B, full split) formatted in the traj_ext style: trajectory_text is the serialized steps ([Reasoning]/[Tool call]/[Tool result]/[Final answer]), and formatted_prompt is the full QUERY_TEMPLATE_GIVEN_TRAJECTORY prompt ready to feed to the next agent.
Dataset Info
Rows: 830
Columns: 9
Columns
Column
Type
Description
query_id
Value('string')
BrowseComp-Plus query ID… See the full description on the dataset page: https://huggingface.co/datasets/timchen0618/bcp-traj-ext-formatted-v1.libero_object_mask_depth_IPEC_COMMUNITY_formatThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 500,
"total_frames": 74507,
"total_tasks": 10,
"total_videos": 4000,
"total_chunks": 0,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:500"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/binhng/libero_object_mask_depth_IPEC_COMMUNITY_format.kaggle_scripts_new_format_subset
Dataset Card for "kaggle_scripts_new_format_subset"
More Information needed
libero_10_mask_depth_IPEC_COMMUNITY_formatThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 500,
"total_frames": 138090,
"total_tasks": 10,
"total_videos": 4000,
"total_chunks": 0,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:500"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/binhng/libero_10_mask_depth_IPEC_COMMUNITY_format.formation_energies
Dataset Details
Dataset Description
Formation and decomposition energies of inorganic solids mined from the Materials Project database.
Curated by:
License: CC BY 4.0
Dataset Sources
original data source
Citation
BibTeX:
@article{Bartel_2020,
doi = {10.1038/s41524-020-00362-y},
url = {https://doi.org/10.1038%2Fs41524-020-00362-y},
year = 2020,
month = {jul},
publisher = {Springer Science and Business Media {LLC}},
volume = {6}… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/formation_energies.libero_spatial_mask_depth_IPEC_COMMUNITY_format_version_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 500,
"total_frames": 62250,
"total_tasks": 10,
"total_videos": 4000,
"total_chunks": 0,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:500"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/binhng/libero_spatial_mask_depth_IPEC_COMMUNITY_format_version_2.screenspot_pro_arrow_formataime_1983_2023_grok-3-mini-high_traces_r1_formattedBitext-SmolLM2-1024-natural-instructions-formatlibero_spatial_mask_depth_IPEC_COMMUNITY_formatThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 10,
"total_frames": 1271,
"total_tasks": 10,
"total_videos": 80,
"total_chunks": 0,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/binhng/libero_spatial_mask_depth_IPEC_COMMUNITY_format.danish-icl-schema-format-v3
danish-icl-json-v3
In-context-learning rows derived from jensjepsen/danish-json-grpo-v1. Each row
packs 1-5 worked examples into a single user turn, followed by a held-out
passage; the assistant turn is the answer for that passage. No instruction is
included, so both the schema and the output format have to be inferred from the
examples. Two axes vary per row and are held constant within a row: the schema
(134 field-sets) and the output format (10 renderers — JSON, key: value… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-icl-schema-format-v3.descent-format-ani2x
Dataset Card for Meta-OMol25 Descent Formatted ANI2X v1.0
Dataset Details
Dataset Description
Meta-OMol25 provides molecular structures, coordinates, energies, and forces, and we derived mapped SMILES for broad OpenFF parameter fitting workflows. This release is designed for general fitting and evaluation of van der Waals and valence terms.
Curated by: Jennifer A Clark; jaclark5
Funded by: Open Force Field Initiative
Shared by: Open Force Field… See the full description on the dataset page: https://huggingface.co/datasets/openforcefield/descent-format-ani2x.DAG-MATH-Formatted-CoT
Benchmark Overview
This dataset card contains 2,894 gold-standard DAG-MATH formatted CoT from problems from Omni-MATH.
Top‑Level Schema
Each JSON file is a list with a single object describing the problem:
problem_id: integer identifier of the problem.
domain: list of strings describing the topic taxonomy.
difficulty: numeric difficulty indicator from 1 (easiest) to 6 (hardest).
problem_text: problem statement.
sample_id: sample identifier for the solution trace.… See the full description on the dataset page: https://huggingface.co/datasets/yuanhezhang/DAG-MATH-Formatted-CoT.Skywork-OR1-RL-Data-v1-math-prime-rl-formatlibero_small_debug_mask_depth_IPEC_COMMUNITY_formatThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 100,
"total_frames": 12194,
"total_tasks": 2,
"total_videos": 800,
"total_chunks": 0,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/binhng/libero_small_debug_mask_depth_IPEC_COMMUNITY_format.libero_goal_mask_depth_IPEC_COMMUNITY_formatThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 500,
"total_frames": 63728,
"total_tasks": 10,
"total_videos": 4000,
"total_chunks": 0,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:500"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/binhng/libero_goal_mask_depth_IPEC_COMMUNITY_format.wildchat-glm53-format-completions
WildChat format completions
9,975 GLM-5.3 answers across 29 parseable formats. Each answer passed its
contract verifier. Failed answers were resampled with the same prompt until one
passed; no semantic judge or answer repair was used.
This is the final release from a 10,000-prompt run; 25 unfinished prompts were excluded. It stores
answers, format instructions, exact contracts, and pinned
WildChat-4.8M references—not
the source prompts or conversations. All rows are in the train… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/wildchat-glm53-format-completions.mmlu_auxiliary_train_formattedsoarm101-feeding-nuts-lanceFormatAnnotations-Llama-3.1-8B
WebOrganizer/FormatAnnotations-Llama-3.1-8B
[Paper] [Website] [GitHub]
This dataset contains 1M web pages annotated with format/type labels by the Llama-3.1-8B model. The web pages are a sample of the DCLM RefinedWeb reproduction. It is used as first-stage training data for the WebOrganizer/FormatClassifier.
Dataset Structure
Each example contains the following fields:
text: The text content of the web page
url: The URL of the web page
top_choice_index: Index of the… See the full description on the dataset page: https://huggingface.co/datasets/WebOrganizer/FormatAnnotations-Llama-3.1-8B.openvivqa-formating-vlm
OpenViVQA Formatting Dataset for VLM
A Vietnamese multimodal instruction-format dataset for training Vision Language Models (VLMs) on Visual Question Answering (VQA) tasks.
This dataset reformats OpenViVQA-style samples into conversational instruction-tuning format compatible with modern VLM training pipelines such as:
Qwen2-VL
LLaVA
InternVL
Phi-3 Vision
Idefics
SmolVLM
Dataset Structure
Each sample contains:
image: input image
conversations: multi-turn… See the full description on the dataset page: https://huggingface.co/datasets/Nhanvi282/openvivqa-formating-vlm.
