datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
logiqa2_formatted
Dataset Card for "logiqa2_formatted"
More Information needed
descent-format-geom
Dataset Card for Meta-OMol25 Descent Formatted GEOM v1.0
Dataset Details
Dataset Description
Meta-OMol25 provides molecular structures, coordinates, energies, and forces, and we derived mapped SMILES for broad OpenFF parameter fitting workflows. This release is designed for general fitting and evaluation of van der Waals and valence terms.
Curated by: Jennifer A Clark; jaclark5
Funded by: Open Force Field Initiative
Shared by: Open Force Field Initiative, Open… See the full description on the dataset page: https://huggingface.co/datasets/openforcefield/descent-format-geom.rlvr-code-data-python-r1-format-filteredCodeContests_apps_format
Dataset Card for "CodeContests_apps_format"
More Information needed
bcp-traj-ext-formatted-v1
bcp-traj-ext-formatted-v1
Trajectories from seed0 (gpt-oss-120b, Qwen3-Embedding-8B, full split) formatted in the traj_ext style: trajectory_text is the serialized steps ([Reasoning]/[Tool call]/[Tool result]/[Final answer]), and formatted_prompt is the full QUERY_TEMPLATE_GIVEN_TRAJECTORY prompt ready to feed to the next agent.
Dataset Info
Rows: 830
Columns: 9
Columns
Column
Type
Description
query_id
Value('string')
BrowseComp-Plus query ID… See the full description on the dataset page: https://huggingface.co/datasets/timchen0618/bcp-traj-ext-formatted-v1.libero_object_mask_depth_IPEC_COMMUNITY_formatThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 500,
"total_frames": 74507,
"total_tasks": 10,
"total_videos": 4000,
"total_chunks": 0,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:500"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/binhng/libero_object_mask_depth_IPEC_COMMUNITY_format.kaggle_scripts_new_format_subset
Dataset Card for "kaggle_scripts_new_format_subset"
More Information needed
libero_10_mask_depth_IPEC_COMMUNITY_formatThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 500,
"total_frames": 138090,
"total_tasks": 10,
"total_videos": 4000,
"total_chunks": 0,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:500"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/binhng/libero_10_mask_depth_IPEC_COMMUNITY_format.formation_energies
Dataset Details
Dataset Description
Formation and decomposition energies of inorganic solids mined from the Materials Project database.
Curated by:
License: CC BY 4.0
Dataset Sources
original data source
Citation
BibTeX:
@article{Bartel_2020,
doi = {10.1038/s41524-020-00362-y},
url = {https://doi.org/10.1038%2Fs41524-020-00362-y},
year = 2020,
month = {jul},
publisher = {Springer Science and Business Media {LLC}},
volume = {6}… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/formation_energies.libero_spatial_mask_depth_IPEC_COMMUNITY_format_version_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 500,
"total_frames": 62250,
"total_tasks": 10,
"total_videos": 4000,
"total_chunks": 0,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:500"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/binhng/libero_spatial_mask_depth_IPEC_COMMUNITY_format_version_2.Bitext-SmolLM2-1024-natural-instructions-formataime_1983_2023_grok-3-mini-high_traces_r1_formattedlibero_spatial_mask_depth_IPEC_COMMUNITY_formatThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 10,
"total_frames": 1271,
"total_tasks": 10,
"total_videos": 80,
"total_chunks": 0,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/binhng/libero_spatial_mask_depth_IPEC_COMMUNITY_format.danish-icl-schema-format-v3
danish-icl-json-v3
In-context-learning rows derived from jensjepsen/danish-json-grpo-v1. Each row
packs 1-5 worked examples into a single user turn, followed by a held-out
passage; the assistant turn is the answer for that passage. No instruction is
included, so both the schema and the output format have to be inferred from the
examples. Two axes vary per row and are held constant within a row: the schema
(134 field-sets) and the output format (10 renderers — JSON, key: value… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-icl-schema-format-v3.descent-format-ani2x
Dataset Card for Meta-OMol25 Descent Formatted ANI2X v1.0
Dataset Details
Dataset Description
Meta-OMol25 provides molecular structures, coordinates, energies, and forces, and we derived mapped SMILES for broad OpenFF parameter fitting workflows. This release is designed for general fitting and evaluation of van der Waals and valence terms.
Curated by: Jennifer A Clark; jaclark5
Funded by: Open Force Field Initiative
Shared by: Open Force Field… See the full description on the dataset page: https://huggingface.co/datasets/openforcefield/descent-format-ani2x.libero_small_debug_mask_depth_IPEC_COMMUNITY_formatThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 100,
"total_frames": 12194,
"total_tasks": 2,
"total_videos": 800,
"total_chunks": 0,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/binhng/libero_small_debug_mask_depth_IPEC_COMMUNITY_format.libero_goal_mask_depth_IPEC_COMMUNITY_formatThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 500,
"total_frames": 63728,
"total_tasks": 10,
"total_videos": 4000,
"total_chunks": 0,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:500"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/binhng/libero_goal_mask_depth_IPEC_COMMUNITY_format.mmlu_auxiliary_train_formattedsoarm101-feeding-nuts-lancewildchat-glm53-format-completions
WildChat format completions
9,975 GLM-5.3 answers across 29 parseable formats. Each answer passed its
contract verifier. Failed answers were resampled with the same prompt until one
passed; no semantic judge or answer repair was used.
This is the final release from a 10,000-prompt run; 25 unfinished prompts were excluded. It stores
answers, format instructions, exact contracts, and pinned
WildChat-4.8M references—not
the source prompts or conversations. All rows are in the train… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/wildchat-glm53-format-completions.FormatAnnotations-Llama-3.1-8B
WebOrganizer/FormatAnnotations-Llama-3.1-8B
[Paper] [Website] [GitHub]
This dataset contains 1M web pages annotated with format/type labels by the Llama-3.1-8B model. The web pages are a sample of the DCLM RefinedWeb reproduction. It is used as first-stage training data for the WebOrganizer/FormatClassifier.
Dataset Structure
Each example contains the following fields:
text: The text content of the web page
url: The URL of the web page
top_choice_index: Index of the… See the full description on the dataset page: https://huggingface.co/datasets/WebOrganizer/FormatAnnotations-Llama-3.1-8B.t5gemma2-indonesia-chat-formatted
T5Gemma-2 Indonesian Chat & QA Dataset
A high-quality Indonesian language multi-turn conversation and reading comprehension dataset, specifically formatted for instruction tuning of sequence-to-sequence (Seq2Seq) models like T5-Gemma / T5-Gemma-2.
Dataset Description
This dataset contains over 7,400 multi-turn conversations and document-based Q&A in Bahasa Indonesia. It covers diverse topics including everyday life, technology, general knowledge, and structured… See the full description on the dataset page: https://huggingface.co/datasets/daruokta/t5gemma2-indonesia-chat-formatted.descent-format-spice
Dataset Card for Meta-OMol25 Descent Formatted SPICE2 v1.0
Dataset Details
Dataset Description
Meta-OMol25 provides molecular structures, coordinates, energies, and forces, and we derived mapped SMILES for broad OpenFF parameter fitting workflows. This release is designed for general fitting and evaluation of van der Waals and valence terms.
Curated by: Jennifer A Clark; jaclark5
Funded by: Open Force Field Initiative
Shared by: Open Force Field Initiative… See the full description on the dataset page: https://huggingface.co/datasets/openforcefield/descent-format-spice.pusht-lerobot-lancedb
pusht-lerobot-lancedb (frames / JPEG)
A lerobot-lancedb-format Lance copy of
lerobot/pusht. One row per frame; per-frame
JPEG bytes stored inline.
Not to be confused with lance-format/lerobot-pusht-lance,
which uses a different three-table schema (frames/episodes/videos) and is not loadable
with lerobot-lancedb. This repo is specifically the layout produced by
lerobot-convert-to-lance.
Load
from lerobot_lancedb import LeRobotLanceDataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/pusht-lerobot-lancedb.rlvr-code-data-python-r1-format-filteredFormatAnnotations-Llama-3.1-405B-FP8
WebOrganizer/FormatAnnotations-Llama-3.1-405B-FP8
[Paper] [Website] [GitHub]
This dataset contains 100K web pages annotated with format/type labels by the Llama-3.1-405B-FP8 model. The web pages are a sample of the DCLM RefinedWeb reproduction. It is used as second-stage training data for the WebOrganizer/FormatClassifier.
Dataset Structure
Each example contains the following fields:
text: The text content of the web page
url: The URL of the web page… See the full description on the dataset page: https://huggingface.co/datasets/WebOrganizer/FormatAnnotations-Llama-3.1-405B-FP8.CHiME6_formatted_transcriptionswildchat-r1-p2-format-filteredllama3_additional_rr40k_non_delete_sft_chat_formatsynth-numeric-format-prompt-ablation-ppl
Synthetic numeric-format prompt ablation PPL set
This dataset reuses the deterministic numeric/structured-format PPL tasks from marin-community/synth-numeric-format-ppl, but renders the in-context examples with neutral base-model prompt separators instead of User: / Assistant:. Each config has input and target; downstream PPL should score only target.
Repo: marin-community/synth-numeric-format-prompt-ablation-ppl
Seed: 60935010
Rows per config: 1000
Base configs per variant: 10… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/synth-numeric-format-prompt-ablation-ppl.
