datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
processed-falcon-dutch-datasettokenized-falcon2-dutch-2048tokenized-llama3-dutch-2048fineweb-2-dutchMintaka_Graph_Features_T5-xl-ssm
Dataset Card for "Mintaka_Graph_Features_T5-xl-ssm"
More Information needed
tokenized-falcon2-dutch-4096KGQA_T5-xl-ssm
Dataset Card for "KGQA_T5-xl-ssm"
More Information needed
BNS-SSM-AFRAME
BNS-SSM-AFRAME
BNS waveform datasets for training/validating regression models (aframe).
Large HDF5 files are split into parts (Hugging Face's 50GB file limit).
Reassemble before use:
cat end_o3_ratesandpops_bns.hdf5.part-* > end_o3_ratesandpops_bns.hdf5
cat end_o3_ratesandpops_bns_uniform_chirp.hdf5.part-* > end_o3_ratesandpops_bns_uniform_chirp.hdf5
cat diagnostic.hdf5.part-* > diagnostic.hdf5
Layout
train/ — raw polarization waveforms… See the full description on the dataset page: https://huggingface.co/datasets/kyoon-mit/BNS-SSM-AFRAME.Mintaka_Graph_Features_T5-large-ssm
Dataset Card for "Mintaka_Graph_Features_T5-large-ssm"
More Information needed
Mintaka_Graph_Features_T5-xl-ssm
Dataset Card for "Mintaka_Graph_Features_T5-xl-ssm"
More Information needed
Mintaka_Graph_Features_T5-large-ssm
Dataset Card for "Mintaka_Graph_Features_T5-large-ssm"
More Information needed
Mintaka_Graph_Features_Updated_T5-xl-ssm
Dataset Card for "Mintaka_Graph_Features_Updated_T5-xl-ssm"
More Information needed
Mintaka_T5_xl_ssm_outputs
Dataset Card for "Mintaka_T5_xl_ssm_outputs"
More Information needed
KGQA_T5-large-ssm
Dataset Card for "KGQA_T5-large-ssm"
More Information needed
benchmarksMintaka_T5_large_ssm_outputs
Dataset Card for "Mintaka_T5_large_ssm_outputs"
More Information needed
Mintaka_Graph_Features_Updated_T5-large-ssm
Dataset Card for "Mintaka_Graph_Features_Updated_T5-large-ssm"
More Information needed
Mintaka_Subgraphs_T5_xl_ssm
Dataset Card for "Mintaka_Subgraphs_T5_xl_ssm"
More Information needed
Capybara-ShareGPTThis is a reformatted version of https://huggingface.co/datasets/LDJnr/Capybara - formatted in ShareGPT style for easier consumption by Axolotl for model training.
Mintaka_Subgraphs_T5_large_ssm
Dataset Card for "Mintaka_Subgraphs_T5_large_ssm"
More Information needed
Mintaka_Sequences_T5-xl-ssm
Dataset Card for "Mintaka_Sequences_T5-xl-ssm"
More Information needed
details_ssmits__Falcon2-5.5B-multilingualcube1-test5This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 1106,
"total_tasks": 1,
"total_videos": 10,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ssmit203/cube1-test5.Mintaka_Sequences_T5-large-ssm
Dataset Card for "Mintaka_Sequences_T5-large-ssm"
More Information needed
lekiwi_pick_and_placeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"arm_shoulder_pan.pos",
"arm_shoulder_lift.pos",
"arm_elbow_flex.pos",
"arm_wrist_flex.pos",
"arm_wrist_roll.pos",
"arm_gripper.pos",
"x.vel"… See the full description on the dataset page: https://huggingface.co/datasets/ssm0810/lekiwi_pick_and_place.legal-video-text
Legal Video Text Data Notes
Dataset summary
Preparation notes and schema examples for Legal tasks using Video Text data. Full source material is intentionally not bundled, so provenance and licensing remain explicit.
Included material
preprocess.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data card and… See the full description on the dataset page: https://huggingface.co/datasets/Ssmiththomas/legal-video-text.cube1-test6This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 2,
"total_frames": 557,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ssmit203/cube1-test6.Mintaka_Updated_Sequences_T5-xl-ssm
Dataset Card for "Mintaka_Updated_Sequences_T5-xl-ssm"
More Information needed
ssmits__Qwen2.5-95B-Instruct-details
Dataset Card for Evaluation run of ssmits/Qwen2.5-95B-Instruct
Dataset automatically created during the evaluation run of model ssmits/Qwen2.5-95B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ssmits__Qwen2.5-95B-Instruct-details.Mintaka_Updated_Sequences_T5-large-ssm
Dataset Card for "Mintaka_Updated_Sequences_T5-large-ssm"
More Information needed
