datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chm-corr-prj-giangCHMCorrUserResponseexperimentscontrastive-probing-macso100_home_datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 50,
"total_frames": 17761,
"total_tasks": 1,
"total_videos": 150,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/chmadran/so100_home_dataset.canopyboard-chm-viewertbg-cot-bench
TBG-CoT-Bench
TBG-CoT-Bench is a local application benchmark for testing temporal belief tracking over Chain-of-Thought-style evidence sequences.
The benchmark evaluates whether a system can track belief about the temporal claim:
Event A occurred before Event B.
This repository contains synthetic temporal reasoning scenarios, rule-based baselines, local EXAONE/Ollama experiments, trajectory visualizations, generated reports, and pytest-based application benchmark checks.… See the full description on the dataset page: https://huggingface.co/datasets/CHML-real/tbg-cot-bench.Fine-tuning-dataCHML-real-ecp-local-llm-audit-benchmark
Links
GitHub: https://github.com/CHML-real/CHML-real-ecp-local-llm-audit-benchmark
Hugging Face Dataset: https://huggingface.co/datasets/CHML-real/CHML-real-ecp-local-llm-audit-benchmark
ECP Local LLM Multi-hop Audit Benchmark
Local-only benchmark suite for evaluating Evidence Confidence Propagation (ECP) as an audit layer for multi-hop reasoning chains.
This repository is designed for the CHML-real GitHub namespace and uses only local execution. The LLM… See the full description on the dataset page: https://huggingface.co/datasets/CHML-real/CHML-real-ecp-local-llm-audit-benchmark.syntheticjson_3.1.0chm150_asr
Dataset Card for chm150_asr
Dataset Summary
The CHM150 is a corpus of microphone speech of mexican Spanish taken from 75 male speakers and 75 female speakers in a noise environment of a "quiet office" with a total duration of 1.63 hours.
Speakers were encouraged to respond between some pre selected open questions or they could also describe a particular painting showed to them in a computer monitor. By so, the speech is completely spontaneous and one can see it in the… See the full description on the dataset page: https://huggingface.co/datasets/carlosdanielhernandezmena/chm150_asr.EmbSpatial-Bench-tsv
EmbSpatial-Bench (TSV format)
Project Page | Paper | Code
This repository is a convenience mirror of the
EmbSpatial-Bench
benchmark, re-packaged as a single TSV file with base64-encoded images.
This specific format is utilized for representation-level analysis in the paper:
"Why Far Looks Up: Probing Spatial Representation in Vision-Language Models".
The TSV format is used by our probing framework
(cheolhong0916/contrastive-probing)
because it streams trivially and avoids… See the full description on the dataset page: https://huggingface.co/datasets/ch-min/EmbSpatial-Bench-tsv.VLMEvalKit-outputsChMap-Data
ChMapData: Chinese Memory-aware Proactive Dataset
Overview
The Chinese Memory-aware Proactive Dataset (ChMapData) is a novel dataset proposed in the paper "Interpersonal Memory Matters: A New Task for Proactive Dialogue Utilizing Conversational History". This dataset focuses on training and evaluating models' capabilities in proactive topic introduction based on conversational history, supporting the memory-aware proactive dialogue framework proposed in the paper.… See the full description on the dataset page: https://huggingface.co/datasets/FrontierLab/ChMap-Data.so100_dataset08This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 50,
"total_frames": 18407,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/chmadran/so100_dataset08.so100_dataset04This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 50,
"total_frames": 18751,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/chmadran/so100_dataset04.RoboSpatial-Home-tsvai4bharat__samanantar_processed_teThis is extracted from telugu subset from https://huggingface.co/datasets/ai4bharat/samanantar - used to create telugu kenLM models for ASR decoding.
tts-audio-inputsso100_dataset07This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 2,
"total_frames": 212,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/chmadran/so100_dataset07.treesatai-chm
TreeSatAI-CHM
Extension of the TreeSatAI benchmark with canopy height model (CHM) labels derived from German LiDAR airborne surveys (LGLN). Supports joint prediction of tree species (15 genera) and canopy height from Sentinel-1 SAR. Study area: Lower Saxony, Germany.
Code and documentation: GitHub
Files
treesatai-chm/
├── training-data-60m/ # NumPy arrays, ready to use (33 MB)
│ ├── train_x.npy # (40292, 4, 6, 6) float32
│ ├── train_y.npy # (40292… See the full description on the dataset page: https://huggingface.co/datasets/siyux1927/treesatai-chm.openassistant-mouryarecord-testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 0,
"total_frames": 0,
"total_tasks": 0,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/chma3ylma3-2026/record-test.ERQA-tsvCHM-Corr-DatadddNuminaMath-1.5-integer-30kmatchup_finetuning_korjdsf_chm
