datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
my-storagejitteredwebsites-merged-224-paraphrasedjitsuwaoresaikyoudeshita
Bangumi Image Base of Jitsu Wa Ore, Saikyou Deshita?
This is the image base of bangumi Jitsu wa Ore, Saikyou deshita?, we detected 72 characters, 4946 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1%… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/jitsuwaoresaikyoudeshita.easyr1-114k-hard-qwen7b-easy-gta1-4MP-nores-jedi-fix-synced-aug-jitter-tokenizedSWE-lite-trace100
SWE-lite trace100
Recorded agent workloads for fixed-trace inference-serving performance regression.
This dataset contains the first 100 instances in a frozen SWE-bench Lite test
selection, with 6,490 model requests, 6,390 inter-turn delays, and
1,468,636 recorded completion tokens. It was reconstructed from existing
GLM-4.7-Flash / upstream vLLM MiniSWE runs, without running the agents again.
There are 55 Submitted and 45 LimitsExceeded terminal trajectories. These are
agent… See the full description on the dataset page: https://huggingface.co/datasets/JitaiHao/SWE-lite-trace100.Methods2Test_java_unit_test_code
Dataset Description
Microsoft created this large dataset of Java Junit test cases with its corresponding focal methods.
It contains 780k pairs of JUnit test cases and focal methods which were extracted from a total of 91K
Java open source project hosted on GitHub.
The mapping between test case and focal methods are based heuristics rules and Java developer's best practice.
More information could be found here:
methods2test Github repo
Methods2Test: A dataset of focal methods… See the full description on the dataset page: https://huggingface.co/datasets/jitx/Methods2Test_java_unit_test_code.ariel-2025-jitter-decorrelated-cachejits-legal-dataset
JITS Legal Dataset
The only major open Indian legal dataset with citation-graph extraction, statutory section tagging, and IPC/CrPC→BNS/BNSS transition mapping — verified against the schemas of every comparable open dataset in this space (including one at 17.1M rows), which provide raw text and/or task labels but not structured legal extraction.
Overview
Disclaimer: This dataset is independently created for research and engineering use. It is not an official… See the full description on the dataset page: https://huggingface.co/datasets/Viverun/jits-legal-dataset.JitOPD-OpenR1-Math-220k-Teacher-Memory
JitOPD OpenR1-Math-220k Teacher Prefix-Logit Memory
This dataset contains sparse teacher next-token logits collected for JitOPD
retrieval-augmented decoding. The source prompts are the default configuration
of open-r1/OpenR1-Math-220k,
and the teacher is
Qwen/Qwen2.5-Math-7B-Instruct.
Only teacher trajectories whose final boxed answer passes both a numeric
signature prefilter and Math-Verify are retained. This release contains raw
teacher prefix/logit memory and does not contain… See the full description on the dataset page: https://huggingface.co/datasets/sadadasdasdas/JitOPD-OpenR1-Math-220k-Teacher-Memory.jits-legal-dataset
JITS Legal Dataset
A production-ready, deterministic pipeline for processing Indian legal judgments into structured, high-quality legal datasets — with comprehensive extraction, self-citation exclusion, and multi-act statutory section detection.
Overview
Disclaimer: This dataset is independently created for research and engineering use. It is not an official government or judicial release and does not constitute legal advice.
The JITS Legal Dataset currently… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/jits-legal-dataset.easyr1-114k-hard-qwen7b-easy-gta1-4MP-nores-jedi-fix-synced-aug-jitter-coord-grid
easyr1-114k-hard-qwen7b-easy-gta1-4MP-nores-jedi-fix-synced-aug-jitter-coord-grid
Augmented version of easyr1-114k-hard-qwen7b-easy-gta1-4MP-nores-jedi-fix-synced-aug-jitter with a fixed 100px coordinate grid overlay.
Each image is overlaid with vertical and horizontal grid lines every 100
pixels at native resolution. Major ticks (every 1 steps)
are emphasized and axis labels show pixel values to help models localize
precise coordinates.
Summary
Generated on:… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-114k-hard-qwen7b-easy-gta1-4MP-nores-jedi-fix-synced-aug-jitter-coord-grid.easyr1-126k-nores-jedi-fix-synced-ui-vision-manually-labeled-icon-data-from-yt-4MP-aug-jitter
easyr1-126k-nores-jedi-fix-synced-ui-vision-manually-labeled-icon-data-from-yt-4MP-aug-jitter
Augmented version of datasets/easyr1-63k-nores-jedi-fix-synced-ui-vision-manually-labeled-icon-data-from-yt-4MP with coordinate jitter.
For each original example, 1 additional copies were created. Each copy
randomly jitters the target coordinate by ±1 pixel in both X and Y. The
assistant coordinate in messages is updated, and bbox/normalized_bbox
are shifted when present.… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-126k-nores-jedi-fix-synced-ui-vision-manually-labeled-icon-data-from-yt-4MP-aug-jitter.tlcThai Literature Corpora (TLC): Corpora of machine-ingestible Thai classical literature texts.
Release: 6/25/19
It consists of two datasets:
## TLC set
It is texts from [Vajirayana Digital Library](https://vajirayana.org/), stored by chapters and stanzas (non-tokenized).
tlc v.2.0 (6/17/19 : a total of 34 documents, 292,270 lines, 31,790,734 characters)
tlc v.1.0 (6/11/19 : a total of 25 documents, 113,981 lines, 28,775,761 characters)
## TNHC set
It is texts from Thai National Historical Corpus, stored by lines (manually tokenized).
tnhc v.1.0 (6/25/19 : a total of 47 documents, 756,478 lines, 13,361,142 characters)easyr1-114k-hard-qwen7b-easy-gta1-4MP-nores-jedi-fix-synced-aug-jitter
easyr1-114k-hard-qwen7b-easy-gta1-4MP-nores-jedi-fix-synced-aug-jitter
Augmented version of easyr1-57k-hard-qwen7b-easy-gta1-4MP-nores-jedi-fix-synced with coordinate jitter.
For each original example, 1 additional copies were created. Each copy
randomly jitters the target coordinate by ±1 pixel in both X and Y. The
assistant coordinate in messages is updated, and bbox/normalized_bbox
are shifted when present.
Summary
Generated on: 2025-09-07 17:36:27 UTC
Source… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-114k-hard-qwen7b-easy-gta1-4MP-nores-jedi-fix-synced-aug-jitter.ariel-2025-jitter-decorrelated-robust-cachejit-meta-harness
JIT Meta-Harness
JIT Meta-Harness contains a subset of the training data used in the JIT-Agent project. Each example pairs a task instruction with a structured agent harness represented by four Python modules and a prompt configuration.
Dataset Summary
The repository provides a single training split in JSON Lines format.
Statistic
Value
Training examples
2,852
Unique task instructions
1,427
Harnesses per instruction
1--3
All examples contain… See the full description on the dataset page: https://huggingface.co/datasets/JIT-Agent/jit-meta-harness.franka_search_not_avoid_calculator_compass10_close_mixed_jitter20x_v1_seed20260626_molmobotfranka_pick_and_place_avoid_calculator_compass10_close_mixed_jitter20x_v2_seed20260626_molmobotfranka_avoid_calculator_compass10_close_mixed_jitter20x_v2_seed20260626_classic_overlay_molmobotjitteredwebsites-merged-224-paraphrased-pairedso100_flip_jitter_100nyaya-manual-assetsCVPR2025_EditAR_releasedawa-voice-kirundi-fr
Dawa Voice — Kirundi/French Benchmark Audio
Code-switched Kirundi/French audio clips recorded for clinical triage STT benchmarking.
Challenge: MLC Africa × Intron Agentic Voice AI Challenge 2026 — Health categoryProject: Dawa Voice — Voice-driven clinical triage for Burundi
Dataset Details
Language pair: Kirundi (rn) ↔ French (fr)
Domain: Healthcare — symptom descriptions, emergency phrases
Country: Burundi
Total duration: ~46 seconds (10 clips)
Format: M4A (AAC… See the full description on the dataset page: https://huggingface.co/datasets/Jitimay/dawa-voice-kirundi-fr.CVPR2022_CoordGAN_releaseAlargeDatabaseethercat-master-jitter-logs
EtherCAT Master Jitter Logs
这是 motor_jitter_viewer 的实机 EtherCAT 周期日志数据集,用于横向查看 EC-Master、SOEM 和 IgH 三种主站下,右臂 D18–D24 的位置跟踪与 1 ms 控制周期波动。
配套查看器和分析说明:https://github.com/zhuzq2020/motor_jitter_viewer
数据结构
raw/
ecmaster/cycle_ecmaster_20260908_141617.csv
soem/cycle_soem_20260908_141257.csv
igh/cycle_igh_20260908_140958.csv
metadata/
runs.csv
三次记录使用相同的 fixed_end_comparison_demo 固定右臂末端轨迹,峰值转速设为 3 RPM;控制周期为 1 ms,网口为 enp2s0,记录关节 D18–D24。metadata/runs.csv… See the full description on the dataset page: https://huggingface.co/datasets/zaki2022/ethercat-master-jitter-logs.DATAENV
Trading Forecasting Dataset
Upload the contents of this dataset folder to a separate Hugging Face Dataset repository.
Expected Dataset repo root:
Data/
Alt Data/
README.md
.gitattributes
The Hugging Face Space backend expects these folders to hydrate into its research_runtime folder:
Data/
Alt Data/
After uploading this dataset repo, set this Space environment variable:
HF_DATASET_REPO_ID=your-hf-username/your-forecasting-dataset
Optional:
HF_DATASET_REVISION=main
The backend… See the full description on the dataset page: https://huggingface.co/datasets/Jitendra12421/DATAENV.deltakv_qwen_train_v3.0_num40000_seqlen8192jithinraj-voice
