datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RefRef_additionalRefRef: A Synthetic Dataset and Benchmark for Reconstructing Refractive and Reflective Objects
Yue Yin ·
Enze Tao ·
Weijian Deng ·
Dylan Campbell
About
This repository provides additional data for the RefRef dataset.
Citation
@misc{yin2025refrefsyntheticdatasetbenchmark,
title={RefRef: A Synthetic Dataset and Benchmark for Reconstructing Refractive and Reflective Objects},
author={Yue Yin and Enze Tao and… See the full description on the dataset page: https://huggingface.co/datasets/yinyue27/RefRef_additional.beat2-additional-annotations
BEAT2 Official Release + Additional Annotations
This is a fork of H-Liu1997/BEAT2
that adds annotations contributed by the
RAG-Gesture (CVPR 2025)
and MIBURI (CVPR 2026) projects.
The base BEAT2-English data (motion, audio, TextGrids, semantic labels,
pretrained motion-autoencoder weights) is inherited verbatim from upstream;
the additional annotations from RAG-Gesture and MIBURI are pushed on top.
Citations
If you use only the original BEAT2 dataset, please cite… See the full description on the dataset page: https://huggingface.co/datasets/m-hamza-mughal/beat2-additional-annotations.camera_pizza_additionalwizardlm8x22b-logical-math-coding-sft_additional
自動生成したテキスト
WizardLM 8x22bで生成した論理・数学・コード系のデータです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
bank-additional-fullarithmetic_additionaddition-datasetworldedit_addition_v2addition_dataset
Addition Dataset
Addition problems in the format {a} + {b} = {c}.
Subsets
test: 5K held-out evaluation examples (operands >= 10, i.e. min 2 digits)
1BT: 85M training examples (1 billion tokens under Llama-3 tokenizer)
10BT: 850M training examples (10 billion tokens)
3MT-3digit: Exhaustive single-token addition: all (a, b) with a, b in [0, 999] and a+b <= 999. 500,500 ordered pairs, ~3M tokens. All of a, b, c are single tokens. Symmetry-safe train/test split (10% test).… See the full description on the dataset page: https://huggingface.co/datasets/deqing/addition_dataset.20251206-redcube-additionalGenText-Forensics_third_place_additional_materials
GenText-Forensics 2026 — Third-Place Additional Materials (Team MSU)
Model weights, code, and reproduction artifacts for Team MSU's third-place
solution to the ACM MM 2026 GenText-Forensics challenge
(Codabench).
The method is a decomposed chain-of-thought pipeline for detecting,
localizing, typing, and explaining forgeries in multilingual document text
images:
DTD (Document Tampering Detector) — an external pixel-level visual
tampering detector that produces a tampering… See the full description on the dataset page: https://huggingface.co/datasets/cmcshnik/GenText-Forensics_third_place_additional_materials.bank-marketing-additional
Dataset Card for Bank Marketing (additional)
This dataset is a precise version of UCI Bank Marketing
We first created the default bank marketing dataset, as seen here. Then we further run the following Python script to create this additional portion.
# Define feature types
continuous_columns = ["age", "duration", "campaign", "pdays", "previous",
"emp.var.rate", "cons.price.idx", "cons.conf.idx",
"euribor3m", "nr.employed"]… See the full description on the dataset page: https://huggingface.co/datasets/cestwc/bank-marketing-additional.ny_test_frames_green_addition_30This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "sam_evt2",
"total_episodes": 30,
"total_frames": 17686,
"total_tasks": 1,
"total_videos": 120,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 60,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/1g0rrr/ny_test_frames_green_addition_30.task753_svamp_addition_question_answering
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task753_svamp_addition_question_answering
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task753_svamp_addition_question_answering.eval_results_pi05_basic_golden_additional_seed100kslam_stage2_additional_dataopenbookqa_additional_promptsourcepersian-ocr-bench-submitted10-additional4-results
Persian OCR benchmark
This run evaluates 4 vision OCR lanes over 193 bbox crops from Reza2kn/persian-ocr-bench-submitted10-bbox-crops.
Each row in results.jsonl preserves the crop identity and current gold content_text, then records the model output, latency, usage, provider hint, HTTP status, and normalized OCR metrics. Failures are retained.
OpenRouter constraints: Grok uses xai/zdr, GPT uses openai, GLM uses baseten/fp8, Gemma 4 26B uses cloudflare, and Gemma 4 31B uses… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-ocr-bench-submitted10-additional4-results.llama3_additional_rr40k_non_delete_sftworldedit_addition_v1libero_goal_iiwa_additionalCams_failuresThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "iiwa",
"total_episodes": 116,
"total_frames": 17299,
"total_tasks": 9,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:116"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/OliverHausdoerfer/libero_goal_iiwa_additionalCams_failures.piper_isaacsim_top_wrist_D1_120ep_additionedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "isaac_piper",
"total_episodes": 120,
"total_frames": 29989,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:120"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/shenj/piper_isaacsim_top_wrist_D1_120ep_additioned.additional_fontsAdditional_Yoruba_Databpe-single-multi-token-additionpiper_isaacsim_top_wrist_D1_20ep_additionThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "isaac_piper",
"total_episodes": 20,
"total_frames": 4979,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/shenj/piper_isaacsim_top_wrist_D1_20ep_addition.teleop-trail-mix-23-cotrain-50fpsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "addition-openarm",
"total_episodes": 521,
"total_frames": 201382,
"total_tasks": 32,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 50,
"splits": {
"train": "0:521"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/addition-robotics/teleop-trail-mix-23-cotrain-50fps.quirky_addition_increment0
Dataset Card for "quirky_addition_increment0"
More Information needed
llama3_additional_rr80k_NON_balanced_sftaddition_decimal
