datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
beat2-additional-annotations
BEAT2 Official Release + Additional Annotations
This is a fork of H-Liu1997/BEAT2
that adds annotations contributed by the
RAG-Gesture (CVPR 2025)
and MIBURI (CVPR 2026) projects.
The base BEAT2-English data (motion, audio, TextGrids, semantic labels,
pretrained motion-autoencoder weights) is inherited verbatim from upstream;
the additional annotations from RAG-Gesture and MIBURI are pushed on top.
Citations
If you use only the original BEAT2 dataset, please cite… See the full description on the dataset page: https://huggingface.co/datasets/m-hamza-mughal/beat2-additional-annotations.wizardlm8x22b-logical-math-coding-sft_additional
自動生成したテキスト
WizardLM 8x22bで生成した論理・数学・コード系のデータです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
bank-additional-fullarithmetic_additionaddition-datasetworldedit_addition_v2addition_dataset
Addition Dataset
Addition problems in the format {a} + {b} = {c}.
Subsets
test: 5K held-out evaluation examples (operands >= 10, i.e. min 2 digits)
1BT: 85M training examples (1 billion tokens under Llama-3 tokenizer)
10BT: 850M training examples (10 billion tokens)
3MT-3digit: Exhaustive single-token addition: all (a, b) with a, b in [0, 999] and a+b <= 999. 500,500 ordered pairs, ~3M tokens. All of a, b, c are single tokens. Symmetry-safe train/test split (10% test).… See the full description on the dataset page: https://huggingface.co/datasets/deqing/addition_dataset.task753_svamp_addition_question_answering
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task753_svamp_addition_question_answering
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task753_svamp_addition_question_answering.slam_stage2_additional_dataopenbookqa_additional_promptsourcellama3_additional_rr40k_non_delete_sftworldedit_addition_v1Additional_Yoruba_Databpe-single-multi-token-additionquirky_addition_increment0
Dataset Card for "quirky_addition_increment0"
More Information needed
llama3_additional_rr80k_NON_balanced_sftquirky_addition_rawllama31_no_additional_chat_formatllama3_additional_rr40k_non_delete_sft_chat_formatquirky_addition_increment3_bob_hard
Dataset Card for "quirky_addition_increment3_bob_hard"
More Information needed
multilingual-addition
Multilingual Addition Dataset
Synthetic dataset of addition problems of the form a+b=answer, where a
and b are written-form representations of integers in 21 languages, plus
a 22nd split using raw digit strings.
Task format
Each sample contains:
field
type
description
a_str
str
written-form (or digit) representation of a
a_digit
int
integer value of a
b_str
str
written-form (or digit) representation of b
b_digit
int
integer value of b
answer
str… See the full description on the dataset page: https://huggingface.co/datasets/flexitok/multilingual-addition.llama3_additional_rr40k_NON_balanced_sftlemonseed-addition-carry-method
lemonseed-addition-carry-method
LemonSeed — carry-method addition scratchpad (LSB-first, single-digit facts). Teaches digit-level addition with explicit written steps.
Contents
addition.jsonl (10000 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
llama3_additional_rr40k_balanced_sftllama3_additional_rr10k_NON_balanced_sftaddition_dataset_5_milliongaperon-distill-additionsquirky_addition_increment0_alice
Dataset Card for "quirky_addition_increment0_alice"
More Information needed
langchain-additional-resourceswizardlm8x22b-logical-math-coding-sft_additional-ja
自動生成したテキスト
WizardLM 8x22bで生成した論理・数学・コード系のデータを、Calm3-22bで翻訳したものです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました
