datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
symile-m3
Dataset Card for Symile-M3
Symile-M3 is a multilingual dataset of (audio, image, text) samples. The dataset is specifically designed to test a model's ability to capture higher-order information between three distinct high-dimensional data types: by incorporating multiple languages, we construct a task where text and audio are both needed to predict the image, and where, importantly, neither text nor audio alone would suffice.
Paper: https://arxiv.org/abs/2411.01053
GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/arsaporta/symile-m3.ARS
data/ — Directory Structure
All data is gitignored. This file documents what lives here and how it's produced.
paperreview_data/
Crawled ICLR + NeurIPS paper corpus (read-only source of truth).
paperreview_data/
{venue}/ # iclr, neurips
{year}/ # 2017–2026 (ICLR), 2021–2025 (NeurIPS)
papers.jsonl # paper metadata + reviews (official_reviews,
# meta_reviews… See the full description on the dataset page: https://huggingface.co/datasets/Jerry999/ARS.ritual-agent-configsIllusionBencharsma-knowledge-dbFinglish-To-Persian-Dataset-Large
Finglish to Persian Large Dataset
A massive-scale parallel corpus containing over 9.8 million sentence pairs for Finglish (Latin-script Persian) to Persian script transliteration. This dataset provides a robust foundation for training and fine-tuning seq2seq models, normalizing user-generated text, and enhancing Persian input methods.
What is Finglish?
Finglish (also known as Pinglish) is the practice of writing Persian using the Latin alphabet. Because there is… See the full description on the dataset page: https://huggingface.co/datasets/Arshia82sbn/Finglish-To-Persian-Dataset-Large.ar_sarcasm
Dataset Card for ArSarcasm
Dataset Summary
ArSarcasm is a new Arabic sarcasm detection dataset.
The dataset was created using previously available Arabic sentiment analysis
datasets (SemEval 2017
and ASTD) and adds sarcasm and
dialect labels to them.
The dataset contains 10,547 tweets, 1,682 (16%) of which are sarcastic.
For more details, please check the paper
From Arabic Sentiment Analysis to Sarcasm Detection: The ArSarcasm Dataset
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/iabufarha/ar_sarcasm.MotionDecode
ChingMu 1000-Hour Embodied Motion Dataset
High-precision optical motion capture data for humanoid robots, dexterous hands, embodied AI, and virtual production.
Duration
1000+ hours @ 120 Hz
**Scenarios **
15+ real-world scenes
**Tasks **
500+ standardized tasks
Objects
200+ tracked props (6D pose)
Modalities
Skeleton · Finger · Object 6D · Video · Labels
**Formats **
BVH · Retargeted CSV · NPZ
✅ Access note: This dataset is fully open and publicly… See the full description on the dataset page: https://huggingface.co/datasets/arslan219/MotionDecode.alexandria-system
ALEXANDRIA - frozen system artifacts
Every file needed to run the evaluated ALEXANDRIA system: the 41 ZIM archives its
retriever searches at query time, the pre-built index, and the three models.
53 objects, 96.66 GB. Nothing needs to be rebuilt.
Code, installer and verification suite: https://github.com/arsh-imam/alexandria-system
Contents
folder
objects
contents
wikipedia/
9
6 primary archives + wikipedia_en.zim in 3 parts
zim_extra/
34
WikiMed… See the full description on the dataset page: https://huggingface.co/datasets/arshimam/alexandria-system.arsnokyojuu
Bangumi Image Base of Ars No Kyojuu
This is the image base of bangumi Ars no Kyojuu, we detected 70 characters, 5038 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/arsnokyojuu.toxicity_classification_jigsaw
Dataset info
Training Dataset:
You are provided with a large number of Wikipedia comments which have been labeled by human raters for toxic behavior. The types of toxicity are:
toxic
severe_toxic
obscene
threat
insult
identity_hate
The original dataset can be found here: jigsaw_toxic_classification
Our training dataset is a sampled version from the original dataset, containing equal number of samples for both clean and toxic classes.
Dataset creation:… See the full description on the dataset page: https://huggingface.co/datasets/Arsive/toxicity_classification_jigsaw.ArSAS_An_Arabic_Speech-Act_and_Sentiment_Corpus_of_Tweets
ArSAS: An Arabic Speech-Act and Sentiment Corpus of Tweets
Dataset Card for "ArSAS: An Arabic Speech-Act and Sentiment Corpus of Tweets"
Note About Sentiment_label_confidence
"Crowdflower provides a confidence score with each annotated tweet that represents the confidence in the quality of the label. For a three annotators per tweet setup, the confidence score would range between 0.3 and 1 according to two factors: 1) annotator quality level; and 2) agreement… See the full description on the dataset page: https://huggingface.co/datasets/Qanadil/ArSAS_An_Arabic_Speech-Act_and_Sentiment_Corpus_of_Tweets.ArSASars-magna-greatest-hits
Ars Magna Greatest Hits
The funniest and most apt anagrams of people, companies, products, titles, places and phrases, found by Ars Magna and kept by hand.
Every row is a real anagram: the words use exactly the input's letters, checked against a pinned revision of English OpenList (368bf0e4460461c985fca8bde49e4062d56c1516), and every word is in the tier the row names. Accented letters fold to their base letter, so Beyoncé has three e's. Nothing typed is ever replaced by… See the full description on the dataset page: https://huggingface.co/datasets/ryanjosephkamp/ars-magna-greatest-hits.WallpapersarsivTranslation-Dataset-Large
Translation-Dataset_Large 🌍
A massive, unified multilingual parallel corpus for Persian, English, and Arabic NMT research.
Translation-Dataset_Large aggregates and normalizes millions of sentence pairs across three language directions — English↔Persian (en-fa), Arabic↔English (ar-en), and Arabic↔Persian (ar-fa) — into a single, deduplicated, research-ready .parquet dataset.
Dataset Summary
Property
Value
Languages
Persian (fa), English (en), Arabic… See the full description on the dataset page: https://huggingface.co/datasets/Arshia82sbn/Translation-Dataset-Large.ARSG-110K
ARSG-110K
Project Page | Paper | GitHub
ARSG-110K is a large-scale scene-level dataset comprising over 110K diverse scenes and 3M annotated images with high-fidelity 3D ground truth. It is designed to support the training and evaluation of compositional 3D scene generation and in-place completion models. The dataset provides accurate 3D object-level ground-truth, layout, and annotations.
This dataset was introduced as part of the paper: 3D-Fixer: Coarse-to-Fine In-place Completion… See the full description on the dataset page: https://huggingface.co/datasets/HorizonRobotics/ARSG-110K.record-test
record-test
This dataset was generated using phosphobot.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot.
To get started in robotics, get your own phospho starter pack..
arsentd_levThe Arabic Sentiment Twitter Dataset for Levantine dialect (ArSenTD-LEV) contains 4,000 tweets written in Arabic and equally retrieved from Jordan, Lebanon, Palestine and Syria.POWDER_CoChannel_Protocol_Dataset
POWDER Co-Channel Protocol (PCP) Dataset
768 real-world over-the-air (OTA) IQ captures for multi-label RF fingerprinting under co-channel interference, with heterogeneous waveforms (802.11a Wi‑Fi, 4G LTE, and 5G NR) collected on the POWDER PAWR testbed at the University of Utah by the CREDIT Center, Prairie View A&M University.
Each capture records the superposition of up to six simultaneously transmitting USRP radios on a shared 20 MHz channel at 2.425 GHz. Every .bin IQ file… See the full description on the dataset page: https://huggingface.co/datasets/T-Arshad/POWDER_CoChannel_Protocol_Dataset.ArSAS
Dataset Card for "ArSAS"
More Information needed
kana-sounds
Kana Sounds
147 short spoken clips, one for every hiragana and katakana character used by
Kana Trainer: the 46 seion, 20 dakuon,
5 handakuon, 33 yoon and 43 tokushon. They come from a single reader on
FUN Japanese Learning.
Dataset structure
audio/
seion/ 46 clips a.mp3, i.mp3, ka.mp3, ... n.mp3
dakuon/ 20 clips ga.mp3, za.mp3, ji.mp3, ... bo.mp3
handakuon/ 5 clips pa.mp3, pi.mp3, pu.mp3, pe.mp3, po.mp3
yoon/ 33 clips kya.mp3… See the full description on the dataset page: https://huggingface.co/datasets/arsalan-anwari/kana-sounds.omni-dreams-samples
AlpaDreams Samples
Curated single-view driving sequences for evaluating the
nvidia/alpadreams-dit world model.
Layout
data/
└── single_view/
├── <clip-id>/
| ├── <clip-id_...>.mp4 # ground truth video
│ ├── <clip-id_..._hdmap>.mp4 # HD-map rasterized conditioning video
│ ├── first_frame.png # RGB first frame, extracted from ground truth video
│ └── prompt.txt # text prompt
└──… See the full description on the dataset page: https://huggingface.co/datasets/Arsh9210/omni-dreams-samples.details_arshadshk__Mistral-Hinglish-7B-Instruct-v0.2
Dataset Card for Evaluation run of arshadshk/Mistral-Hinglish-7B-Instruct-v0.2
Dataset automatically created during the evaluation run of model arshadshk/Mistral-Hinglish-7B-Instruct-v0.2 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_arshadshk__Mistral-Hinglish-7B-Instruct-v0.2.SWE-Zero-openhands-trajectories
SWE-Zero Trajectories: Execution-free Fine-tuning for Software Engineering Agents
Data Overview
SWE-ZERO Trajectories is an agentic instruction tuning dataset designed to advance the capabilities of LLMs in software engineering. This dataset comprises 318k agent
trajectories collected using the OpenHands framework. The trajectories
were synthesized using Qwen3-Coder-480B-A35B-Instruct, specifically curated for supervised fine-tuning (SFT),
aiming to improve… See the full description on the dataset page: https://huggingface.co/datasets/Arsh9210/SWE-Zero-openhands-trajectories.Open-SWE-Traces
Open-SWE-Traces: Advancing Distillation for Software Engineering Agents
Data Overview
Open-SWE-Traces is an agentic instruction tuning dataset designed to advance the capabilities of LLMs in software engineering. This dataset comprises 200k+ agent
trajectories collected using the SWE-agent and OpenHands framework. The trajectories
were synthesized using Minimax-M2.5 (with thinking) and Qwen3.5-122B-A10B
(without thinking) and specifically curated for supervised… See the full description on the dataset page: https://huggingface.co/datasets/Arsh9210/Open-SWE-Traces.egora-benchmarks
EgoRA Benchmark Results
Comprehensive benchmark results for EgoRA (Entropy-Governed Orthogonality Regularization for Adaptation) across multiple model scales, domains, and architectures.
📦 Package: egora on PyPI
💻 Code: ArsSocratica/EgoRA on GitHub
📄 Paper: arXiv:2602.05192
🔖 DOI: 10.5281/zenodo.19398709
Dataset Structure
llama-3.2-1b/, llama-3.2-3b/, llama-3.1-8b/
Fine-tuning results across 3 model scales, 2 domains (Alpaca general, Medical), 4… See the full description on the dataset page: https://huggingface.co/datasets/ArsSocratica/egora-benchmarks.ar-sa-tts-speakers-synthetic
Deprecated -- consolidated
This repo's data files stay in place, but the rows now live as named config(s) on the single Salesteq synthetic-speech dataset:
speakers-v1 on Salesteq/ar-sa-tts-corpus-synthetic
Load from there rather than this repo going forward.
ar-sa-tts-speakers — multi-speaker Najdi Arabic TTS (synthetic)
Ten single-speaker synthetic Najdi Arabic TTS sets unified into one dataset, distinguished by the speaker column. 91,350 clips across 10… See the full description on the dataset page: https://huggingface.co/datasets/Salesteq/ar-sa-tts-speakers-synthetic.kl3m-data-dotgov-www.ars.usda.gov
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.ars.usda.gov.
