CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01arsaporta /symile-m3 Dataset Card for Symile-M3 Symile-M3 is a multilingual dataset of (audio, image, text) samples. The dataset is specifically designed to test a model's ability to capture higher-order information between three distinct high-dimensional data types: by incorporating multiple languages, we construct a task where text and audio are both needed to predict the image, and where, importantly, neither text nor audio alone would suffice. Paper: https://arxiv.org/abs/2411.01053 GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/arsaporta/symile-m3.audiozero-shot-classification10M<n<100M8 likes23k downloads2y agoHugging Face02Jerry999 /ARS data/ — Directory Structure All data is gitignored. This file documents what lives here and how it's produced. paperreview_data/ Crawled ICLR + NeurIPS paper corpus (read-only source of truth). paperreview_data/ {venue}/ # iclr, neurips {year}/ # 2017–2026 (ICLR), 2021–2025 (NeurIPS) papers.jsonl # paper metadata + reviews (official_reviews, # meta_reviews… See the full description on the dataset page: https://huggingface.co/datasets/Jerry999/ARS.0 likes4.2k downloads28d agoHugging Face03arshan-ritual /ritual-agent-configs2 likes2.4k downloads23d agoHugging Face04arshiahemmat /IllusionBenchimage10K<n<100K3 likes676 downloads2y agoHugging Face05houlab /arsma-knowledge-dbtext100K<n<1M0 likes672 downloads2d agoHugging Face06Arshia82sbn /Finglish-To-Persian-Dataset-Large Finglish to Persian Large Dataset A massive-scale parallel corpus containing over 9.8 million sentence pairs for Finglish (Latin-script Persian) to Persian script transliteration. This dataset provides a robust foundation for training and fine-tuning seq2seq models, normalizing user-generated text, and enhancing Persian input methods. What is Finglish? Finglish (also known as Pinglish) is the practice of writing Persian using the Latin alphabet. Because there is… See the full description on the dataset page: https://huggingface.co/datasets/Arshia82sbn/Finglish-To-Persian-Dataset-Large.texttranslation10M<n<100M1 likes621 downloads2mo agoHugging Face07iabufarha /ar_sarcasm Dataset Card for ArSarcasm Dataset Summary ArSarcasm is a new Arabic sarcasm detection dataset. The dataset was created using previously available Arabic sentiment analysis datasets (SemEval 2017 and ASTD) and adds sarcasm and dialect labels to them. The dataset contains 10,547 tweets, 1,682 (16%) of which are sarcastic. For more details, please check the paper From Arabic Sentiment Analysis to Sarcasm Detection: The ArSarcasm Dataset Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/iabufarha/ar_sarcasm.texttext-classification10K<n<100K18 likes476 downloads3y agoHugging Face08arslan219 /MotionDecode ChingMu 1000-Hour Embodied Motion Dataset High-precision optical motion capture data for humanoid robots, dexterous hands, embodied AI, and virtual production. Duration 1000+ hours @ 120 Hz **Scenarios ** 15+ real-world scenes **Tasks ** 500+ standardized tasks Objects 200+ tracked props (6D pose) Modalities Skeleton · Finger · Object 6D · Video · Labels **Formats ** BVH · Retargeted CSV · NPZ ✅ Access note: This dataset is fully open and publicly… See the full description on the dataset page: https://huggingface.co/datasets/arslan219/MotionDecode.0 likes429 downloads2mo agoHugging Face09arshimam /alexandria-system ALEXANDRIA - frozen system artifacts Every file needed to run the evaluated ALEXANDRIA system: the 41 ZIM archives its retriever searches at query time, the pre-built index, and the three models. 53 objects, 96.66 GB. Nothing needs to be rebuilt. Code, installer and verification suite: https://github.com/arsh-imam/alexandria-system Contents folder objects contents wikipedia/ 9 6 primary archives + wikipedia_en.zim in 3 parts zim_extra/ 34 WikiMed… See the full description on the dataset page: https://huggingface.co/datasets/arshimam/alexandria-system.text10B<n<100B0 likes405 downloads18d agoHugging Face10BangumiBase /arsnokyojuu Bangumi Image Base of Ars No Kyojuu This is the image base of bangumi Ars no Kyojuu, we detected 70 characters, 5038 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability). Here is the… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/arsnokyojuu.image1K<n<10K0 likes393 downloads2y agoHugging Face11Arsive /toxicity_classification_jigsaw Dataset info Training Dataset: You are provided with a large number of Wikipedia comments which have been labeled by human raters for toxic behavior. The types of toxicity are: toxic severe_toxic obscene threat insult identity_hate The original dataset can be found here: jigsaw_toxic_classification Our training dataset is a sampled version from the original dataset, containing equal number of samples for both clean and toxic classes. Dataset creation:… See the full description on the dataset page: https://huggingface.co/datasets/Arsive/toxicity_classification_jigsaw.tabulartext-classification100K<n<1M5 likes392 downloads3y agoHugging Face12Qanadil /ArSAS_An_Arabic_Speech-Act_and_Sentiment_Corpus_of_Tweets ArSAS: An Arabic Speech-Act and Sentiment Corpus of Tweets Dataset Card for "ArSAS: An Arabic Speech-Act and Sentiment Corpus of Tweets" Note About Sentiment_label_confidence "Crowdflower provides a confidence score with each annotated tweet that represents the confidence in the quality of the label. For a three annotators per tweet setup, the confidence score would range between 0.3 and 1 according to two factors: 1) annotator quality level; and 2) agreement… See the full description on the dataset page: https://huggingface.co/datasets/Qanadil/ArSAS_An_Arabic_Speech-Act_and_Sentiment_Corpus_of_Tweets.tabulartext-classification10K<n<100K1 likes390 downloads2y agoHugging Face13Osaleh /ArSAStext10K<n<100K0 likes377 downloads4y agoHugging Face14ryanjosephkamp /ars-magna-greatest-hits Ars Magna Greatest Hits The funniest and most apt anagrams of people, companies, products, titles, places and phrases, found by Ars Magna and kept by hand. Every row is a real anagram: the words use exactly the input's letters, checked against a pinned revision of English OpenList (368bf0e4460461c985fca8bde49e4062d56c1516), and every word is in the tier the row names. Accented letters fold to their base letter, so Beyoncé has three e's. Nothing typed is ever replaced by… See the full description on the dataset page: https://huggingface.co/datasets/ryanjosephkamp/ars-magna-greatest-hits.texttext-classificationn<1K0 likes360 downloads6h agoHugging Face15Ars-btw /Wallpapersimagen<1K0 likes316 downloads8d agoHugging Face16benolanben /arsivaudion<1K0 likes291 downloads1mo agoHugging Face17Arshia82sbn /Translation-Dataset-Large Translation-Dataset_Large 🌍 A massive, unified multilingual parallel corpus for Persian, English, and Arabic NMT research. Translation-Dataset_Large aggregates and normalizes millions of sentence pairs across three language directions — English↔Persian (en-fa), Arabic↔English (ar-en), and Arabic↔Persian (ar-fa) — into a single, deduplicated, research-ready .parquet dataset. Dataset Summary Property Value Languages Persian (fa), English (en), Arabic… See the full description on the dataset page: https://huggingface.co/datasets/Arshia82sbn/Translation-Dataset-Large.texttranslation10M<n<100M0 likes291 downloads2mo agoHugging Face18HorizonRobotics /ARSG-110K ARSG-110K Project Page | Paper | GitHub ARSG-110K is a large-scale scene-level dataset comprising over 110K diverse scenes and 3M annotated images with high-fidelity 3D ground truth. It is designed to support the training and evaluation of compositional 3D scene generation and in-place completion models. The dataset provides accurate 3D object-level ground-truth, layout, and annotations. This dataset was introduced as part of the paper: 3D-Fixer: Coarse-to-Fine In-place Completion… See the full description on the dataset page: https://huggingface.co/datasets/HorizonRobotics/ARSG-110K.3dimage-to-3d2 likes281 downloads6mo agoHugging Face19arsenxeno /record-test record-test This dataset was generated using phosphobot. This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot. To get started in robotics, get your own phospho starter pack.. tabularrobotics10K<n<100K0 likes206 downloads1y agoHugging Face20ramybaly /arsentd_levThe Arabic Sentiment Twitter Dataset for Levantine dialect (ArSenTD-LEV) contains 4,000 tweets written in Arabic and equally retrieved from Jordan, Lebanon, Palestine and Syria.text-classification1K<n<10K5 likes196 downloads3y agoHugging Face21T-Arshad /POWDER_CoChannel_Protocol_Dataset POWDER Co-Channel Protocol (PCP) Dataset 768 real-world over-the-air (OTA) IQ captures for multi-label RF fingerprinting under co-channel interference, with heterogeneous waveforms (802.11a Wi‑Fi, 4G LTE, and 5G NR) collected on the POWDER PAWR testbed at the University of Utah by the CREDIT Center, Prairie View A&M University. Each capture records the superposition of up to six simultaneously transmitting USRP radios on a shared 20 MHz channel at 2.425 GHz. Every .bin IQ file… See the full description on the dataset page: https://huggingface.co/datasets/T-Arshad/POWDER_CoChannel_Protocol_Dataset.textn<1K0 likes188 downloads4d agoHugging Face22arbml /ArSAS Dataset Card for "ArSAS" More Information needed text10K<n<100K1 likes187 downloads4y agoHugging Face23arsalan-anwari /kana-sounds Kana Sounds 147 short spoken clips, one for every hiragana and katakana character used by Kana Trainer: the 46 seion, 20 dakuon, 5 handakuon, 33 yoon and 43 tokushon. They come from a single reader on FUN Japanese Learning. Dataset structure audio/ seion/ 46 clips a.mp3, i.mp3, ka.mp3, ... n.mp3 dakuon/ 20 clips ga.mp3, za.mp3, ji.mp3, ... bo.mp3 handakuon/ 5 clips pa.mp3, pi.mp3, pu.mp3, pe.mp3, po.mp3 yoon/ 33 clips kya.mp3… See the full description on the dataset page: https://huggingface.co/datasets/arsalan-anwari/kana-sounds.audioaudio-classificationn<1K1 likes162 downloads18d agoHugging Face24Arsh9210 /omni-dreams-samples AlpaDreams Samples Curated single-view driving sequences for evaluating the nvidia/alpadreams-dit world model. Layout data/ └── single_view/ ├── <clip-id>/ | ├── <clip-id_...>.mp4 # ground truth video │ ├── <clip-id_..._hdmap>.mp4 # HD-map rasterized conditioning video │ ├── first_frame.png # RGB first frame, extracted from ground truth video │ └── prompt.txt # text prompt └──… See the full description on the dataset page: https://huggingface.co/datasets/Arsh9210/omni-dreams-samples.imageimage-to-videon<1K0 likes160 downloads2mo agoHugging Face25open-llm-leaderboard-old /details_arshadshk__Mistral-Hinglish-7B-Instruct-v0.2 Dataset Card for Evaluation run of arshadshk/Mistral-Hinglish-7B-Instruct-v0.2 Dataset automatically created during the evaluation run of model arshadshk/Mistral-Hinglish-7B-Instruct-v0.2 on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_arshadshk__Mistral-Hinglish-7B-Instruct-v0.2.0 likes152 downloads3y agoHugging Face26Arsh9210 /SWE-Zero-openhands-trajectories SWE-Zero Trajectories: Execution-free Fine-tuning for Software Engineering Agents Data Overview SWE-ZERO Trajectories is an agentic instruction tuning dataset designed to advance the capabilities of LLMs in software engineering. This dataset comprises 318k agent trajectories collected using the OpenHands framework. The trajectories were synthesized using Qwen3-Coder-480B-A35B-Instruct, specifically curated for supervised fine-tuning (SFT), aiming to improve… See the full description on the dataset page: https://huggingface.co/datasets/Arsh9210/SWE-Zero-openhands-trajectories.text100K<n<1M0 likes149 downloads2mo agoHugging Face27Arsh9210 /Open-SWE-Traces Open-SWE-Traces: Advancing Distillation for Software Engineering Agents Data Overview Open-SWE-Traces is an agentic instruction tuning dataset designed to advance the capabilities of LLMs in software engineering. This dataset comprises 200k+ agent trajectories collected using the SWE-agent and OpenHands framework. The trajectories were synthesized using Minimax-M2.5 (with thinking) and Qwen3.5-122B-A10B (without thinking) and specifically curated for supervised… See the full description on the dataset page: https://huggingface.co/datasets/Arsh9210/Open-SWE-Traces.text100K<n<1M0 likes146 downloads2mo agoHugging Face28ArsSocratica /egora-benchmarks EgoRA Benchmark Results Comprehensive benchmark results for EgoRA (Entropy-Governed Orthogonality Regularization for Adaptation) across multiple model scales, domains, and architectures. 📦 Package: egora on PyPI 💻 Code: ArsSocratica/EgoRA on GitHub 📄 Paper: arXiv:2602.05192 🔖 DOI: 10.5281/zenodo.19398709 Dataset Structure llama-3.2-1b/, llama-3.2-3b/, llama-3.1-8b/ Fine-tuning results across 3 model scales, 2 domains (Alpaca general, Medical), 4… See the full description on the dataset page: https://huggingface.co/datasets/ArsSocratica/egora-benchmarks.text-generation1K<n<10K0 likes127 downloads6mo agoHugging Face29Salesteq /ar-sa-tts-speakers-synthetic Deprecated -- consolidated This repo's data files stay in place, but the rows now live as named config(s) on the single Salesteq synthetic-speech dataset: speakers-v1 on Salesteq/ar-sa-tts-corpus-synthetic Load from there rather than this repo going forward. ar-sa-tts-speakers — multi-speaker Najdi Arabic TTS (synthetic) Ten single-speaker synthetic Najdi Arabic TTS sets unified into one dataset, distinguished by the speaker column. 91,350 clips across 10… See the full description on the dataset page: https://huggingface.co/datasets/Salesteq/ar-sa-tts-speakers-synthetic.tabular100K<n<1M0 likes121 downloads14d agoHugging Face30alea-institute /kl3m-data-dotgov-www.ars.usda.gov KL3M Data Project Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper. Description This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models. Dataset Details Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.ars.usda.gov.text10K<n<100K0 likes113 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.