CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01arshiaafshani /persian-natural-fluently Persian scientific dataset I have prepared a great and natural persian dataset of scientific datas including chemistry, physics, mathematics (including algebra & etc) , biology. The content of the dataset has been generated by : Human, Grok3, DeepSeek R1. License This dataset is licensed under apache-2.0. texttext-generationn<1K14 likes41 downloads1y agoHugging Face02Arsh9210 /Nemotron-RL-litmus-bench-v0.1 Dataset Description: Litmus-Bench v0.1 is an open dataset for training and evaluating chemical reasoning in language models. It includes 5,232 training questions and 482 test questions, each in short-answer format and was created from the ChEMBL dataset with RDKit descriptors requiring short answers. The dataset is for RL training. This dataset is released as part of NVIDIA NeMo-Gym, an open-source library within the NVIDIA NeMo framework, designed for large-scale, verifiable… See the full description on the dataset page: https://huggingface.co/datasets/Arsh9210/Nemotron-RL-litmus-bench-v0.1.textreinforcement-learning1K<n<10K0 likes40 downloads2mo agoHugging Face03ArsParadox /Viel-v1.0 This is Viel An Industrial Robot turned into a robot assistant. A bit rough around the edges and would rather enter sleep mode than helping you with your inane request. This dataset contains the entire 52k Alpaca Dataset modified to mimic Viel's personality and speaking pattern and a hint of her background. . . . Have fun then . . . Finetune coming soon, just putting her dataset here for easier Collab Access textquestion-answering10K<n<100K0 likes20 downloads2y agoHugging Face04ArSyra /arsyra-questionsgated 🧠 ArSyra Question Bank — 25K Smart Arabic Dialect Prompts Comprehensive question bank for Arabic dialect data collection and NLP training. Dataset Summary A collection of 25,000 linguistically diverse questions designed for Arabic dialect data collection. Each question includes MSA source text, Arabic prompts, category metadata, difficulty level (beginner/intermediate/advanced), and topic tags. Expertly crafted prompts across 13 categories: dialect translation… See the full description on the dataset page: https://huggingface.co/datasets/ArSyra/arsyra-questions.text-generation10K<n<100K0 likes4 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.