CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AlexCuadron /SWE-Bench-Verified-O1-native-tool-calling-reasoning-high-results SWE-Bench Verified O1 Dataset Executive Summary This repository contains verified reasoning traces from the O1 model evaluating software engineering tasks. Using OpenHands + CodeAct v2.2, we tested O1's bug-fixing capabilities using their native tool calling capabilities on the SWE-Bench Verified dataset, achieving a 45.8% success rate across 500 test instances. Overview This dataset was generated using the CodeAct framework, which aims to improve code… See the full description on the dataset page: https://huggingface.co/datasets/AlexCuadron/SWE-Bench-Verified-O1-native-tool-calling-reasoning-high-results.textquestion-answeringn<1K4 likes1.7k downloads2y agoHugging Face02OALL /AlGhafa-Arabic-LLM-Benchmark-Native AlGhafa Arabic LLM Benchmark New fix: Normalized whitespace characters and ensured consistency across all datasets for improved data quality and compatibility. Multiple-choice evaluation benchmark for zero- and few-shot evaluation of Arabic LLMs, we adapt the following tasks: Belebele Ar MSA Bandarkar et al. (2023): 900 entries Belebele Ar Dialects Bandarkar et al. (2023): 5400 entries COPA Ar: 89 entries machine-translated from English COPA and verified by native Arabic… See the full description on the dataset page: https://huggingface.co/datasets/OALL/AlGhafa-Arabic-LLM-Benchmark-Native.text10K<n<100K7 likes1.6k downloads3y agoHugging Face03Archangel-system /glaive-function-calling-v2-openai-native glaive-function-calling-v2-openai-native glaiveai/glaive-function-calling-v2 restructured into the native OpenAI / TRL format: tools is a typed column and tool_calls[].function.arguments is a real object — not JSON inside a string. The original is widely used (69k downloads/month) but inactive for ~3 years, and ships tool calls as <functioncall> text blobs with Python-quoted arguments. Existing repackagings either keep ShareGPT with tools as a string, or carry no license at all.… See the full description on the dataset page: https://huggingface.co/datasets/Archangel-system/glaive-function-calling-v2-openai-native.texttext-generation10K<n<100K1 likes713 downloads13d agoHugging Face04SultanR /AraMix-Native AraMix-Native A native-Arabic-filtered version of AdaMLLab/AraMix (minhash_deduped), derived from SultanR/AraMix-Translation-Scores: machine-translated and garbled-MT documents removed, 162,887,010 rows kept of 178,883,241 (91.06%). All columns preserved. Filter rules A document is kept iff all of: mmbert_translated_score < 0.1, or a classical-text rescue: diacritic (tashkeel) ratio ≥ 0.02 over Arabic letters and ≥ 3 distinct diacritic classes (fully/partially… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/AraMix-Native.tabular100M<n<1B0 likes499 downloads2mo agoHugging Face05DTU54DL /common-native-proc Dataset Card for [Dataset Name] Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information Needed] Source Data… See the full description on the dataset page: https://huggingface.co/datasets/DTU54DL/common-native-proc.texttoken-classification10K<n<100K0 likes496 downloads4y agoHugging Face06asaverren /native-sft native-sft A format-alignment remix, not new instruction data. Conversations come from AllenAI Dolci (ODC-By) and NVIDIA Nemotron-Post-Training-Dataset-v1 (CC BY 4.0). Each family config re-renders those chats through a real 2026 instruct template so SFT can keep native special tokens / think / tools markers. Trainers get prompt + completion, so they do not need {% generation %} in jinja. v1 2026-08-31: ~9609 canonical conversations; 57 unique-hash family configs; 539,326… See the full description on the dataset page: https://huggingface.co/datasets/asaverren/native-sft.texttext-generation100K<n<1M2 likes449 downloads26d agoHugging Face07gayanin /babylon-native-v8-noise-op-wisetext10K<n<100K0 likes409 downloads3y agoHugging Face08deepearth /central-florida-native-plants DeepEarth Central Florida Native Plants Dataset v0.2.0 🌿 Dataset Summary A comprehensive multimodal dataset featuring 33,665 observations of 232 native plant species from Central Florida. This dataset combines citizen science observations with state-of-the-art vision and language embeddings for advancing multimodal self-supervised ecological intelligence research. Key Features 🌍 Spatiotemporal Coverage: Complete GPS coordinates and timestamps for all… See the full description on the dataset page: https://huggingface.co/datasets/deepearth/central-florida-native-plants.tabularimage-classification10K<n<100K0 likes303 downloads1y agoHugging Face09gayanin /kaggle-native-v8-noise-op-wisetext10K<n<100K0 likes243 downloads3y agoHugging Face10arcee-globe /AlGhafa-Arabic-LLM-Benchmark-Native-10percenttext1K<n<10K0 likes159 downloads2y agoHugging Face11gayanin /gcd-native-v8-noise-op-wisetext1K<n<10K0 likes98 downloads3y agoHugging Face12Archangel-system /medmcqa-openai-native MedMCQA — OpenAI-native, with a usable test split MedMCQA is one of the most downloaded medical QA datasets on the Hub. Its test split has been unusable since release: all 6,150 rows carry cop=-1 (no label) and an empty explanation. You cannot score a model on it. This release rebuilds a labelled, leak-free test split and converts everything to the native messages format, so it loads straight into TRL with no custom parsing. What was actually wrong Measured on the… See the full description on the dataset page: https://huggingface.co/datasets/Archangel-system/medmcqa-openai-native.textquestion-answering100K<n<1M0 likes90 downloads11d agoHugging Face13Archangel-system /oasst2-openai-native oasst2-openai-native A deterministic, native OpenAI/TRL reconstruction of OpenAssistant/oasst2. It turns the original flat parent_id message table into two directly usable configs without LLM transformation: multilingual SFT conversations and ranked DPO preference pairs. At a glance Config Train Test Unit sft 12,717 671 alternating conversation ending in assistant dpo 42,639 2,284 prompt + chosen/rejected assistant pair The data is multilingual:… See the full description on the dataset page: https://huggingface.co/datasets/Archangel-system/oasst2-openai-native.texttext-generation10K<n<100K0 likes82 downloads13d agoHugging Face14yacdev /somali-100k-native-conversations 🇸🇴 Somali High-Diversity Multi-Turn Conversational SFT Dataset A state-of-the-art, 100.00% unique (zero duplicate responses) multi-turn conversational dataset in authentic Somali (Af-Soomaali) across 25 real-world knowledge domains. 🌟 Quality Standards: 100% Unique Assistant Responses: Guaranteed zero template repetition (23,334 / 23,334 unique turns). Grounded Knowledge: Spanning Python coding, web dev, Git/Linux, cybersecurity, diabetes & health, business… See the full description on the dataset page: https://huggingface.co/datasets/yacdev/somali-100k-native-conversations.texttext-generation10K<n<100K0 likes77 downloads24d agoHugging Face15deepearth /central-florida-native-plants-language-embeddings Central Florida Native Plants Language Embeddings This dataset contains language embeddings for 232 native plant species from Central Florida, extracted using the DeepSeek-V3 language model. Dataset Summary This dataset provides pre-computed language embeddings for Central Florida plant species. Each species has been encoded using the prompt "Ecophysiology of {species_name}:" to capture semantic information about the plant's ecological characteristics. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/deepearth/central-florida-native-plants-language-embeddings.tabularfeature-extraction1K<n<10K0 likes74 downloads1y agoHugging Face16tiiuae /NativeQA 3LM Native STEM Arabic Benchmark Dataset Summary The 3LM Native STEM dataset contains 865 multiple-choice questions (MCQs) curated from real Arabic educational sources. It targets mid- to high-school level content in Biology, Chemistry, Physics, Mathematics, and Geography. This benchmark is designed to evaluate Arabic large language models on structured, domain-specific knowledge. Motivation While Arabic NLP has seen growth in cultural and linguistic tasks… See the full description on the dataset page: https://huggingface.co/datasets/tiiuae/NativeQA.textn<1K2 likes71 downloads7mo agoHugging Face17Archangel-system /hh-rlhf-dpo-native hh-rlhf-dpo-native Anthropic/hh-rlhf in the native TRL conversational-preference format, with zero-gradient and unparsable pairs removed. The original dataset ships two raw strings (chosen, rejected) containing the entire conversation serialized with \n\nHuman: / \n\nAssistant: separators. Every user has to write their own parser, and that parser has to make a judgement call on ~2% of rows that are corrupted. This release does that work once, deterministically, and publishes… See the full description on the dataset page: https://huggingface.co/datasets/Archangel-system/hh-rlhf-dpo-native.texttext-generation100K<n<1M0 likes65 downloads11d agoHugging Face18Archangel-system /helpsteer2-preference-openai-native HelpSteer2 Preference — OpenAI Native Format A deterministic, training-ready repackaging of the preference split of nvidia/HelpSteer2. Why use this What it is for. Preference optimisation — DPO, ORPO, SimPO, KTO — and reward modelling, on 7,051 pairs that come from paid human annotators, not from an LLM judge. Each pair carries a graded strength from 1 to 3 rather than a bare binary label, so you can weight the loss by how strongly humans actually disagreed, or… See the full description on the dataset page: https://huggingface.co/datasets/Archangel-system/helpsteer2-preference-openai-native.textreinforcement-learning1K<n<10K0 likes64 downloads13d agoHugging Face19Archangel-system /codealpaca-openai-native CodeAlpaca OpenAI Native This is a deterministic, lossless-formatting derivative of sahil2801/CodeAlpaca-20k, modernized with a typed OpenAI/TRL messages column and decontaminated against the HumanEval and MBPP test sets. The original Alpaca columns remain available for backward compatibility. Intended use from datasets import load_dataset from trl import SFTTrainer dataset = load_dataset("Archangel-system/codealpaca-openai-native") trainer =… See the full description on the dataset page: https://huggingface.co/datasets/Archangel-system/codealpaca-openai-native.texttext-generation10K<n<100K0 likes64 downloads11d agoHugging Face20tiiuae /NativeQA-RDP 3LM Native STEM Arabic Benchmark - RDP version Dataset Summary The 3LM Native STEM dataset contains 865 multiple-choice questions (MCQs) curated from real Arabic educational sources. It targets mid- to high-school level content in Biology, Chemistry, Physics, Mathematics, and Geography. This benchmark is designed to evaluate Arabic large language models on structured, domain-specific knowledge. In this "RDP - Robustness under Distractor Perturbation" version, 25% of… See the full description on the dataset page: https://huggingface.co/datasets/tiiuae/NativeQA-RDP.textn<1K0 likes59 downloads7mo agoHugging Face21gayanin /babylon-native-v8-vocab-noisedtext1K<n<10K0 likes58 downloads3y agoHugging Face22NLP-FBK /e3c-sentences-EU-nativetext1K<n<10K0 likes56 downloads2y agoHugging Face23werty1248 /s1.1-Ko-Native-resulttext1K<n<10K0 likes55 downloads2y agoHugging Face24nativemind /kene_multimodal_gift Kene Multimodal Gift Dataset (Enhanced with Ethnic Languages) Описание Мультимодальный духовный датасет с ИКАРОС на испанском, Джив Джаго на хинди и языками народностей России, СНГ и Украины. Обновления ✅ Добавлены ИКАРОС на испанском языке ✅ Добавлен Джив Джаго на хинди ✅ НОВОЕ: Добавлены языки народностей России, СНГ и Украины ✅ Улучшены мультимодальные данные ✅ Расширена поддержка языков до 50 примеров Языки Духовные языки Русский:… See the full description on the dataset page: https://huggingface.co/datasets/nativemind/kene_multimodal_gift.audion<1K1 likes55 downloads11mo agoHugging Face25werty1248 /qwen-s1.1-Ko-Native-resulttext1K<n<10K0 likes48 downloads2y agoHugging Face26NLP-FBK /e3c-sentences-IT-nativetext1K<n<10K0 likes43 downloads2y agoHugging Face27gayanin /babylon-native-mixedtext10K<n<100K0 likes38 downloads3y agoHugging Face28dmedhi /common-native-en-bpeaudio10K<n<100K0 likes38 downloads1y agoHugging Face29praneethv35 /react-native-codetext10K<n<100K1 likes28 downloads7mo agoHugging Face30DTU54DL /common-native Dataset Card for [Dataset Name] Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information Needed] Source Data… See the full description on the dataset page: https://huggingface.co/datasets/DTU54DL/common-native.audiotoken-classification10K<n<100K0 likes26 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.