CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tonyALTR /3D_native_rot_gentextn<1K0 likes5.6k downloads1y agoHugging Face02tonyALTR /3D_native_gentextn<1K0 likes2.7k downloads1y agoHugging Face03OALL /AlGhafa-Arabic-LLM-Benchmark-Native AlGhafa Arabic LLM Benchmark New fix: Normalized whitespace characters and ensured consistency across all datasets for improved data quality and compatibility. Multiple-choice evaluation benchmark for zero- and few-shot evaluation of Arabic LLMs, we adapt the following tasks: Belebele Ar MSA Bandarkar et al. (2023): 900 entries Belebele Ar Dialects Bandarkar et al. (2023): 5400 entries COPA Ar: 89 entries machine-translated from English COPA and verified by native Arabic… See the full description on the dataset page: https://huggingface.co/datasets/OALL/AlGhafa-Arabic-LLM-Benchmark-Native.text10K<n<100K7 likes2.7k downloads3y agoHugging Face04AlexCuadron /SWE-Bench-Verified-O1-native-tool-calling-reasoning-high-results SWE-Bench Verified O1 Dataset Executive Summary This repository contains verified reasoning traces from the O1 model evaluating software engineering tasks. Using OpenHands + CodeAct v2.2, we tested O1's bug-fixing capabilities using their native tool calling capabilities on the SWE-Bench Verified dataset, achieving a 45.8% success rate across 500 test instances. Overview This dataset was generated using the CodeAct framework, which aims to improve code… See the full description on the dataset page: https://huggingface.co/datasets/AlexCuadron/SWE-Bench-Verified-O1-native-tool-calling-reasoning-high-results.textquestion-answeringn<1K4 likes1.9k downloads2y agoHugging Face05Archangel-system /glaive-function-calling-v2-openai-native glaive-function-calling-v2-openai-native glaiveai/glaive-function-calling-v2 restructured into the native OpenAI / TRL format: tools is a typed column and tool_calls[].function.arguments is a real object — not JSON inside a string. The original is widely used (69k downloads/month) but inactive for ~3 years, and ships tool calls as <functioncall> text blobs with Python-quoted arguments. Existing repackagings either keep ShareGPT with tools as a string, or carry no license at all.… See the full description on the dataset page: https://huggingface.co/datasets/Archangel-system/glaive-function-calling-v2-openai-native.texttext-generation10K<n<100K1 likes622 downloads9d agoHugging Face06RickyHFR /carla-ue5-native-erp-2048x1024-5m-track-100k0 likes572 downloads3mo agoHugging Face07SultanR /AraMix-Native AraMix-Native A native-Arabic-filtered version of AdaMLLab/AraMix (minhash_deduped), derived from SultanR/AraMix-Translation-Scores: machine-translated and garbled-MT documents removed, 162,887,010 rows kept of 178,883,241 (91.06%). All columns preserved. Filter rules A document is kept iff all of: mmbert_translated_score < 0.1, or a classical-text rescue: diacritic (tashkeel) ratio ≥ 0.02 over Arabic letters and ≥ 3 distinct diacritic classes (fully/partially… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/AraMix-Native.tabular100M<n<1B0 likes532 downloads2mo agoHugging Face08DTU54DL /common-native-proc Dataset Card for [Dataset Name] Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information Needed] Source Data… See the full description on the dataset page: https://huggingface.co/datasets/DTU54DL/common-native-proc.texttoken-classification10K<n<100K0 likes508 downloads4y agoHugging Face09asaverren /native-sft native-sft A format-alignment remix, not new instruction data. Conversations come from AllenAI Dolci (ODC-By) and NVIDIA Nemotron-Post-Training-Dataset-v1 (CC BY 4.0). Each family config re-renders those chats through a real 2026 instruct template so SFT can keep native special tokens / think / tools markers. Trainers get prompt + completion, so they do not need {% generation %} in jinja. v1 2026-08-31: ~9609 canonical conversations; 57 unique-hash family configs; 539,326… See the full description on the dataset page: https://huggingface.co/datasets/asaverren/native-sft.texttext-generation100K<n<1M2 likes396 downloads22d agoHugging Face10gayanin /babylon-native-v8-noise-op-wisetext10K<n<100K0 likes379 downloads3y agoHugging Face11deepearth /central-florida-native-plants DeepEarth Central Florida Native Plants Dataset v0.2.0 🌿 Dataset Summary A comprehensive multimodal dataset featuring 33,665 observations of 232 native plant species from Central Florida. This dataset combines citizen science observations with state-of-the-art vision and language embeddings for advancing multimodal self-supervised ecological intelligence research. Key Features 🌍 Spatiotemporal Coverage: Complete GPS coordinates and timestamps for all… See the full description on the dataset page: https://huggingface.co/datasets/deepearth/central-florida-native-plants.tabularimage-classification10K<n<100K0 likes274 downloads1y agoHugging Face12gayanin /kaggle-native-v8-noise-op-wisetext10K<n<100K0 likes221 downloads3y agoHugging Face13QCRI /LlamaLens-Arabic-Native LlamaLens: Specialized Multilingual LLM Dataset Overview LlamaLens is a specialized multilingual LLM designed for analyzing news and social media content. It focuses on 18 NLP tasks, leveraging 52 datasets across Arabic, English, and Hindi. LlamaLens This repo includes scripts needed to run our full pipeline, including data preprocessing and sampling, instruction dataset creation, model fine-tuning, inference and evaluation. Features… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/LlamaLens-Arabic-Native.texttext-classification1M<n<10M0 likes213 downloads2y agoHugging Face14arcee-globe /AlGhafa-Arabic-LLM-Benchmark-Native-10percenttext1K<n<10K0 likes169 downloads2y agoHugging Face15OpenMLRL /BFCL-V4-Parallel-Native BFCL V4 Parallel Native Native BFCL v4 single-turn parallel function-calling rows for decentralized multi-agent collaboration. Source data comes from the official Berkeley Function Calling Leaderboard v4 data and possible-answer files. Fields id official_category task_type user_prompt function ground_truth Categories live_parallel live_parallel_multiple parallel parallel_multiple Counts train: 352 rows eval: 88 rows total: 440… See the full description on the dataset page: https://huggingface.co/datasets/OpenMLRL/BFCL-V4-Parallel-Native.texttext-generationn<1K1 likes156 downloads3mo agoHugging Face16kkomyoeminaung /myanmar-native-text Dataset Card for myanmar-native-text Dataset Summary ဒီ dataset က myanmar-native-text အတွက် ဖန်တီးထားတာပါ။ Languages Myanmar (my) / English (en) Dataset Structure Data Instances { "text": "နမူနာ စာသား", "label": "အညွှန်း" } Data Fields text: main content, label: optional. Data Splits Split Files train data/native_00000.jsonl, data/native_00001.jsonl, data/native_00002.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/kkomyoeminaung/myanmar-native-text.1 likes150 downloads2mo agoHugging Face17adarshxs /voxcpm2-native-generated-audio-user-ref VoxCPM2 Native Generated Audio (User Ref) Raw audio files generated from the native VoxCPM2 path in sglang-omni using a user-provided reference clip. Contents 9 generated .wav files metadata.json with prompt text, mode, status, size, and latency Source Reference Audio Reference clip used for the reference-mode generations: https://huggingface.co/datasets/adarshxs/voxcpm2-native-test-samples/resolve/main/data/audio.wav Files ref_expressive.wav… See the full description on the dataset page: https://huggingface.co/datasets/adarshxs/voxcpm2-native-generated-audio-user-ref.audiotext-to-speechn<1K0 likes142 downloads5mo agoHugging Face18djroytburg /auditbench-graft-vs-native-eval-results AuditBench — graft vs native organisms, evaluation results Numeric evaluation results for the AuditBench model-organism grid on two model families: Qwen3-14B and Llama-3.3-70B-Instruct. The organisms themselves are published separately (djroytburg/auditbench-qwen3-14b-*, djroytburg/auditbench-llama33-70b-*). The design Each cell compares three arms on the same eval, served together: arm meaning bare the untouched instruct model native SDF… See the full description on the dataset page: https://huggingface.co/datasets/djroytburg/auditbench-graft-vs-native-eval-results.0 likes118 downloads2mo agoHugging Face19sangamdas /Execution-Finality-Enforcement-for-AI-Agents-AI-Native-Telecom-Financial-Systems Identity Is Not Authority: Execution-Finality Enforcement for AI Agents, 6G, Financial, Digital, and Autonomous Systems Non-Routable, Non-Bearer Virtual Identity Bound to a Protected Compliance Jurisdiction Structure, Exact-Act Validation, LAVR, Execution Handle, and Finality Sink Verification Author: Sangam DasTechnical domain: AI security, agentic AI, trusted computing, execution governance, digital identity, 6G/telecommunications, financial infrastructure… See the full description on the dataset page: https://huggingface.co/datasets/sangamdas/Execution-Finality-Enforcement-for-AI-Agents-AI-Native-Telecom-Financial-Systems.documentn<1K0 likes118 downloads1mo agoHugging Face20AlazarM /hey-native-wakeword Hey Native — wake-word dataset Synthetic 16 kHz mono audio for training a "Hey Native" wake-word detector. data/positive/ — utterances of "Hey Native" (label 1) data/negative/ — general speech, not the wake word (label 0) data/hard_negative/ — near-miss confusables, e.g. "hey navy", "hey maybe" (label 0) metadata.csv — file_name, label, label_name, text, source_model, mos_p808 5,000 positives · 6,000 negatives · 1,050 hard negatives. audioaudio-classification10K<n<100K0 likes114 downloads4d agoHugging Face21Praxel /psp-native-centroids Praxel/psp-native-centroids Native-speaker reference artefacts for the PSP (Phoneme Substitution Profile) benchmark for Indic text-to-speech accent evaluation. Companion to the paper PSP: An Interpretable Per-Dimension Accent Benchmark for Indic Text-to-Speech (Menta, 2026). This dataset is a scoring reference, not a training corpus. It contains pre-computed acoustic references extracted from publicly-licensed native-speaker speech corpora, used by the psp-eval package to score TTS… See the full description on the dataset page: https://huggingface.co/datasets/Praxel/psp-native-centroids.text-to-speech1K<n<10K0 likes106 downloads5mo agoHugging Face22dreamdifferent /vam-cross-target-widowx250-native-corner-frontThis dataset was created using LeRobot. Dataset Description 30.10 minutes (160 episodes, 160 marked successful) of the VAM-Cross target collection for widowx250 in MuJoCo at 30 Hz using the widowx-texture robot appearance. Observations include RGB, joint state, gripper state, and achieved EE pose; actions include the full commanded EE pose and gripper command. Task assets are derived from the MolmoSpaces THOR 20251117 asset release (MolmoSpaces commit… See the full description on the dataset page: https://huggingface.co/datasets/dreamdifferent/vam-cross-target-widowx250-native-corner-front.tabularrobotics10K<n<100K0 likes103 downloads7d agoHugging Face23QCRI /LlamaLens-Hindi-Native LlamaLens: Specialized Multilingual LLM Dataset Overview LlamaLens is a specialized multilingual LLM designed for analyzing news and social media content. It focuses on 18 NLP tasks, leveraging 52 datasets across Arabic, English, and Hindi. LlamaLens This repo includes scripts needed to run our full pipeline, including data preprocessing and sampling, instruction dataset creation, model fine-tuning, inference and evaluation. Features… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/LlamaLens-Hindi-Native.texttext-classification100K<n<1M0 likes101 downloads2y agoHugging Face24rs545837 /entity-native-agent-sessions Entity-Native vs File-Native Agent Sessions on SWE-bench Verified Full session logs from a controlled A/B experiment measuring how a coding agent's retrieval substrate changes its behaviour, cost, and success rate on real software-engineering tasks. Both arms run the same model (Claude Sonnet 4.5), on the same tasks, from the same repository state. The only difference is how the agent is allowed to find code. Arm Label Tools available A file-native Bash, Read, Grep… See the full description on the dataset page: https://huggingface.co/datasets/rs545837/entity-native-agent-sessions.tabulartext-generationn<1K0 likes99 downloads20d agoHugging Face25gayanin /gcd-native-v8-noise-op-wisetext1K<n<10K0 likes97 downloads3y agoHugging Face26open-llm-leaderboard-old /details_chavinlo__alpaca-native Dataset Card for Evaluation run of chavinlo/alpaca-native Dataset Summary Dataset automatically created during the evaluation run of model chavinlo/alpaca-native on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_chavinlo__alpaca-native.0 likes89 downloads3y agoHugging Face27Archangel-system /medmcqa-openai-native MedMCQA — OpenAI-native, with a usable test split MedMCQA is one of the most downloaded medical QA datasets on the Hub. Its test split has been unusable since release: all 6,150 rows carry cop=-1 (no label) and an empty explanation. You cannot score a model on it. This release rebuilds a labelled, leak-free test split and converts everything to the native messages format, so it loads straight into TRL with no custom parsing. What was actually wrong Measured on the… See the full description on the dataset page: https://huggingface.co/datasets/Archangel-system/medmcqa-openai-native.textquestion-answering100K<n<1M0 likes89 downloads7d agoHugging Face28webis /generative-native-ads Webis Generated Native Ads 2024 Dataset Summary This dataset was created to train ad blocking systems on the task of identifying advertisements in responses of conversational search engines. There are two dataset dictionaries available: responses.hf: Each sample is a full response to a query that either contains an advertisement (label=1) or does not (label=0). sentence_pairs.hf: Each sample is a pair of two sentences taken from the responses. If one of them… See the full description on the dataset page: https://huggingface.co/datasets/webis/generative-native-ads.text-classification1 likes87 downloads2mo agoHugging Face29schneiderkamplab /dfm11-toolace-native-tool-use-repaired dfm11-toolace-native-tool-use-repaired ToolACE conversations with declared-name parsing and complete parallel result binding. This is a DFM11 replacement for schneiderkamplab/dfm10-toolace-native-tool-use. All rows pass exhaustive structural validation. See metadata/manifest.json. text10K<n<100K0 likes83 downloads18d agoHugging Face30deepearth /central-florida-native-plants-language-embeddings Central Florida Native Plants Language Embeddings This dataset contains language embeddings for 232 native plant species from Central Florida, extracted using the DeepSeek-V3 language model. Dataset Summary This dataset provides pre-computed language embeddings for Central Florida plant species. Each species has been encoded using the prompt "Ecophysiology of {species_name}:" to capture semantic information about the plant's ecological characteristics. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/deepearth/central-florida-native-plants-language-embeddings.tabularfeature-extraction1K<n<10K0 likes73 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.