CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PierreColombo /miamMultilingual dIalogAct benchMark is a collection of resources for training, evaluating, and analyzing natural language understanding systems specifically designed for spoken language. Datasets are in English, French, German, Italian and Spanish. They cover a variety of domains including spontaneous speech, scripted scenarios, and joint task completion. Some datasets additionally include emotion and/or sentimant labels.text-generation10K<n<100K6 likes172 downloads3y agoHugging Face02Miaowuawa /ChineseNovels 中文小说数据集 包含内容: 网游/系统/重生 言情小说 同人/耽美小说 科幻小说 军事小说 以上加起来共4万本左右 海棠文学城小说:约1000本(未清洗) texttext-generation1K<n<10K22 likes82 downloads2y agoHugging Face03edikka-lab /mia-fr-prompts MIA-FR: French prompts for observing AI visibility 25 French prompts · 11 themes · 5 intent categories · Corpus v1.0 · CC BY 4.0 Lire en français · Field dictionary · Method and limitations · Source protocol MIA-FR is a small, fixed prompt corpus published by Bertrand Morel / Edikka to support repeated observation of brand mentions and source citations in AI search interfaces. It gives practitioners a documented starting point for a French-language collection workflow. This… See the full description on the dataset page: https://huggingface.co/datasets/edikka-lab/mia-fr-prompts.texttext-generationn<1K0 likes53 downloads13d agoHugging Face04minhthien /mia-meeting MIA Meeting E2E Dataset Synthetic meeting dataset for end-to-end experiments: audio to transcript transcript plus roster to action items action item extraction benchmark Splits train: 200 samples, 0 with linked audio validation: 5 samples, 5 with linked audio eval: 205 samples, 5 with linked audio Structure data/*.jsonl # split manifests audio/<split>/* # linked audio files when available transcripts/<split>/*.json #… See the full description on the dataset page: https://huggingface.co/datasets/minhthien/mia-meeting.audioautomatic-speech-recognition0 likes33 downloads4mo agoHugging Face05Miaow-Lab /RUT-Bench Beyond Ideal Instruction: A Comprehensive Framework for Evaluating LLMs in Realistic Interactions This repository contains the RUT-Bench benchmark, which consists of 1638 test samples for evaluating LLM agents under realistic user interactions. Paper: Beyond Ideal Instruction: A Comprehensive Framework for Evaluating LLMs in Realistic Interactions Code: GitHub Collection: Hugging Face Collection 📖 Overview RUT-Bench is a dedicated benchmark designed to assess… See the full description on the dataset page: https://huggingface.co/datasets/Miaow-Lab/RUT-Bench.tabulartext-generation1K<n<10K1 likes32 downloads4mo agoHugging Face06miaomiao64 /tb-explore17-mcode-m3-harness-variancegated Terminal-Bench 2.1 explore-17 — mcode / MiniMax-M3 harness variance Three complete 17-task runs of the same dataset ref with the same agent and model, differing only in execution substrate and concurrency, plus one isolated rerun. The point of the bundle is not the resolve rate — it is how much the resolve rate moves when nothing about the task or the model changes. Same everywhere: dataset ai-solution-finetune/terminal-bench-2-1-explore-17 at… See the full description on the dataset page: https://huggingface.co/datasets/miaomiao64/tb-explore17-mcode-m3-harness-variance.tabulartext-generationn<1K0 likes31 downloads14d agoHugging Face07matthewwicker /shadow-llm-mia-signals Shadow LLM MIA Signals (OLMo-2-1B) Membership Inference Attack (MIA) signal tensors extracted from 128 shadow models fine-tuned from allenai/OLMo-2-0425-1B. Overview This dataset enables research on membership inference attacks against large language models. Each of 128 shadow models was trained on a different random subset of 64 out of 128 candidate documents from the OLMo-mix-1124 pretraining dataset. For each (model, document) pair, we extracted softmax prediction… See the full description on the dataset page: https://huggingface.co/datasets/matthewwicker/shadow-llm-mia-signals.tabulartext-generationn<1K0 likes25 downloads5mo agoHugging Face08Miaow-Lab /SSAE-Dataset Dataset Card This is the official dataset repository for the paper "Step-Level Sparse Autoencoder for Reasoning Process Interpretation". Paper: Arxiv Code: GitHub Collection: HuggingFace Dataset Overview The repository hosts three distinct datasets covering the domains of mathematical reasoning and code generation. Each subset is pre-partitioned into training and validation splits to facilitate reproducible experiments. 1. GSM8K (Math) Description: A… See the full description on the dataset page: https://huggingface.co/datasets/Miaow-Lab/SSAE-Dataset.texttext-generation1M<n<10M0 likes24 downloads7mo agoHugging Face09miaomiao64 /swe-explore-find-dev100-runsgated SWE-Explore find — dev-100 evidence runs Five complete 100-trial runs of the find (fault-localization) arm of SWE-Explore, kept because each one is load-bearing evidence for a specific claim about the scoring fixes on branch fix/swe-explore-find-scoring of harness_bench. Every run is 100 trials of the same dev-100 find subset, run through Harbor with the swe_explore.agents:PiSut agent. index.jsonl has one row per trial (500 rows); the full raw Harbor trial directories are in… See the full description on the dataset page: https://huggingface.co/datasets/miaomiao64/swe-explore-find-dev100-runs.tabulartext-generationn<1K0 likes24 downloads15d agoHugging Face10Lots-of-LoRAs /task896_miam_language_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task896_miam_language_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task896_miam_language_classification.texttext-generation1K<n<10K0 likes15 downloads2y agoHugging Face11miazaitman /CheatCleantabulartext-generation1K<n<10K0 likes13 downloads1y agoHugging Face12miazaitman /nutriswap-healthy-food-alternativestexttext-classification1K<n<10K0 likes2 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.