CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Human-Centric-Machine-Learning /tokenization-multiplicity-data Dataset: Tokenization Multiplicity Leads to Arbitrary Price Variation in LLM-as-a-service This dataset contains the official experiment inference traces for the paper Tokenization Multiplicity Leads to Arbitrary Price Variation in LLM-as-a-service by Ivi Chatzi, Nina Corvelo Benz, Stratis Tsirtsis and Manuel Gomez-Rodriguez. 📂 Dataset Structure The dataset is organized into folders as follows: .\{model}\{task}\{lang}\{seed}_{10*temperature}.jsonl where {model}… See the full description on the dataset page: https://huggingface.co/datasets/Human-Centric-Machine-Learning/tokenization-multiplicity-data.text-generation10K<n<100K2 likes804 downloads6mo agoHugging Face02Human-Centric-Machine-Learning /strategic-ttc-data Dataset: Strategic Test-Time Compute (TTC) This dataset contains the official experiment inference traces for the paper "Test-Time Compute Games" (arXiv:2601.21839). It includes full model generations, token counts, and correctness verifications for various Large Language Models (LLMs) across three major reasoning benchmarks: GSM8K, AIME, and GPQA. This data allows researchers to analyze the relationship between test-time compute and model performance without needing to re-run… See the full description on the dataset page: https://huggingface.co/datasets/Human-Centric-Machine-Learning/strategic-ttc-data.question-answering10K<n<100K2 likes245 downloads8mo agoHugging Face03freococo /myanmar_quran_parallel_dataset_human_vs_ai Myanmar Quran Parallel Dataset: Human vs AI This dataset is a comprehensive multi-parallel corpus of the Holy Qur'an, containing all 6,236 verses. It is designed as a high-quality linguistic resource for evaluating and aligning AI systems on formal, literary, and modern Myanmar (Burmese) language in a religious context. Each verse aligns the original Uthmani Arabic text with trusted human translations and multiple AI-generated translations, enabling fine-grained comparison between… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_quran_parallel_dataset_human_vs_ai.texttranslation1K<n<10K0 likes53 downloads8mo agoHugging Face04DigiGreen /human_curated_qa_dataset Human Curated QA Dataset DigiGreen/human_curated_qa_dataset is a human-verified question-answer dataset designed to support research and development in natural language question answering and agriculture-focused conversational AI. This dataset contains realistic, domain-relevant QA pairs that were manually curated to ensure accurate and contextually rich answers. It can be used to benchmark models for QA generation. 📌 Dataset Overview Name: Human Curated QA Dataset… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/human_curated_qa_dataset.texttext-generation1K<n<10K1 likes40 downloads5mo agoHugging Face05DataOrigin /human-feedback-mentor-sessions-india license: other task_categories: audio-classification automatic-speech-recognition text-generation language: en hi bn ta te ml mr or as pa pretty_name: Human Feedback Mentor Sessions India size_categories: 1K<n<10K Human Feedback Mentor Sessions India Dataset Description A rare and high-value collection of recorded aspirant-mentor interactions capturing real guidance, feedback, and reasoning corrections in the context of Indian government exam preparation. Produced by… See the full description on the dataset page: https://huggingface.co/datasets/DataOrigin/human-feedback-mentor-sessions-india.videoaudio-classificationn<1K0 likes4 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.