CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kaysss /leetcode-problem-detailed LeetCode Scraper Dataset This dataset contains information scraped from LeetCode, including problem details, metadata, and related files. It is designed to assist developers in analyzing LeetCode problems, generating insights, and building tools for competitive programming or educational purposes. questions_deets.csv Contains detailed information about each problem, including problem descriptions, constraints, and examples. Columns: questionFrontendId: Unique problem ID.… See the full description on the dataset page: https://huggingface.co/datasets/kaysss/leetcode-problem-detailed.tabulartext-classification1K<n<10K10 likes669 downloads1y agoHugging Face02open-llm-leaderboard-old /details_Ramikan-BR__tinyllama_PY-CODER-4bit-lora_4k-v12 Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Ramikan-BR__tinyllama_PY-CODER-4bit-lora_4k-v12.text-generation10K<n<100K0 likes369 downloads2y agoHugging Face03d0rj /povarenok_recipes_detail povarenok_recipes_detail Crawled detailed recipes from povarenok.ru website. Structure WIP imagetext-classification100K<n<1M3 likes102 downloads3y agoHugging Face04code-critic-model /critic-sft-cwm-only-detailed-prompt critic-sft-cwm-only-detailed-prompt The detailed-prompt SFT corpus from Steer, Don't Solve: Training Small Critic Models for Large Code Agents. It trains Qwen3-8B-Critic-SFT-Detailed-Prompt, the comparison arm of the prompt ablation in Table 4. Each record is one critique point: a CWM-32B trajectory up to some step, followed by the critique that Claude Opus 4.6 wrote for it. The difference from critic-sft-cwm-only is the teacher prompt. Here the teacher used the detailed prompt… See the full description on the dataset page: https://huggingface.co/datasets/code-critic-model/critic-sft-cwm-only-detailed-prompt.texttext-generation1K<n<10K0 likes69 downloads22d agoHugging Face05YangL1122 /DetailBench DetailBench A Benchmark for Detail Hallucination in Long Regulatory Documents DetailBench is a benchmark for evaluating and mitigating detail hallucination in LLM outputs on long regulatory documents. Overview Large language models frequently produce detail hallucinations—subtle errors in threshold values, units, scopes, obligation levels, and conditions—when processing long regulatory documents. DetailBench provides: 322 source documents (172 real + 150 synthetic) from… See the full description on the dataset page: https://huggingface.co/datasets/YangL1122/DetailBench.text-generation10K<n<100K0 likes55 downloads7mo agoHugging Face06leafspark /DetailedReflection-Claude-v3_5-Sonnet DetailedReflection-Claude-v3_5-Sonnet This is a reflection dataset inspired by OpenAI's o1 model. It contains filtered prompts from anthracite-org/kalo-opus-instruct-22k-no-refusal. The data was generated by Claude 3.5 Sonnet with a custom system prompt and sampling parameters on OpenRouter. texttext-generationn<1K5 likes45 downloads2y agoHugging Face07TTS-AGI /balanced-emotion-dataset-majestrino-withtemporal-detailed-captions Balanced Emotion Dataset — Majestrino with Temporal Detailed Captions An emotion-balanced subset of TTS-AGI/majestrino-unified-detailed-captions-temporal. Overview Total samples: 482,594 Samples per emotion category: 12,997 Number of emotion categories: 40 Format: WebDataset (tar files with FLAC audio + JSON metadata) Number of tar files: 483 Samples per tar: ~1000 Balancing Strategy Samples were selected from the source dataset using keyword matching on… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/balanced-emotion-dataset-majestrino-withtemporal-detailed-captions.audioaudio-classification100K<n<1M0 likes35 downloads6mo agoHugging Face08xeophon /detailbench DetailBench This is the dataset for DetailBench, which answers the question: "How good are current LLMs at finding small errors, when they are not explicitly asked to do so?" Dataset Structure article_title: Name of the Wikipedia article the data is from original_text: Original excerpt from the given Wikipedia article modified_text: Modified version of the original text with a single error (one changed number) introduced original_number: The original number from the text… See the full description on the dataset page: https://huggingface.co/datasets/xeophon/detailbench.texttext-generationn<1K2 likes13 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.