CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Lots-of-LoRAs /task093_conala_normalize_lists Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task093_conala_normalize_lists Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task093_conala_normalize_lists.texttext-generation1K<n<10K0 likes157 downloads2y agoHugging Face02jajostrains /Mathlib-Normalized-Sexpr Mathlib Normalized S-Expressions Lean 4 proof states from Mathlib, paired with the tactic applied at each step, in three representations extracted directly from the Lean kernel: Source-faithful S-expressions of the goal and every hypothesis, as Lean elaborated them. Normalized S-expressions of the same state, with stable local-context indices suitable for model input. Annotated tactic syntax -- the original tactic's syntax tree with identifier leaves resolved to the constants… See the full description on the dataset page: https://huggingface.co/datasets/jajostrains/Mathlib-Normalized-Sexpr.tabulartext-generation100K<n<1M0 likes85 downloads29d agoHugging Face03whoisandy /router-chat-normalized-1m Router Chat Normalized 1M Dataset Description Router Chat Normalized 1M is a multilingual conversational dataset containing 1,353,300 conversations normalized from multiple public chat datasets with automatic language detection. Dataset Structure The dataset contains 2 split(s): train, test. Each conversation includes: conversation_id, messages (list of {role, content} structs), source dataset, detected language, and language confidence score. Source… See the full description on the dataset page: https://huggingface.co/datasets/whoisandy/router-chat-normalized-1m.texttext-generation1M<n<10M0 likes74 downloads4mo agoHugging Face04userdavek /Amharic_news_Normalized Dataset Name Amharic news dataset Dataset Details It is a non-normalized version of news dataset crawled from Amharic news websites and from researchers provided in their works. Dataset Description The dataset is collected from different news websites and from different researchers crawled Amharic news dataset from different NLP downstream tasks. News sites like FanaBC, EthiopianReporter, Zehabesha,Esat Amharic, BBC Amharic are the sources for these news data.… See the full description on the dataset page: https://huggingface.co/datasets/userdavek/Amharic_news_Normalized.textsummarization100K<n<1M0 likes26 downloads11mo agoHugging Face05kiarashrzg /TinyPersianStories_normalizedtexttext-generation100K<n<1M0 likes25 downloads2y agoHugging Face06tuandunghcmut /travelplanner-benchmark-normalized TravelPlanner Benchmark (Normalized) Normalized, typed, parquet-first packaging of the TravelPlanner benchmark for planning-centric agent evaluation. Upstream dataset: osunlp/TravelPlanner Upstream code: OSU-NLP-Group/TravelPlanner Paper: TravelPlanner: A Benchmark for Real-World Planning with Language Agents 1) What is included This dataset repo contains: benchmark config (train/validation/test) in typed parquet. reference_entries config: flattened reference-info… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/travelplanner-benchmark-normalized.tabulartext-generation10K<n<100K0 likes21 downloads7mo agoHugging Face070x7o /dostoevsky_frontier_3k_normalized dostoevsky_frontier_3k_normalized Нормализованная версия 0x7o/dostoevsky_frontier_3k. Нормализация Устранены пунктуационные shortcut-ы, позволяющие модели различать chosen/rejected по артефактам форматирования вместо стиля. Общие (chosen + rejected) ё → е по словарю (книги не используют ё, AI всегда использует — 87.8% accuracy) \xa0 (неразрывный пробел) → обычный пробел … (U+2026) → ... (три точки) – (en dash) → — (em dash) !.. → !..., ?.. → ?...… See the full description on the dataset page: https://huggingface.co/datasets/0x7o/dostoevsky_frontier_3k_normalized.texttext-generation1K<n<10K1 likes19 downloads7mo agoHugging Face080x7o /dostoevsky_frontier_v2_normalized dostoevsky_frontier_v2_normalized Нормализованная версия 0x7o/dostoevsky_frontier_v2. Нормализация Устранены пунктуационные shortcut-ы, позволяющие модели различать chosen/rejected по артефактам форматирования вместо стиля. Общие (chosen + rejected) ё → е по словарю (книги не используют ё, AI всегда использует — 87.8% accuracy) \xa0 (неразрывный пробел) → обычный пробел … (U+2026) → ... (три точки) – (en dash) → — (em dash) !.. → !..., ?.. → ?...… See the full description on the dataset page: https://huggingface.co/datasets/0x7o/dostoevsky_frontier_v2_normalized.texttext-generationn<1K0 likes15 downloads7mo agoHugging Face09TilQazyna /til-kk-normalize-v1gated til-kk-normalize-v1 Қазақша мәтінді қалыпқа келтіру · Нормализация казахского текста · Kazakh text normalization Қазақша · Русский · English Қазақша til-kk-normalize-v1 — қате, регистрі мен тыныс белгілері бұзылған мәтінді түзетуге арналған қазақ тіліндегі instruction-датасет. Көлемі — 5.0 МБ, жалпы саны — 10585 мысал. Деректер instruction fine-tune мен тиісті тапсырманы зерттеуге жарайды. Құрамы мен форматы Бөлік Мысал саны train 10375… See the full description on the dataset page: https://huggingface.co/datasets/TilQazyna/til-kk-normalize-v1.texttext-generation10K<n<100K0 likes13 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.