CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01samchain /bis_central_bank_speeches Dataset Card This dataset consists of central bankers speeches from 1997 to 2025 scrapped automatically from the Bank Of International Settlements website. Each speech is associated to a central bank, a date and a description (a metadata provided by the website). The dataset covers a wide range of topics in economics from monetary policies to world outlooks, financial stability, unemployment, fiscal policies... Credits Full credits to the Bank of International… See the full description on the dataset page: https://huggingface.co/datasets/samchain/bis_central_bank_speeches.tabulartext-generation10K<n<100K5 likes160 downloads2y agoHugging Face02Bisilivan /dataset-ohada-droit-commercial-general-echantillon Dataset OHADA — Droit Commercial Général (AUDCG) — Échantillon Description Échantillon de 10 entrées extraites d'un dataset de fine-tuning juridique en cours de conception, portant sur l'Acte Uniforme relatif au Droit Commercial Général (AUDCG) — le texte fondamental du statut du commerçant, des actes de commerce, de la preuve et de la prescription en matière commerciale dans l'espace OHADA (Organisation pour l'Harmonisation en Afrique du Droit des Affaires — 17… See the full description on the dataset page: https://huggingface.co/datasets/Bisilivan/dataset-ohada-droit-commercial-general-echantillon.texttext-generationn<1K1 likes86 downloads2mo agoHugging Face03Bisher /fadel-diacritizationadapted from: https://arxiv.org/abs/1905.01965 texttext-classification100K<n<1M0 likes52 downloads1y agoHugging Face04pkloats /bishop-chess-dataset Bishop Chess Concepts Dataset Cleaned, training-ready chess text focused on bishop concepts and strategy, distilled from the Waterhorse/chess_data dataset. At a glance 15,677 records · ~18.59M tokens (measured with the GPT-4 cl100k BPE). Game-level-disjoint train/test split (no game leaks across splits), seed 20260708, test fraction 0.02. Split Records Tokens train 15,363 18,152,676 test 314 433,026 Sources (origin): Source Train records… See the full description on the dataset page: https://huggingface.co/datasets/pkloats/bishop-chess-dataset.texttext-generation10K<n<100K0 likes40 downloads2mo agoHugging Face05ardakalayci /bist-tool-sft BIST Tool SFT Fine-tuning data for market-data tool calling (Gemma 4 style), Turkish and English. ~1200 examples, 23 tools, BIST stock-flow heavy. Includes a large manual curated block of natural TR questions (bilanço/kârlılık/KAP/sermaye/teknik/tarama) with agent-style multi-tool plans; EUREN always multi-market search (never FX). Files File Format gemma4_sft.jsonl Alpaca: instruction, input, output gemma4_sft_text.jsonl Full chat as { "text": "..."… See the full description on the dataset page: https://huggingface.co/datasets/ardakalayci/bist-tool-sft.texttext-generation1K<n<10K0 likes36 downloads2mo agoHugging Face06eltociear /bishop-tasks-v1 bishop-tasks-v1 43 exact pattern-recognition tasks for the bishop-env RL environment, on the topics of Pattern Recognition and Machine Learning (Bishop): probability and Bayes, information theory, linear regression and ridge, naive Bayes, Bernoulli mixtures and EM, k-means, conjugate priors, and d-separation in directed graphical models. field meaning task_id bi-000 … bi-042 category probability / information-theory / regression / classification / bayesian-inference… See the full description on the dataset page: https://huggingface.co/datasets/eltociear/bishop-tasks-v1.texttext-generationn<1K0 likes32 downloads2mo agoHugging Face07bisonnetworking /Medical-Reasoning-SFT-GPT-OSS-120B Medical-Reasoning-SFT-GPT-OSS-120B A high-quality synthetic dataset of medical reasoning conversations generated using OpenAI's gpt-oss-120B model with reasoning effort set to high, designed for supervised fine-tuning of large language models in healthcare applications. I used Intelligent-Internet/II-Medical-Reasoning-SFT as a seed dataset, so I would like to thank the authors and Intelligent-Internet for their great work. Dataset Statistics Total Samples: 200,927… See the full description on the dataset page: https://huggingface.co/datasets/bisonnetworking/Medical-Reasoning-SFT-GPT-OSS-120B.texttext-generation100K<n<1M0 likes28 downloads9mo agoHugging Face08Bisilivan /dataset-ohada-droit-societes-echantillon Dataset OHADA — Droit des sociétés commerciales (AUSCGIE) — Échantillon Description Échantillon de 10 entrées extraites d'un dataset de fine-tuning juridique de 40 entrées portant sur les dispositions générales de l'Acte Uniforme relatif au Droit des Sociétés Commerciales et du Groupement d'Intérêt Économique (AUSCGIE), le texte fondamental du droit des sociétés dans l'espace OHADA (Organisation pour l'Harmonisation en Afrique du Droit des Affaires — 17 pays… See the full description on the dataset page: https://huggingface.co/datasets/Bisilivan/dataset-ohada-droit-societes-echantillon.texttext-generationn<1K0 likes22 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.