CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01badlogicgames /pi-diff-review Coding agent session traces for badlogicgames/pi-diff-review This dataset contains redacted coding agent session traces collected while working on https://github.com/badlogic/pi-diff-review.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review. Data description Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line… See the full description on the dataset page: https://huggingface.co/datasets/badlogicgames/pi-diff-review.tabulartext-generationn<1K8 likes229 downloads6mo agoHugging Face02armand0e /badlogicgames-pi-mono-opus-filteredFiltered version of badlogicgames/pi-mono - Only opus traces, dropped invalid sessions as well. All traces present are training safe and teich compatible tabularn<1K2 likes200 downloads4mo agoHugging Face03h0ney-badger /civic-records-distill civic-records-distill Training data for a local model that helps a private citizen use public-records law: draft requests that are hard to stall, turn an angry draft into a letter an official has to engage with, look things up instead of inventing them, and escalate correctly when stonewalled. Grounded in Florida (ch. 119 Public Records Act, ch. 286 Sunshine Law, and the ALPR-specific s. 316.0777) and Texas (ch. 552 Public Information Act, ch. 551 Open Meetings Act). Pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/h0ney-badger/civic-records-distill.texttext-generation1K<n<10K0 likes116 downloads20d agoHugging Face04julhaphap /BADAG01textn<1K0 likes106 downloads5mo agoHugging Face05julhaphap /BADAG03textn<1K0 likes105 downloads5mo agoHugging Face06julhaphap /BADAG04textn<1K0 likes104 downloads5mo agoHugging Face07julhaphap /BADAG05textn<1K0 likes104 downloads5mo agoHugging Face08julhaphap /BADAG02textn<1K0 likes100 downloads5mo agoHugging Face09julhaphap /BADAG18textn<1K0 likes100 downloads5mo agoHugging Face10julhaphap /BADAG06textn<1K0 likes99 downloads5mo agoHugging Face11julhaphap /BADAG08textn<1K0 likes99 downloads5mo agoHugging Face12julhaphap /BADAG07textn<1K0 likes96 downloads5mo agoHugging Face13julhaphap /BADAG11textn<1K0 likes94 downloads5mo agoHugging Face14julhaphap /BADAG13textn<1K0 likes93 downloads5mo agoHugging Face15julhaphap /BADAG21textn<1K0 likes93 downloads5mo agoHugging Face16julhaphap /BADAG10textn<1K0 likes92 downloads5mo agoHugging Face17julhaphap /BADAG15textn<1K0 likes92 downloads5mo agoHugging Face18julhaphap /BADAG16textn<1K0 likes91 downloads5mo agoHugging Face19julhaphap /BADAG19textn<1K0 likes91 downloads5mo agoHugging Face20julhaphap /BADAG12textn<1K0 likes88 downloads5mo agoHugging Face21julhaphap /BADAG20textn<1K0 likes87 downloads5mo agoHugging Face22julhaphap /BADAG09textn<1K0 likes86 downloads5mo agoHugging Face23julhaphap /BADAG17textn<1K0 likes86 downloads5mo agoHugging Face24julhaphap /BADAG14textn<1K0 likes83 downloads5mo agoHugging Face25BadDepartment /FLAN-Small FLAN-Small This repository is a reduced version of the data provided by the hardwork of: https://huggingface.co/datasets/imone/OpenOrca_FLAN. FLAN-Small amounts to ~10m examples sampled to approximately hold to the FLAN's final "submix" of: { 'flan': 0.4, 't0': 0.32, 'niv2': 0.20, 'cot': 0.05, 'dialog': 0.03 } Since the cot data is rather small -- this was sampled with replacement; consequently there are some duplicates. Some token length… See the full description on the dataset page: https://huggingface.co/datasets/BadDepartment/FLAN-Small.text1M<n<10M1 likes81 downloads3y agoHugging Face26ystemsrx /Bad_Data_Alpaca中文 README for bad_data.json Dataset Updated on 2024.8.22: Important: For security reasons, the current dataset is an abridged version. See Bad_Data. bad_data.json Overview The bad_data.json dataset is a collection of text data specifically curated for training and evaluating language models on challenging and sensitive content. The dataset covers a wide range of topics, including ethical dilemmas, illegal activities, pornographic content, and… See the full description on the dataset page: https://huggingface.co/datasets/ystemsrx/Bad_Data_Alpaca.texttext-generationn<1K47 likes55 downloads2y agoHugging Face27badaranta /python-codes-25k License MIT This is a Cleaned Python Dataset Covering 25,000 Instructional Tasks Overview The dataset has 4 key features (fields): instruction, input, output, and text.It's a rich source for Python codes, tasks, and extends into behavioral aspects. Dataset Statistics Total Entries: 24,813 Unique Instructions: 24,580 Unique Inputs: 3,666 Unique Outputs: 24,581 Unique Texts: 24,813 Average Tokens per example: 508 Features… See the full description on the dataset page: https://huggingface.co/datasets/badaranta/python-codes-25k.texttext-classification10K<n<100K0 likes48 downloads8mo agoHugging Face28badashin /lince-sa-refined LINCE SA Refined — Cultural-Context Relabeling for Spanish-English Code-Switching Sentiment Analysis This repository releases 763 sentiment label refinements for the Spanish-English code-switching subset (sa_spaeng) of the LINCE benchmark (Aguilar et al., 2020). Refinements were produced by a trilingual annotator (Spanish / English / Korean) drawing on Hispanic-American social media conventions, and validated through controlled mBERT experiments. This work received an Honorable… See the full description on the dataset page: https://huggingface.co/datasets/badashin/lince-sa-refined.texttext-classificationn<1K1 likes45 downloads5mo agoHugging Face29miugod /qw35-27b-badcase-searchtext10K<n<100K0 likes37 downloads6mo agoHugging Face30lvogel /badedit-train-itsmtextn<1K0 likes36 downloads7d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.