datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pi-diff-review
Coding agent session traces for badlogicgames/pi-diff-review
This dataset contains redacted coding agent session traces collected while working on https://github.com/badlogic/pi-diff-review.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line… See the full description on the dataset page: https://huggingface.co/datasets/badlogicgames/pi-diff-review.badlogicgames-pi-mono-opus-filteredFiltered version of badlogicgames/pi-mono - Only opus traces, dropped invalid sessions as well.
All traces present are training safe and teich compatible
civic-records-distill
civic-records-distill
Training data for a local model that helps a private citizen use public-records
law: draft requests that are hard to stall, turn an angry draft into a letter an
official has to engage with, look things up instead of inventing them, and
escalate correctly when stonewalled.
Grounded in Florida (ch. 119 Public Records Act, ch. 286 Sunshine Law, and
the ALPR-specific s. 316.0777) and Texas (ch. 552 Public Information Act,
ch. 551 Open Meetings Act).
Pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/h0ney-badger/civic-records-distill.BADAG01BADAG03BADAG04BADAG05BADAG02BADAG18BADAG06BADAG08BADAG07BADAG11BADAG13BADAG21BADAG10BADAG15BADAG16BADAG19BADAG12BADAG20BADAG09BADAG17BADAG14FLAN-Small
FLAN-Small
This repository is a reduced version of the data provided by the hardwork of: https://huggingface.co/datasets/imone/OpenOrca_FLAN.
FLAN-Small amounts to ~10m examples sampled to approximately hold to the FLAN's final "submix" of:
{
'flan': 0.4,
't0': 0.32,
'niv2': 0.20,
'cot': 0.05,
'dialog': 0.03
}
Since the cot data is rather small -- this was sampled with replacement; consequently there are some duplicates.
Some token length… See the full description on the dataset page: https://huggingface.co/datasets/BadDepartment/FLAN-Small.Bad_Data_Alpaca中文
README for bad_data.json Dataset
Updated on 2024.8.22: Important: For security reasons, the current dataset is an abridged version. See Bad_Data.
bad_data.json
Overview
The bad_data.json dataset is a collection of text data specifically curated for training and evaluating language models on challenging and sensitive content. The dataset covers a wide range of topics, including ethical dilemmas, illegal activities, pornographic content, and… See the full description on the dataset page: https://huggingface.co/datasets/ystemsrx/Bad_Data_Alpaca.python-codes-25k
License
MIT
This is a Cleaned Python Dataset Covering 25,000 Instructional Tasks
Overview
The dataset has 4 key features (fields): instruction, input, output, and text.It's a rich source for Python codes, tasks, and extends into behavioral aspects.
Dataset Statistics
Total Entries: 24,813
Unique Instructions: 24,580
Unique Inputs: 3,666
Unique Outputs: 24,581
Unique Texts: 24,813
Average Tokens per example: 508
Features… See the full description on the dataset page: https://huggingface.co/datasets/badaranta/python-codes-25k.lince-sa-refined
LINCE SA Refined — Cultural-Context Relabeling for Spanish-English Code-Switching Sentiment Analysis
This repository releases 763 sentiment label refinements for the Spanish-English code-switching subset (sa_spaeng) of the LINCE benchmark (Aguilar et al., 2020). Refinements were produced by a trilingual annotator (Spanish / English / Korean) drawing on Hispanic-American social media conventions, and validated through controlled mBERT experiments.
This work received an Honorable… See the full description on the dataset page: https://huggingface.co/datasets/badashin/lince-sa-refined.qw35-27b-badcase-searchbadedit-train-itsm
