datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nepali-bias-dataset
Nepali Bias Language Dataset
Dataset Description
A synthetic dataset of Nepali sentences labeled for
bias categories including gender, religion, caste,
regional, appearance, social status, political, age,
and disability bias. Sentences were first labeled by
LLMs (ChatGPT, Grok) prompted with real Nepali news
context, then manually reviewed and corrected by human
annotators.
Dataset Summary
Split
Examples
Train
1,362
Validation
292… See the full description on the dataset page: https://huggingface.co/datasets/ios-ioe/nepali-bias-dataset.ios-risk-finetune-v3
IOS Risk Fine-Tune Dataset v3
Quality-gated instruction-tuning data for financial fraud, AML typologies, and
Bank Secrecy Act regulatory recall. This is a research dataset assembled from
public data, official public regulations, deterministic synthetic scenarios,
and validated model-assisted rewrites. It is not production transaction evidence.
Composition
Source
Records
Description
Public tabular benchmark
9,242
ULB/Kaggle credit-card examples; record… See the full description on the dataset page: https://huggingface.co/datasets/Etherlabs/ios-risk-finetune-v3.1232122Scramjet
Scramjet is an interception-based web proxy designed to bypass arbitrary web browser restrictions, support a wide range of sites, and act as middleware for open-source projects. It prioritizes security, developer friendliness, and performance.
Supported Sites
Scramjet has CAPTCHA support! Some of the popular websites that Scramjet supports include:
Google
Twitter
Instagram
Youtube
Spotify
Discord
Reddit
GeForce NOW
Ensure you are not hosting on a… See the full description on the dataset page: https://huggingface.co/datasets/IOSIII123212/1232122.iosfinalioscoder2setiosunsloth3Refined-Anime-Text
Refined Anime Text for Continual Pre-training of Language Models
This is a subset of our novel synthetic dataset of anime-themed text, containing over 1M entries, ~440M GPT-4/3.5 tokens. This dataset has never been publicly released before. We are releasing this subset due to the community's interest in anime culture, which is underrepresented in general-purpose datasets, and the low quality of raw text due to the prevalence of internet slang and irrelevant content, making it… See the full description on the dataset page: https://huggingface.co/datasets/iosscu/Refined-Anime-Text.iosdataimprovedioscoder4fixsetiosasioscoder6.5Iosessentialsiosv5Ioscoder4set
