datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
world-signals
World Signals — a daily cross-country snapshot of attention
One folder per day under data/YYYY-MM-DD/, and the same files copied to latest/.
Built every morning (JST) by the EmpireOS world model. Nothing is generated by a model; every row is a measurement from a public source.
file
what
source
search_trends.csv
rising searches, 30 countries, with approximate traffic and the headline that drove them
Google Trends daily RSS
podcast_charts.csv
top-100 podcasts, 30… See the full description on the dataset page: https://huggingface.co/datasets/Lilambd/world-signals.cold-french-law
Collaborative Open Legal Data (COLD) - French Law
COLD French Law is a dataset containing over 800 000 french law articles, filtered and extracted from France's LEGI dataset and formatted as a single CSV file.
This dataset focuses on articles (codes, lois, décrets, arrêtés ...) identified as currently applicable french law.
A large portion of this dataset comes with machine-generated english translations, provided by Casetext, Part of Thomson Reuters using OpenAI's GPT-4.
This… See the full description on the dataset page: https://huggingface.co/datasets/harvard-lil/cold-french-law.world-seeds
World Seeds — every "by country" table, keyed by ISO 3166-1 alpha-2
Wikipedia has hundreds of "... by country" articles. The numbers live inside article tables, keyed by country names that differ from article to article. This dataset re-keys every such table to ISO2 so they join.
One CSV per source article under tables/. Columns: iso2, country, <original column names>. Values are kept exactly as printed (*_num twin columns hold the parsed number where one could be read).… See the full description on the dataset page: https://huggingface.co/datasets/Lilambd/world-seeds.V2-Space-Datasetaligners-datasets
Project information
Paper: https://arxiv.org/pdf/2403.04224
Repository: https://github.com/lilianngweta/aligners
womens-clothing-reviewsfinancial_regulatory_qac-guard
C-Guard: A Constitution-Grid Instrument for Data-Efficient RL Alignment
Released data and constitution for the paper "A Constitution-Grid Instrument for Data-Efficient RL Alignment (C-Guard)" (COLM 2026 Efficient Reasoning Workshop).
Guards run inline on every LLM turn, so the job is high-volume, short-prompt, and latency-bound. C-Guard asks: given a fixed 4B base model and a read-only XSTest eval, can targeted synthetic data + GRPO shrink over-refusal without opening disguise… See the full description on the dataset page: https://huggingface.co/datasets/lilyzhng/c-guard.emo_motiv_tomitomi_sileodmanipulative-chinese-dataset
Dataset for Emotionally Manipulative vs. Non-Manipulative Chinese Texts
Summary
This dataset consists of 10,000 Chinese texts evenly split into 5,000 manipulative and 5,000 non-manipulative samples. It was constructed to support research on detecting emotionally manipulative language.
Dataset Structure
Non-manipulative texts (5,000)
2,700 adapted from common Chinese sentence patterns
2,300 from a Chinese social media corpus on Hugging Face
Manipulative… See the full description on the dataset page: https://huggingface.co/datasets/LilithHu/manipulative-chinese-dataset.modified_erotic_literature_collectionbig-data-movies-datasetSmall_100story_emotion_classificationlilab_llm_testemo_motivV1-Space-Datasetstory_emotion_inferenceemo_motiv_sileodiqtisadiyyat_datasetKenyanrecipedatasetOLA
OLA: Output Language Alignment Benchmark
OLA is a benchmark designed to evaluate LLMs' Output Language Alignment in code-switched interactions
Dataset Structure
OLA consists of two settings: Simple and Complex.
Simple Setting
The Simple setting focuses on intra-sentential code-switching, where the expected response language is the matrix language—the language providing the core grammatical structure into which elements from another language are embedded.… See the full description on the dataset page: https://huggingface.co/datasets/lilolyhh/OLA.test3tehsil_datasetsolke_datasetsfinancial_media_sentimentcomputerequipmentpricesA dataset containing stock prices and equipment details from various brands, covering a range of products and price points.
Feel free to use :>
test2story_motivation_inference
