CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Lilambd /world-signals World Signals — a daily cross-country snapshot of attention One folder per day under data/YYYY-MM-DD/, and the same files copied to latest/. Built every morning (JST) by the EmpireOS world model. Nothing is generated by a model; every row is a measurement from a public source. file what source search_trends.csv rising searches, 30 countries, with approximate traffic and the headline that drove them Google Trends daily RSS podcast_charts.csv top-100 podcasts, 30… See the full description on the dataset page: https://huggingface.co/datasets/Lilambd/world-signals.tabular1K<n<10K0 likes1.1k downloads10h agoHugging Face02harvard-lil /cold-french-law Collaborative Open Legal Data (COLD) - French Law COLD French Law is a dataset containing over 800 000 french law articles, filtered and extracted from France's LEGI dataset and formatted as a single CSV file. This dataset focuses on articles (codes, lois, décrets, arrêtés ...) identified as currently applicable french law. A large portion of this dataset comes with machine-generated english translations, provided by Casetext, Part of Thomson Reuters using OpenAI's GPT-4. This… See the full description on the dataset page: https://huggingface.co/datasets/harvard-lil/cold-french-law.tabular100K<n<1M21 likes351 downloads2y agoHugging Face03Lilambd /world-seeds World Seeds — every "by country" table, keyed by ISO 3166-1 alpha-2 Wikipedia has hundreds of "... by country" articles. The numbers live inside article tables, keyed by country names that differ from article to article. This dataset re-keys every such table to ISO2 so they join. One CSV per source article under tables/. Columns: iso2, country, <original column names>. Values are kept exactly as printed (*_num twin columns hold the parsed number where one could be read).… See the full description on the dataset page: https://huggingface.co/datasets/Lilambd/world-seeds.tabular1K<n<10K0 likes167 downloads11h agoHugging Face04Lil-R /V2-Space-Datasettextn<1K0 likes75 downloads2y agoHugging Face05lilianngweta /aligners-datasets Project information Paper: https://arxiv.org/pdf/2403.04224 Repository: https://github.com/lilianngweta/aligners text100K<n<1M0 likes67 downloads2y agoHugging Face06lilu007 /womens-clothing-reviewstabular10K<n<100K0 likes45 downloads2y agoHugging Face07li-long /financial_regulatory_qatext10K<n<100K1 likes37 downloads2y agoHugging Face08lilyzhng /c-guard C-Guard: A Constitution-Grid Instrument for Data-Efficient RL Alignment Released data and constitution for the paper "A Constitution-Grid Instrument for Data-Efficient RL Alignment (C-Guard)" (COLM 2026 Efficient Reasoning Workshop). Guards run inline on every LLM turn, so the job is high-volume, short-prompt, and latency-bound. C-Guard asks: given a fixed 4B base model and a read-only XSTest eval, can targeted synthetic data + GRPO shrink over-refusal without opening disguise… See the full description on the dataset page: https://huggingface.co/datasets/lilyzhng/c-guard.texttext-classificationn<1K0 likes35 downloads1mo agoHugging Face09lilyray /emo_motiv_tomitext100K<n<1M0 likes30 downloads2y agoHugging Face10lilyray /tomi_sileodtext10K<n<100K0 likes30 downloads2y agoHugging Face11LilithHu /manipulative-chinese-dataset Dataset for Emotionally Manipulative vs. Non-Manipulative Chinese Texts Summary This dataset consists of 10,000 Chinese texts evenly split into 5,000 manipulative and 5,000 non-manipulative samples. It was constructed to support research on detecting emotionally manipulative language. Dataset Structure Non-manipulative texts (5,000) 2,700 adapted from common Chinese sentence patterns 2,300 from a Chinese social media corpus on Hugging Face Manipulative… See the full description on the dataset page: https://huggingface.co/datasets/LilithHu/manipulative-chinese-dataset.texttext-classification10K<n<100K1 likes26 downloads1y agoHugging Face12li-long /modified_erotic_literature_collectiontext10K<n<100K5 likes25 downloads2y agoHugging Face13LiliDuenas /big-data-movies-datasetimage1M<n<10M0 likes22 downloads6mo agoHugging Face14Lil-R /Small_100textn<1K0 likes20 downloads2y agoHugging Face15lilyray /story_emotion_classificationtext100K<n<1M0 likes17 downloads2y agoHugging Face16plm3332 /lilab_llm_testtextn<1K0 likes15 downloads3y agoHugging Face17lilyray /emo_motivtext100K<n<1M0 likes14 downloads2y agoHugging Face18Lil-R /V1-Space-Datasettextn<1K0 likes14 downloads2y agoHugging Face19lilyray /story_emotion_inferencetext100K<n<1M0 likes12 downloads2y agoHugging Face20lilyray /emo_motiv_sileodtext100K<n<1M0 likes11 downloads2y agoHugging Face21lilas12 /iqtisadiyyat_datasettext1K<n<10K0 likes9 downloads1y agoHugging Face22Lilian5 /Kenyanrecipedatasettabularn<1K0 likes8 downloads2y agoHugging Face23lilolyhh /OLA OLA: Output Language Alignment Benchmark OLA is a benchmark designed to evaluate LLMs' Output Language Alignment in code-switched interactions Dataset Structure OLA consists of two settings: Simple and Complex. Simple Setting The Simple setting focuses on intra-sentential code-switching, where the expected response language is the matrix language—the language providing the core grammatical structure into which elements from another language are embedded.… See the full description on the dataset page: https://huggingface.co/datasets/lilolyhh/OLA.texttext-generation1K<n<10K0 likes7 downloads5mo agoHugging Face24lilpad /test3tabularn<1K0 likes6 downloads3y agoHugging Face25lilas12 /tehsil_datasetstext1K<n<10K0 likes6 downloads1y agoHugging Face26lilas12 /olke_datasetstext1K<n<10K0 likes6 downloads1y agoHugging Face27li-long /financial_media_sentimenttext10K<n<100K1 likes5 downloads1y agoHugging Face28LilaMartinez4322 /computerequipmentpricesA dataset containing stock prices and equipment details from various brands, covering a range of products and price points. Feel free to use :> tabularn<1K0 likes3 downloads2y agoHugging Face29lilpad /test2tabularn<1K0 likes2 downloads3y agoHugging Face30lilyray /story_motivation_inferencegatedtext100K<n<1M1 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.