CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01skylenage /FilmBench FilmBench — Video Generation Benchmark Dataset 📢 Update (2026-08-02) Added English prompt files: filmbench_prompts_en.csv is now available with English prompts. filmbench_prompts_en.csv (1,169 rows): Prompt-level table with English prompts. Columns: uid, task, movie_type (English), en_prompt, reference_url. 📢 Update (2026-07-31) Fixed a batch of misaligned prompts in filmbench_videos.csv: the zh_prompt column has been recalibrated against the… See the full description on the dataset page: https://huggingface.co/datasets/skylenage/FilmBench.text1K<n<10K3 likes6.6k downloads2mo agoHugging Face02aieng-lab /genter-ajibawa-name-filled GENTER Ajibawa Name-Filled This dataset expands aieng-lab/genter-ajibawa by inserting concrete names into each template. It provides nested Hugging Face configs with 1, 2, 5, or 10 names per gender and template. For every template sentence, names are sampled independently from NAMEXACT (matching split; frequency-weighted), using K female and K male names in config nK. It is intended for experiments that need concrete text rather than [NAME]/[MASK] placeholders while still… See the full description on the dataset page: https://huggingface.co/datasets/aieng-lab/genter-ajibawa-name-filled.text1M<n<10M1 likes591 downloads1mo agoHugging Face03samansmink /test_many_filesn<1K0 likes305 downloads2y agoHugging Face04jennyota /filing-boards POPS4 Filing Boards Nine tables built from United States SEC filings, refreshed every business day, each row linked to the original document on EDGAR. Live boards: https://www.pops4.com/boards Source repository: https://github.com/HuangGoodmanAgency/filing-boards Dataset home: https://huggingface.co/datasets/jennyota/filing-boards What makes this dataset unusual Nothing in it is written. Every value either comes from a document that the row links to, or is a… See the full description on the dataset page: https://huggingface.co/datasets/jennyota/filing-boards.tabulartabular-classification1K<n<10K0 likes261 downloads8h agoHugging Face05pushpender-23 /traffic-demand-csv-files-1tabular100K<n<1M0 likes254 downloads4mo agoHugging Face06sabir15 /osworld_tasks_filesdocumentn<1K0 likes179 downloads9mo agoHugging Face07chandrabhuma /vlmevalkit_filestext10K<n<100K0 likes159 downloads9mo agoHugging Face08Sukratii /jailbreak-filter-outputstabularn<1K1 likes111 downloads5mo agoHugging Face09rhyliieee /tagalog-filipino-english-translationThis dataset is a Tagalog-English translation data. It is a compiled comma-separated values dataset from different existing HuggingFace and External dataset. Here are the collected and compiled data: saillab/alpaca_tamil_taco DIBT/MPEP_FILIPINO Nag, S., Ma, S., Ntalli, A., & Dulay, K. M. (2024, June 10). TalkTogether. https://doi.org/10.17605/OSF.IO/3ZDFN texttranslation100K<n<1M6 likes91 downloads2y agoHugging Face10mapsoriano /2016_2022_hate_speech_filipino Dataset Card for 2016 and 2022 Hate Speech in Filipino Dataset Summary Contains a total of 27,383 tweets that are labeled as hate speech (1) or non-hate speech (0). Split into 80-10-10 (train-validation-test) with a total of 21,773 tweets for training, 2,800 tweets for validation, and 2,810 tweets for testing. Created by combining hate_speech_filipino and a newly crawled 2022 Philippine Presidential Elections-related Tweets Hate Speech Dataset. This dataset has an almost… See the full description on the dataset page: https://huggingface.co/datasets/mapsoriano/2016_2022_hate_speech_filipino.texttext-classification10K<n<100K1 likes85 downloads2y agoHugging Face11TianfuXinqu /filesystem_huggingface_yahoo-finance_excel_terminal_8590_watchlisttextn<1K0 likes84 downloads28d agoHugging Face12Roy229 /excel_filesystem_terminal_huggingface_2307_narc9ibmtabularn<1K0 likes80 downloads1mo agoHugging Face13raj-jha /osworld_tasks_filesdocumentn<1K0 likes74 downloads9mo agoHugging Face14vitor-cirilo-santos /osworld_tasks_filestabularn<1K0 likes72 downloads11mo agoHugging Face15Sanjay4HF /osworld_tasks_filestextn<1K0 likes70 downloads1y agoHugging Face16TianfuXinqu /huggingface_filesystem_terminal_12696_orders_f567d517 Order Data Lake - f567d517 Source repository containing order master data and the nightly audit log. textn<1K0 likes70 downloads28d agoHugging Face17Roy229 /pdf-tools_huggingface_terminal_filesystem_7936_procurement_agreements_rgtfq5 Northwind Trading Co. - Procurement Agreements A curated dataset of procurement agreements from Northwind Trading Co.'s Q3 2026 contract audit batch. Records include supply agreements, purchase agreements, master procurement agreements and framework supply agreements. This dataset is intended for vendor-risk classification model training. textn<1K0 likes67 downloads29d agoHugging Face18TianfuXinqu /huggingface_filesystem_terminal_12696_orders_5c9af4bb Order Data Lake - 5c9af4bb Source repository containing order master data and the nightly audit log. textn<1K0 likes64 downloads29d agoHugging Face19TianfuXinqu /huggingface_filesystem_terminal_12696_results_f567d517Results textn<1K0 likes64 downloads28d agoHugging Face20filipeasm18 /crypto-related-labelingtext10K<n<100K0 likes63 downloads7mo agoHugging Face21Telugu-LLM-Labs /telugu_alpaca_yahma_cleaned_filtered_romanizedtext10K<n<100K19 likes59 downloads3y agoHugging Face22TianfuXinqu /filesystem_huggingface_9831_jpwp2z Product Catalog Retail distributor product catalog with inventory levels used for reorder planning. tabularn<1K0 likes58 downloads1mo agoHugging Face23TianfuXinqu /filesystem_huggingface_9840_z1xmjtic_support_tickets Support Ticket Export Fresh export of support tickets from the company's customer support system. Each record contains ticket metadata, priority, assignment, SLA deadline and whether the ticket has breached its SLA window. This is the source dataset for the support-ticket triage workflow. File: tickets.csv textn<1K0 likes58 downloads1mo agoHugging Face24NotShrirang /email-spam-filtertabulartext-classification1K<n<10K11 likes56 downloads1y agoHugging Face25Verah /JParaCrawl-Filtered-English-Japanese-Parallel-Corpus Introduction This is a LLM-filtered set of the first 1M rows from ntt's JParaCrawl v3 large English-Japanese parallel corpus. The original JParaCrawl corpus was put together by automated means - aligning Japanese texts with their apparent English translations that were found in-the-wild, on the internet. Whilst manually browsing the original data, I noticed that there were obvious quality issues that made me anxious about using the dataset at all. Poorly aligned translations… See the full description on the dataset page: https://huggingface.co/datasets/Verah/JParaCrawl-Filtered-English-Japanese-Parallel-Corpus.tabulartranslation1M<n<10M3 likes49 downloads3y agoHugging Face26filwsyl /ascend Dataset Card for ASCEND Dataset Summary ASCEND (A Spontaneous Chinese-English Dataset) introduces a high-quality resource of spontaneous multi-turn conversational dialogue Chinese-English code-switching corpus collected in Hong Kong. ASCEND consists of 10.62 hours of spontaneous speech with a total of ~12.3K utterances. The corpus is split into 3 sets: training, validation, and test with a ratio of 8:1:1 while maintaining a balanced gender proportion on each set.… See the full description on the dataset page: https://huggingface.co/datasets/filwsyl/ascend.audioautomatic-speech-recognition1K<n<10K1 likes45 downloads4y agoHugging Face27TianfuXinqu /huggingface_filesystem_terminal_12696_orders_6b361bae Order Data Lake - 6b361bae Source repository containing order master data and the nightly audit log. textn<1K0 likes45 downloads1mo agoHugging Face28bpm2007 /shelfglance-ucp-category-filter Shopify's agent-commerce category filter didn't filter on any of 190 stores Every Shopify store answers an agent-commerce endpoint whose schema declares a category filter. Sent to 200 live stores with a price-filter control: 186 ignored it, 4 rejected every value, 0 filtered. One row per Shopify storefront that answered its own agent-commerce endpoint (POST /api/ucp/mcp) on 2 September 2026: 190 of 200 sampled from a corpus of 10,099. Each row records what search_catalog… See the full description on the dataset page: https://huggingface.co/datasets/bpm2007/shelfglance-ucp-category-filter.tabularn<1K0 likes45 downloads17d agoHugging Face29zhuq41 /filesystem_fetch_hf_playwright_googlemap_terminal_github_scholarly_8016_bwdelreg_1zytb4 BlueWave Logistics Delivery-Point Registry This dataset maintains the delivery-point registry for BlueWave Logistics (regional freight & dispatch). Files registry.csv — the master delivery-point registry. review_decisions.csv — the latest Q3 2026 review report (published by the operations analyst). registry.csv schema Columns: id,branch,address,city,state,status,review_month id: delivery-point identifier (e.g. DP-101). branch: operations branch… See the full description on the dataset page: https://huggingface.co/datasets/zhuq41/filesystem_fetch_hf_playwright_googlemap_terminal_github_scholarly_8016_bwdelreg_1zytb4.textn<1K0 likes44 downloads29d agoHugging Face30Weyaxi /HelpSteer-filtered HelpSteer-filtered This dataset is a highly filtered version of the nvidia/HelpSteer dataset. ❓ How this dataset was filtered: I calculated the sum of the columns ["helpfulness," "correctness," "coherence," "complexity," "verbosity"] and created a new column named sum. I changed some column names and added a empty column to match the Alpaca format. The dataset was then filtered to include only those entries with a sum greater than or equal to 16. 🧐 More… See the full description on the dataset page: https://huggingface.co/datasets/Weyaxi/HelpSteer-filtered.tabular1K<n<10K4 likes43 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.