datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FilmBench
FilmBench — Video Generation Benchmark Dataset
📢 Update (2026-08-02)
Added English prompt files: filmbench_prompts_en.csv is now available with English prompts.
filmbench_prompts_en.csv (1,169 rows): Prompt-level table with English prompts. Columns: uid, task, movie_type (English), en_prompt, reference_url.
📢 Update (2026-07-31)
Fixed a batch of misaligned prompts in filmbench_videos.csv: the zh_prompt column has been recalibrated against the… See the full description on the dataset page: https://huggingface.co/datasets/skylenage/FilmBench.genter-ajibawa-name-filled
GENTER Ajibawa Name-Filled
This dataset expands aieng-lab/genter-ajibawa by inserting concrete names into each template.
It provides nested Hugging Face configs with 1, 2, 5, or 10 names per gender and template.
For every template sentence, names are sampled independently from NAMEXACT (matching split; frequency-weighted), using K female and K male names in config nK.
It is intended for experiments that need concrete text rather than [NAME]/[MASK] placeholders while still… See the full description on the dataset page: https://huggingface.co/datasets/aieng-lab/genter-ajibawa-name-filled.test_many_filesfiling-boards
POPS4 Filing Boards
Nine tables built from United States SEC filings, refreshed every business day, each row linked to the
original document on EDGAR.
Live boards: https://www.pops4.com/boards
Source repository: https://github.com/HuangGoodmanAgency/filing-boards
Dataset home: https://huggingface.co/datasets/jennyota/filing-boards
What makes this dataset unusual
Nothing in it is written. Every value either comes from a document that the row links to,
or is a… See the full description on the dataset page: https://huggingface.co/datasets/jennyota/filing-boards.traffic-demand-csv-files-1osworld_tasks_filesvlmevalkit_filesjailbreak-filter-outputstagalog-filipino-english-translationThis dataset is a Tagalog-English translation data. It is a compiled comma-separated values dataset from different
existing HuggingFace and External dataset.
Here are the collected and compiled data:
saillab/alpaca_tamil_taco
DIBT/MPEP_FILIPINO
Nag, S., Ma, S., Ntalli, A., & Dulay, K. M. (2024, June 10). TalkTogether. https://doi.org/10.17605/OSF.IO/3ZDFN
2016_2022_hate_speech_filipino
Dataset Card for 2016 and 2022 Hate Speech in Filipino
Dataset Summary
Contains a total of 27,383 tweets that are labeled as hate speech (1) or non-hate speech (0). Split into 80-10-10 (train-validation-test) with a total of 21,773 tweets for training, 2,800 tweets for validation, and 2,810 tweets for testing.
Created by combining hate_speech_filipino and a newly crawled 2022 Philippine Presidential Elections-related Tweets Hate Speech Dataset.
This dataset has an almost… See the full description on the dataset page: https://huggingface.co/datasets/mapsoriano/2016_2022_hate_speech_filipino.filesystem_huggingface_yahoo-finance_excel_terminal_8590_watchlistexcel_filesystem_terminal_huggingface_2307_narc9ibmosworld_tasks_filesosworld_tasks_filesosworld_tasks_fileshuggingface_filesystem_terminal_12696_orders_f567d517
Order Data Lake - f567d517
Source repository containing order master data and the nightly audit log.
pdf-tools_huggingface_terminal_filesystem_7936_procurement_agreements_rgtfq5
Northwind Trading Co. - Procurement Agreements
A curated dataset of procurement agreements from Northwind Trading Co.'s Q3 2026
contract audit batch. Records include supply agreements, purchase agreements,
master procurement agreements and framework supply agreements. This dataset is
intended for vendor-risk classification model training.
huggingface_filesystem_terminal_12696_orders_5c9af4bb
Order Data Lake - 5c9af4bb
Source repository containing order master data and the nightly audit log.
huggingface_filesystem_terminal_12696_results_f567d517Results
crypto-related-labelingtelugu_alpaca_yahma_cleaned_filtered_romanizedfilesystem_huggingface_9831_jpwp2z
Product Catalog
Retail distributor product catalog with inventory levels used for reorder planning.
filesystem_huggingface_9840_z1xmjtic_support_tickets
Support Ticket Export
Fresh export of support tickets from the company's customer support system.
Each record contains ticket metadata, priority, assignment, SLA deadline and
whether the ticket has breached its SLA window. This is the source dataset for
the support-ticket triage workflow.
File: tickets.csv
email-spam-filterJParaCrawl-Filtered-English-Japanese-Parallel-Corpus
Introduction
This is a LLM-filtered set of the first 1M rows from ntt's JParaCrawl v3 large English-Japanese parallel corpus.
The original JParaCrawl corpus was put together by automated means - aligning Japanese texts with their apparent English translations that were found in-the-wild, on the internet.
Whilst manually browsing the original data, I noticed that there were obvious quality issues that made me anxious about using the dataset at all. Poorly aligned translations… See the full description on the dataset page: https://huggingface.co/datasets/Verah/JParaCrawl-Filtered-English-Japanese-Parallel-Corpus.ascend
Dataset Card for ASCEND
Dataset Summary
ASCEND (A Spontaneous Chinese-English Dataset) introduces a high-quality resource of spontaneous multi-turn conversational dialogue Chinese-English code-switching corpus collected in Hong Kong. ASCEND consists of 10.62 hours of spontaneous speech with a total of ~12.3K utterances. The corpus is split into 3 sets: training, validation, and test with a ratio of 8:1:1 while maintaining a balanced gender proportion on each set.… See the full description on the dataset page: https://huggingface.co/datasets/filwsyl/ascend.huggingface_filesystem_terminal_12696_orders_6b361bae
Order Data Lake - 6b361bae
Source repository containing order master data and the nightly audit log.
shelfglance-ucp-category-filter
Shopify's agent-commerce category filter didn't filter on any of 190 stores
Every Shopify store answers an agent-commerce endpoint whose schema declares a category filter. Sent to 200 live stores with a price-filter control: 186 ignored it, 4 rejected every value, 0 filtered.
One row per Shopify storefront that answered its own agent-commerce endpoint (POST /api/ucp/mcp) on 2 September 2026: 190 of 200 sampled from a corpus of 10,099. Each row records what search_catalog… See the full description on the dataset page: https://huggingface.co/datasets/bpm2007/shelfglance-ucp-category-filter.filesystem_fetch_hf_playwright_googlemap_terminal_github_scholarly_8016_bwdelreg_1zytb4
BlueWave Logistics Delivery-Point Registry
This dataset maintains the delivery-point registry for BlueWave Logistics (regional freight & dispatch).
Files
registry.csv — the master delivery-point registry.
review_decisions.csv — the latest Q3 2026 review report (published by the operations analyst).
registry.csv schema
Columns: id,branch,address,city,state,status,review_month
id: delivery-point identifier (e.g. DP-101).
branch: operations branch… See the full description on the dataset page: https://huggingface.co/datasets/zhuq41/filesystem_fetch_hf_playwright_googlemap_terminal_github_scholarly_8016_bwdelreg_1zytb4.HelpSteer-filtered
HelpSteer-filtered
This dataset is a highly filtered version of the nvidia/HelpSteer dataset.
❓ How this dataset was filtered:
I calculated the sum of the columns ["helpfulness," "correctness," "coherence," "complexity," "verbosity"] and created a new column named sum.
I changed some column names and added a empty column to match the Alpaca format.
The dataset was then filtered to include only those entries with a sum greater than or equal to 16.
🧐 More… See the full description on the dataset page: https://huggingface.co/datasets/Weyaxi/HelpSteer-filtered.
