CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jstonge1 /folktalestext1K<n<10K0 likes430 downloads9mo agoHugging Face02jsfactory /mental_health_reddit_poststext10K<n<100K11 likes315 downloads5y agoHugging Face03jstet /quotes-500kTaken from Kaggle: https://www.kaggle.com/datasets/manann/quotes-500k?resource=download It was upload there from this repo: https://github.com/ShivaliGoel/Quotes-500K Paper: Goel, S., Madhok, R., & Garg, S. (2018). Proposing Contextually Relevant Quotes for Images. Advances in Information Retrieval. Springer. doi: 10.1007/978-3-319-76941-7_49 text100K<n<1M4 likes224 downloads3y agoHugging Face04SalehAhmad /Intiial-Knowledge-And-Detailed-Assessment-JSON-Format-Datatext1K<n<10K0 likes188 downloads2y agoHugging Face05jean-jsj /CARD CARD — Causal Recovery of Demand Can a model that fits observed demand well still recover causal price response, substitution, and counterfactual outcomes when prices and promotions are endogenous? CARD pairs synthetic retail scanner panels with marketing-copy product descriptions that carry the true substitution geometry. Demand is simulated from a known data-generating process; in half the cells, promotion depth responds to a hidden demand shock, so estimators that ignore… See the full description on the dataset page: https://huggingface.co/datasets/jean-jsj/CARD.tabular10K<n<100K0 likes155 downloads2mo agoHugging Face06jsdfghdhtrseriu /synthetic-indian-logical-reasoning-CoTyes text1K<n<10K1 likes132 downloads2mo agoHugging Face07nogyxo /question-answering-ukrainian-json-answerstext100K<n<1M5 likes97 downloads3y agoHugging Face08shubh303 /Invoice-to-Json Invoice-to-Json Dataset Dataset Description Dataset Summary Invoice-to-Json is a dataset designed for document understanding and information extraction tasks. It consists of document images paired with questions and answers, specifically focused on extracting structured information (JSON format) from documents. Supported Tasks Document Question Answering: The dataset supports training models to answer questions about document content Information… See the full description on the dataset page: https://huggingface.co/datasets/shubh303/Invoice-to-Json.imagedocument-question-answering10K<n<100K4 likes79 downloads2y agoHugging Face09MasterControlAIML /JSON-Unstructured-StructuredDataset Contains Synthetically Generated Unstructured Text, Set of Rules for Schema Creation, Filled Structured JSON Can be used for any unstructured to structured tasks text10K<n<100K11 likes40 downloads2y agoHugging Face10acoustichao /Invoice-to-Json Invoice-to-Json Dataset Dataset Description Dataset Summary Invoice-to-Json is a dataset designed for document understanding and information extraction tasks. It consists of document images paired with questions and answers, specifically focused on extracting structured information (JSON format) from documents. Supported Tasks Document Question Answering: The dataset supports training models to answer questions about document content Information… See the full description on the dataset page: https://huggingface.co/datasets/acoustichao/Invoice-to-Json.imagedocument-question-answering10K<n<100K0 likes39 downloads4mo agoHugging Face11JScharp /genz-slang-pairs-1k Gen Z Slang Pairs Corpus (1 K) The Gen Z Slang Pairs Corpus (1 K) contains 1,000 everyday English sentences alongside their Gen&nbsp;Z–style slang rewrites. This dataset is designed for style-transfer, informal-language generation, and paraphrasing research. Use it to train models that transform formal or neutral sentences into expressive, youth‑oriented slang. Dataset Details This dataset was generated programmatically using OpenAI GPT-4.1 Nano. Language:… See the full description on the dataset page: https://huggingface.co/datasets/JScharp/genz-slang-pairs-1k.texttext-generation1K<n<10K0 likes39 downloads28d agoHugging Face12jsurrea /ek100-mir-demo-assetsimage10K<n<100K0 likes29 downloads4mo agoHugging Face13Ebullioscopic /Raw-Web-Scraped-to-JSONtextn<1K0 likes27 downloads2y agoHugging Face14SulthanAbiyyu /herman-json-mode Herman: Indonesian Single-Turn JSON Mode Herman is an Indonesian language dataset specifically designed for training LLMs using a single-turn JSON mode. This dataset is used in Supervised Fine-Tuning (SFT) to improve JSON parsing capabilities in LLMs. Herman was obtained from Hermes and translated into Indonesian for the purpose of training Indonesian language models. Code used for constructing Herman can be found here. Schema Format The desired JSON schema can… See the full description on the dataset page: https://huggingface.co/datasets/SulthanAbiyyu/herman-json-mode.texttext-generation1K<n<10K1 likes27 downloads2y agoHugging Face15jsisonou /narrative-engine-emotion-5c Try the PV Peak/Valley Explorer🔗 PV Radar (Beta) Space: https://huggingface.co/spaces/jsisonou/narrative-engine-pv-radar-betaUse this dataset’s sample files to test: Curve Mode: upload book_curve.scene.csv → Run Text Mode: paste one scene per line → RunYou’ll get pv_pred (per-scene labels), arc_summary (global peak/valley), and score curves.Assistive only; human-in-the-loop. No model weights or training recipes are exposed. ⚠️ This repository is no longer maintained.👉 Please visit the… See the full description on the dataset page: https://huggingface.co/datasets/jsisonou/narrative-engine-emotion-5c.tabularn<1K0 likes26 downloads1y agoHugging Face16404NotF0und /MtG-json-to-ForgeScripttext10K<n<100K0 likes24 downloads3y agoHugging Face17jshmatt /DinoV2-YGO-card-embeddingstabular10K<n<100K0 likes23 downloads5mo agoHugging Face18MCES10-Software /JS-Code-Solutions Python Code Solutions Features 1000k of JS Code Solutions for Text Generation and Question Answering JS Coding Problems labelled by topic and difficulty Recommendations Train your Model on Logical Operations and Mathematical Problems Before Training it on this. This is optional for Fine Tuning 2B parameter + models. Format the prompts in a orderly way when formatting data eg. {question} Solution: {solution} Topic: {topic} textquestion-answering10K<n<100K1 likes21 downloads1y agoHugging Face19jslin09 /Fraud_Case_Verdictsgated The "Crime Facts" of "Offenses of Fraudulence" in Judicial Yuan Verdicts Dataset This data set is based on the judgments of "Offenses of Fraudulence" cases published by the Judicial Yuan. The data range of the dataset is from January 1, 2011, to December 31, 2021. 74,823 pieces of original data (judgments and rulings) were collected. We only took the contents of the "criminal facts" field of the judgment. This dataset is divided into three parts. The training dataset has 59,858… See the full description on the dataset page: https://huggingface.co/datasets/jslin09/Fraud_Case_Verdicts.texttext-generation10K<n<100K7 likes20 downloads2y agoHugging Face20nruia /JSON_Preferencetext1K<n<10K0 likes18 downloads1y agoHugging Face21mouwjone /J-Shuwagated J-Shuwa J-Shuwa is a parallel corpus for Japanese Sign Language (JSL) and Japanese, collected from YouTube videos accessible as of approximately June 2023. It is designed to support research on Japanese Sign Language translation and related multimodal language tasks. Because the original videos and associated textual content cannot be redistributed, this Hugging Face release provides only the redistributable metadata layer: YouTube video IDs, segment timestamps, and a source… See the full description on the dataset page: https://huggingface.co/datasets/mouwjone/J-Shuwa.tabular100K<n<1M1 likes18 downloads3mo agoHugging Face22pyto-p /Evol-Instruct-JS-Code-500-v1textn<1K3 likes16 downloads2y agoHugging Face23Bojian92 /JSON_Preference_decomposed JSON_Preference_decomposed A length / syntax / semantic decomposition of the original nruia/JSON_Preference dataset. For each preference pair (y1, y2), two intermediate responses y2'' (double prime) and y2' (prime) are added so that the total alignment gap G(y1, y2) = log P(y1 | x) - log P(y2 | x) can be decomposed along a path of intermediate latent representations: Step Quantity Interpretation 1 `log P(y2'' x) - log P(y2 2 `log P(y2' x) - log P(y2'' 3 `log P(y1 x)… See the full description on the dataset page: https://huggingface.co/datasets/Bojian92/JSON_Preference_decomposed.texttext-classification1K<n<10K0 likes16 downloads4mo agoHugging Face24jsisonou /webnovel-emotion-5c-free Try the PV Peak/Valley Explorer🔗 PV Radar (Beta) Space: https://huggingface.co/spaces/jsisonou/narrative-engine-pv-radar-betaUse this dataset’s sample files to test: Curve Mode: upload book_curve.scene.csv → Run Text Mode: paste one scene per line → RunYou’ll get pv_pred (per-scene labels), arc_summary (global peak/valley), and score curves.Assistive only; human-in-the-loop. No model weights or training recipes are exposed. ⚠️ This repository is no longer maintained.👉 Please visit the… See the full description on the dataset page: https://huggingface.co/datasets/jsisonou/webnovel-emotion-5c-free.tabularn<1K0 likes15 downloads1y agoHugging Face25jsonfin17 /financial_conversation_summary Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/jsonfin17/financial_conversation_summary.textn<1K0 likes14 downloads3y agoHugging Face26Minervus00 /Race-text-to-quiz-jsontext1K<n<10K1 likes14 downloads2y agoHugging Face27shanaka95 /json_with_nulltextn<1K0 likes11 downloads7mo agoHugging Face28Mihir1108 /json_datatextn<1K0 likes10 downloads3y agoHugging Face29sujitvasanth /jsonsearch2textn<1K0 likes10 downloads3y agoHugging Face30shripadkrishna /E-Commerce_Customer_Support_Conversations_JSON_Outputtext1K<n<10K2 likes10 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.