datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
folktalesmental_health_reddit_postsquotes-500kTaken from Kaggle: https://www.kaggle.com/datasets/manann/quotes-500k?resource=download
It was upload there from this repo: https://github.com/ShivaliGoel/Quotes-500K
Paper:
Goel, S., Madhok, R., & Garg, S. (2018). Proposing Contextually Relevant Quotes for Images. Advances in Information Retrieval. Springer. doi: 10.1007/978-3-319-76941-7_49
Intiial-Knowledge-And-Detailed-Assessment-JSON-Format-DataCARD
CARD — Causal Recovery of Demand
Can a model that fits observed demand well still recover causal price response, substitution, and counterfactual outcomes when prices and promotions are endogenous?
CARD pairs synthetic retail scanner panels with marketing-copy product descriptions that carry the true substitution geometry. Demand is simulated from a known data-generating process; in half the cells, promotion depth responds to a hidden demand shock, so estimators that ignore… See the full description on the dataset page: https://huggingface.co/datasets/jean-jsj/CARD.synthetic-indian-logical-reasoning-CoTyes
question-answering-ukrainian-json-answersInvoice-to-Json
Invoice-to-Json Dataset
Dataset Description
Dataset Summary
Invoice-to-Json is a dataset designed for document understanding and information extraction tasks. It consists of document images paired with questions and answers, specifically focused on extracting structured information (JSON format) from documents.
Supported Tasks
Document Question Answering: The dataset supports training models to answer questions about document content
Information… See the full description on the dataset page: https://huggingface.co/datasets/shubh303/Invoice-to-Json.JSON-Unstructured-StructuredDataset Contains Synthetically Generated Unstructured Text, Set of Rules for Schema Creation, Filled Structured JSON
Can be used for any unstructured to structured tasks
Invoice-to-Json
Invoice-to-Json Dataset
Dataset Description
Dataset Summary
Invoice-to-Json is a dataset designed for document understanding and information extraction tasks. It consists of document images paired with questions and answers, specifically focused on extracting structured information (JSON format) from documents.
Supported Tasks
Document Question Answering: The dataset supports training models to answer questions about document content
Information… See the full description on the dataset page: https://huggingface.co/datasets/acoustichao/Invoice-to-Json.genz-slang-pairs-1k
Gen Z Slang Pairs Corpus (1 K)
The Gen Z Slang Pairs Corpus (1 K) contains 1,000 everyday English sentences alongside their Gen Z–style slang rewrites. This dataset is designed for style-transfer, informal-language generation, and paraphrasing research. Use it to train models that transform formal or neutral sentences into expressive, youth‑oriented slang.
Dataset Details
This dataset was generated programmatically using OpenAI GPT-4.1 Nano.
Language:… See the full description on the dataset page: https://huggingface.co/datasets/JScharp/genz-slang-pairs-1k.ek100-mir-demo-assetsRaw-Web-Scraped-to-JSONherman-json-mode
Herman: Indonesian Single-Turn JSON Mode
Herman is an Indonesian language dataset specifically designed
for training LLMs using a single-turn JSON mode. This dataset
is used in Supervised Fine-Tuning (SFT) to improve JSON parsing
capabilities in LLMs. Herman was obtained from Hermes and translated
into Indonesian for the purpose of training Indonesian language models.
Code used for constructing Herman can be found here.
Schema Format
The desired JSON schema can… See the full description on the dataset page: https://huggingface.co/datasets/SulthanAbiyyu/herman-json-mode.narrative-engine-emotion-5c
Try the PV Peak/Valley Explorer🔗 PV Radar (Beta) Space: https://huggingface.co/spaces/jsisonou/narrative-engine-pv-radar-betaUse this dataset’s sample files to test:
Curve Mode: upload book_curve.scene.csv → Run
Text Mode: paste one scene per line → RunYou’ll get pv_pred (per-scene labels), arc_summary (global peak/valley), and score curves.Assistive only; human-in-the-loop. No model weights or training recipes are exposed.
⚠️ This repository is no longer maintained.👉 Please visit the… See the full description on the dataset page: https://huggingface.co/datasets/jsisonou/narrative-engine-emotion-5c.MtG-json-to-ForgeScriptDinoV2-YGO-card-embeddingsJS-Code-Solutions
Python Code Solutions
Features
1000k of JS Code Solutions for Text Generation and Question Answering
JS Coding Problems labelled by topic and difficulty
Recommendations
Train your Model on Logical Operations and Mathematical Problems Before Training it on this. This is optional for Fine Tuning 2B parameter + models.
Format the prompts in a orderly way when formatting data eg. {question} Solution: {solution} Topic: {topic}
Fraud_Case_Verdicts
The "Crime Facts" of "Offenses of Fraudulence" in Judicial Yuan Verdicts Dataset
This data set is based on the judgments of "Offenses of Fraudulence" cases published by the Judicial Yuan. The data range of the dataset is from January 1, 2011, to December 31, 2021. 74,823 pieces of original data (judgments and rulings) were collected. We only took the contents of the "criminal facts" field of the judgment. This dataset is divided into three parts. The training dataset has 59,858… See the full description on the dataset page: https://huggingface.co/datasets/jslin09/Fraud_Case_Verdicts.JSON_PreferenceJ-Shuwa
J-Shuwa
J-Shuwa is a parallel corpus for Japanese Sign Language (JSL) and Japanese, collected from YouTube videos accessible as of approximately June 2023. It is designed to support research on Japanese Sign Language translation and related multimodal language tasks.
Because the original videos and associated textual content cannot be redistributed, this Hugging Face release provides only the redistributable metadata layer: YouTube video IDs, segment timestamps, and a source… See the full description on the dataset page: https://huggingface.co/datasets/mouwjone/J-Shuwa.Evol-Instruct-JS-Code-500-v1JSON_Preference_decomposed
JSON_Preference_decomposed
A length / syntax / semantic decomposition of the original
nruia/JSON_Preference
dataset. For each preference pair (y1, y2), two intermediate responses
y2'' (double prime) and y2' (prime) are added so that the total
alignment gap
G(y1, y2) = log P(y1 | x) - log P(y2 | x)
can be decomposed along a path of intermediate latent representations:
Step
Quantity
Interpretation
1
`log P(y2''
x) - log P(y2
2
`log P(y2'
x) - log P(y2''
3
`log P(y1
x)… See the full description on the dataset page: https://huggingface.co/datasets/Bojian92/JSON_Preference_decomposed.webnovel-emotion-5c-free
Try the PV Peak/Valley Explorer🔗 PV Radar (Beta) Space: https://huggingface.co/spaces/jsisonou/narrative-engine-pv-radar-betaUse this dataset’s sample files to test:
Curve Mode: upload book_curve.scene.csv → Run
Text Mode: paste one scene per line → RunYou’ll get pv_pred (per-scene labels), arc_summary (global peak/valley), and score curves.Assistive only; human-in-the-loop. No model weights or training recipes are exposed.
⚠️ This repository is no longer maintained.👉 Please visit the… See the full description on the dataset page: https://huggingface.co/datasets/jsisonou/webnovel-emotion-5c-free.financial_conversation_summary
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/jsonfin17/financial_conversation_summary.Race-text-to-quiz-jsonjson_with_nulljson_datajsonsearch2E-Commerce_Customer_Support_Conversations_JSON_Output
