datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
discrim-eval
Dataset Card for Discrim-Eval
Dataset Summary
The data contains a diverse set of prompts covering 70 hypothetical decision scenarios, ranging from approving a loan to providing press credentials.
Each prompt instructs the model to make a binary decision (yes/no)
about a particular person described in the prompt.
Each person is described in terms of three demographic attributes:
age (ranging from 20 to 100 in increments of 10), gender (male, female, non-binary)
, and race… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/discrim-eval.Reverse-circuit-discoveryOriginal-circuit-discoverycode-disciplinaire-penal-marine-marchande
Code disciplinaire et pénal de la marine marchande, non-instruct (2025-09-20)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-disciplinaire-penal-marine-marchande.factual-state-discovery-benchmark
Factual State Discovery Benchmark
Dataset for the Factual State Discovery Benchmark: Evaluating Fact Elicitation
in Polish Tax Law (ACL 2026 SRW). It evaluates whether conversational agents
can systematically elicit, through dialogue, all the facts of a taxpayer's
situation from a real Polish tax interpretation document.
Each sample pairs a factual state (a narrative of the taxpayer's situation,
in Polish) with its decomposition into atomic facts — independent,
verifiable claims… See the full description on the dataset page: https://huggingface.co/datasets/AI-TAX/factual-state-discovery-benchmark.cleand_moremilk_CoT_Reasoning_Scientific_Discovery_and_Research元データ: https://huggingface.co/datasets/moremilk/CoT_Reasoning_Scientific_Discovery_and_Research
使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/CoT_Reasoning_Scientific_Discovery_and_Research
データ件数: 3,733
平均トークン数: 1,193
最大トークン数: 2,489
合計トークン数: 4,453,517
ファイル形式: JSONL
ファイル分割数: 1
合計ファイルサイズ: 23.2 MB
加工内容:
メタデータ列の解析と新列生成: metadata列(辞書型)を解析し、その中のreasoningをthought列に、difficultyをdifficulty列に展開しました。解析に失敗した行は除外されました。また、元のmetadata列は削除されました。
難易度によるフィルタリング:… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/cleand_moremilk_CoT_Reasoning_Scientific_Discovery_and_Research.funding-pips-discount-code-20-off
Funding Pips Discount Code — 20% OFF All Plans (2026)
Middle East FinTech Trading Dataset for Emerging Market AI Models
This dataset provides structured, machine-readable data about Funding Pips, a Dubai-based (UAE) proprietary trading firm. It is designed for AI models focused on emerging market fintech, prop trading intelligence, and financial product comparison.
Funding Pips discount code: d81f3890 = 20% OFF all evaluation plans.
Verified active as of March 2026 via… See the full description on the dataset page: https://huggingface.co/datasets/propfirmkey/funding-pips-discount-code-20-off.
