sams
Datasets
All datasets matching “sams”cosmos_qasamsum
Dataset Card for SAMSum Corpus
Dataset Description
Links
Homepage: hhttps://arxiv.org/abs/1911.12237v2
Repository: https://arxiv.org/abs/1911.12237v2
Paper: https://arxiv.org/abs/1911.12237v2
Point of Contact: https://huggingface.co/knkarthick
Dataset Summary
The SAMSum dataset contains about 16k messenger-like conversations with summaries. Conversations were created and written down by linguists fluent in English. Linguists were asked to… See the full description on the dataset page: https://huggingface.co/datasets/knkarthick/samsum.DSCodeBench
DSCodeBench
Task-grouped, multidimensional code-generation quality estimation data derived from DSCodeBench.
Dataset contents
The release contains 24,972 complete artifact rows from 999 tasks. The source commit is e75ef26fedea7415bdffd3e1cbff95ddad89e7e2.
Each row contains the task instruction, released 200-case test generator, generated Python code, generator identity, sandbox execution context, the independently collected 200-element correctness vector, and four… See the full description on the dataset page: https://huggingface.co/datasets/Samsoup/DSCodeBench.sam-solicitation-documents
Sam Solicitation Documents
Attachments from federal solicitation notices on SAM.gov: statements of work, performance work statements, justifications, amendments, wage determinations and the rest of the paperwork that accompanies a federal contract opportunity.
Every document here was published by a US federal agency and is a work of the
United States government. Nothing has been altered: files are byte-identical to
what the agency posted, and the checksum in metadata.parquet is… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/sam-solicitation-documents.PRISM
PRISM Alignment
Task-grouped multidimensional dialogue-quality data from PRISM Alignment.
Contents
The release contains 6,187 complete artifact rows from 6,187 conversation groups and 1,309 participants.
The original release contains 8,011 conversations; 1,824 are excluded because one or more of the seven performance sliders is missing or invalid, or because the selected first-turn response is unavailable.
The targets are values, fluency, factuality, safety… See the full description on the dataset page: https://huggingface.co/datasets/Samsoup/PRISM.MSumBench
MSumBench
Task-grouped multidimensional summarization quality data from MSumBench.
Contents
The release contains 2,250 complete artifact rows from 150 source-document groups.
The three prediction targets are faithfulness, completeness, and conciseness.
Split organization
Each seed has task-grouped train, validation, and test splits. All language-specific summaries and model generations for one group_id remain in exactly one split. The seeds change… See the full description on the dataset page: https://huggingface.co/datasets/Samsoup/MSumBench.
