datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ILSA-LLM-Extractor-Dataset
ILSA LLM Extractor Dataset
Project website: https://dedemerve.github.io/ILSA-LLM-Extractor/
Dataset Description
This dataset contains structured metadata automatically extracted from 1,756 peer-reviewed articles and reports covering International Large-Scale Assessments (IEA: TIMSS, PIRLS, ICCS; OECD: PISA, TALIS, PIAAC). The extraction pipeline combines PDF parsing, LLM-based structured extraction, and RAG-based synthesis.
Pipeline stages:
Stage 1: LLM-based… See the full description on the dataset page: https://huggingface.co/datasets/dedemerve/ILSA-LLM-Extractor-Dataset.word_extractorlora-adapters-are-good-feature-extractors
LORA Adapters are Good Feature Extractors Dataset
This dataset contains images of two sets of categories that are not safe for work (hentai and porn, labelled as 0 and 2 correspondingly) and one neutral category, labelled as 2.
The dataset is the source data for training a zoo of LORA adapters on sample images from each category. Adapters representations will then be used as input data to a weight-space model
in an experiment to verify whether WS models operating in low rank… See the full description on the dataset page: https://huggingface.co/datasets/jacekduszenko/lora-adapters-are-good-feature-extractors.sentence-relevance-extractor
Sentence Relevance Extractor (SRE)
Sentence Relevance Extractor (SRE) is a large-scale dataset for binary evidence selection in multi-document, multi-hop question answering.
The goal:
Given a question and a sentence from the context, predict whether this sentence is relevant evidence ("Yes") or irrelevant ("No").
This dataset is suitable for training:
Sentence-level RAG rerankers
Binary relevance classifiers
Optimization-based truth discovery systems
Multi-hop QA evidence… See the full description on the dataset page: https://huggingface.co/datasets/MinaGabriel/sentence-relevance-extractor.3gpp-innovation-extractor-dsmodel_card_extractor_qwq
Dataset card for model_card_extractor_qwq
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"modelId": "digiplay/XtReMixAnimeMaster_v1",
"author": "digiplay",
"last_modified": "2024-03-16 00:22:41+00:00",
"downloads": 232,
"likes": 2,
"library_name": "diffusers",
"tags": [
"diffusers",
"safetensors",
"stable-diffusion",
"stable-diffusion-diffusers",
"text-to-image"… See the full description on the dataset page: https://huggingface.co/datasets/aravind-selvam/model_card_extractor_qwq.
