datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Our1-2b-Datasetno-name-dataset
no-name-dataset
Dataset Description
Authors: xxxxxxxx
Data source: ARTFL Encyclopédie Project, University of Chicago
Git repository: xxxxxxxx
Language: French
License: cc-by-nc-4.0
Dataset Summary
This dataset contains 2,750 labeled entries from the Encyclopédie of Diderot and d’Alembert, all classified under Geography.
The no-name-dataset provides a set of features used at different stages of our knowledge graph construction pipeline, which is illustrated… See the full description on the dataset page: https://huggingface.co/datasets/no-name-research/no-name-dataset.research-papers-dataset-mixtral7B-processed2
Research Papers Dataset - Processed with Train/Test/Valid Splits
This dataset contains preprocessed research papers with the following enhancements, split into train/test/validation sets.
Dataset Splits:
Train: 7,328 entries (85.0%)
Test: 431 entries (5.0%)
Valid: 863 entries (10.0%)
Preprocessing Applied:
Section Splitting: Papers are split into logical sections (Abstract, Introduction, Methods, Results, etc.)
Whitespace Normalization: Excessive whitespace… See the full description on the dataset page: https://huggingface.co/datasets/abhi26/research-papers-dataset-mixtral7B-processed2.research-papers-dataset-mixtral7B-processed
Research Papers Dataset - Processed
This dataset contains preprocessed research papers with the following enhancements:
Preprocessing Applied:
Section Splitting: Papers are split into logical sections (Abstract, Introduction, Methods, Results, etc.)
Whitespace Normalization: Excessive whitespace removed and normalized
Punctuation Fixing: Missing spaces after punctuation marks corrected
Sentence Boundary Fixing: Proper sentence boundaries established
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/abhi26/research-papers-dataset-mixtral7B-processed.
