datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
oasismri-oasis-1-ixi-pre
OASIS-1 + IXI Preprocessed MRI
Preprocessed T1-weighted brain MRI volumes from two widely-used public neuroimaging datasets,
ready for machine learning and 3-D visualisation.
Volumes are stored as float32 NumPy arrays of shape (96 × 128 × 96) (D × H × W),
so loading is as simple as np.load("OAS1_0001_MR1_t1.npy").
Datasets
OASIS-1 — Open Access Series of Imaging Studies
Property
Value
Subjects
436 (ages 18–96)
Population
Cognitively… See the full description on the dataset page: https://huggingface.co/datasets/pzarzycki/mri-oasis-1-ixi-pre.oasisoasis-distillation-dataset-3dOasis
Oasis: One Image is All You Need for Multimodal Instruction Data Synthesis
This dataset contains Oasis-500k dataset.
[Read the Paper] | [Github Repo]
All images come from Cambrian-10M. Instructions and responses are generated by MLLM.
oasis-siting-cacheLLM-Oasis_unfactual_text_generation
Babelscape/LLM-Oasis_unfactual_text_generation
Dataset Description
LLM-Oasis_unfactual_text_generation is part of the LLM-Oasis suite and contains unfactual texts generated from a set of falsified claims extracted from a Wikipedia passage and its paraphrase.
This dataset corresponds to the unfactual text generation step described in Section 3.4 of the LLM-Oasis paper. Please refer to our GitHub repository for more information on the overall data generation pipeline of… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/LLM-Oasis_unfactual_text_generation.telly-oasis-1Telly-Oasis-1-dataset
A high-quality, multi-turn conversational dataset designed for LLM fine-tuning and instruction-following.This dataset is used to train the Telly Oasis 1 language model Which is curently under development
Telly Oasis 1 is a conversational instruction dataset stored in JSONL format. It is designed for training, fine-tuning, and evaluating Large Language Models (LLMs), particularly chat and instruction-following models.
The dataset contains multi-turn conversations between a… See the full description on the dataset page: https://huggingface.co/datasets/Trytellypls/telly-oasis-1.LLM-Oasis_paraphrase_generation
Babelscape/LLM-Oasis_paraphrase_generation
Dataset Description
LLM-Oasis_paraphrase_generation is part of the LLM-Oasis suite and contains paraphrases generated from a set of claims extracted from a Wikipedia passage.
This dataset supports the paraphrase generation step described in Section 3.3 of the LLM-Oasis paper. Please refer to our GitHub repository for more information on the overall data generation pipeline of LLM-Oasis.
Features
title: The title… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/LLM-Oasis_paraphrase_generation.oasis-datasetOASIS
OASIS: A Multilingual and Multimodal Dataset for Culturally Grounded Spoken Visual QA
Dataset Description
OASIS is a large-scale culturally grounded multimodal question answering dataset covering images, text, and speech. It is designed to evaluate multimodal models beyond object recognition, with emphasis on pragmatic, commonsense, and culturally grounded reasoning in real-world scenarios.
Large-scale multimodal models achieve strong results on tasks such as Visual… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/OASIS.trifuse-ad-oasis1
TriFuse-AD: Honest Multimodal Benchmark for Three-Stage Dementia Staging on OASIS-1
Code, processed data, results, and paper for a leakage-free benchmark of three-stage
cognitive classification (CN / VMD / AD) on the OASIS-1 cross-sectional cohort, plus
the proposed TriFuse-AD model (tri-planar CNN + slice-plane Transformer + gated
demographic fusion).
Key result (honest / negative)
On an age-restricted cohort (≥60, 198 subjects) with subject-level repeated 5-fold… See the full description on the dataset page: https://huggingface.co/datasets/minhy112/trifuse-ad-oasis1.LLM-Oasis_claim_falsification
Babelscape/LLM-Oasis_claim_falsification
Dataset Description
LLM-Oasis_claim_falsification is part of the LLM-Oasis suite and contains the outcomes of the claim falsification process.
This dataset provides pairs of factual and falsified claims from a given Wikipedia text as described in Section 3.2 of the LLM-Oasis paper. Please refer to our GitHub repository for more information on the overall data generation pipeline of LLM-Oasis.
Features
title: The… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/LLM-Oasis_claim_falsification.oasis-alzheimers-multi-classLLM-Oasis_claim_extraction
Babelscape/LLM-Oasis_claim_extraction
Dataset Description
LLM-Oasis_claim_extraction is part of the LLM-Oasis suite and contains text-claim pairs extracted from Wikipedia pages.
It provides the data used to train the claim extraction system described in Section 3.1 of the LLM-Oasis paper. Please refer to our GitHub repository for more information on the overall data generation pipeline of LLM-Oasis.
Features
title: The title of the Wikipedia page.
text: A… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/LLM-Oasis_claim_extraction.oasisLLM-Oasis_claim_verification
Babelscape/LLM-Oasis_claim_verification
Dataset Description
LLM-Oasis_claim_verification is part of the LLM-Oasis suite and contains the gold-standard dataset for verifying the veracity of individual claims against provided evidence.
This dataset supports the claim verification task described in Section 4.2 of the LLM-Oasis paper. Please refer to our GitHub repository for additional information on the LLM-Oasis data generation pipeline.
Features
id: A… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/LLM-Oasis_claim_verification.oasisThis dataset is designed to generate lyrics with HuggingArtists.oasis1Open-Oasis-Partial-Docs
Open Oasis Partial Docs Crawl
A partial crawl of https://docs.oasis-open.org/.
Whole thing couldnt be crawled because it'd take too long, it's surprisingly large. I don't know how useful this will be, it's probably a small fraction of the actual size, as of now its 19.63GB in this TAR archive.
Oasis-Corpus
Dataset Card for Oasis-Corpus
Dataset Description
Oasis-Corpus is a 783GB high-quality bilingual corpus.
All data in Oasis-Corpus are built by Oasis and sourced from Common Crawl.
It consists of 374GB of Chinese from 17 recent dumps and 409GB of English textual data from 5 dumps.
Languages
English(409GB, 70,121,125 lines) and Chinese(374GB, 110,580,964 lines)
Data Splits
Language
Dump
docs
size
Chinese
cc-may-jun-2023-zh
5,627,020
19.31 GB… See the full description on the dataset page: https://huggingface.co/datasets/Oasis-Team/Oasis-Corpus.oasissimpdataset
OasisSimp: An Open-source Asian-English Sentence Simplification Dataset
task_categories:
- text-generation
language:
- en
- si
- ps
- ta
- th
OasisSimp Dataset (https://oasissimpdataset.github.io/)
Each language has its own folder containing two JSONL files corresponding to valid and test set.
Directory Structure
Folder contains:
OasisSimp-EN/
├── OasisSimp-EN_valid.jsonl
├── OasisSimp-EN_test.jsonl
OasisSimp-PS/
├── OasisSimp-PS_valid.jsonl
├──… See the full description on the dataset page: https://huggingface.co/datasets/ravishekhar/oasissimpdataset.africa-tunisia-oasis-evolution-des-superficies-oasis-en-hectare-4c7af0f0
Oasis Evolution Des Superficies Oasis En Hectare | Africa (Tunisia Open Data)
64 rows - 1 Africa country/area - 1980-2016 - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 64 rows from Tunisia Open Data, covering Oasis Evolution Des Superficies Oasis En Hectare. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-tunisia-oasis-evolution-des-superficies-oasis-en-hectare-4c7af0f0.africa-tunisia-oasis-evolution-de-la-production-des-dattes-1000-t-d8981b70
Oasis Evolution De La Production Des Dattes 1000 T | Africa (Tunisia Open Data)
44 rows - 1 Africa country/area - detected - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 44 rows from Tunisia Open Data, covering Oasis Evolution De La Production Des Dattes 1000 T. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-tunisia-oasis-evolution-de-la-production-des-dattes-1000-t-d8981b70.OASISOasisLyricsOasisSplitedQuazim0t0__Oasis-14B-ties-details
Dataset Card for Evaluation run of Quazim0t0/Oasis-14B-ties
Dataset automatically created during the evaluation run of model Quazim0t0/Oasis-14B-ties
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Quazim0t0__Oasis-14B-ties-details.OASIS-2-axialcours-oasis-de-cergy
Cours Oasis de Cergy
[!NOTE]
Ce jeu de données Hugging Face est vide. Cette carte sert seulement à référencer le jeu de données Cours Oasis de Cergy qui est disponible à l'adresse https://www.data.gouv.fr/datasets/676d62090d3bb78fba302778
Description
Emplacement des cours Oasis sur Cergy (seulement au groupe scolaire de la Justice actuellement)
Ces données peuvent être visibles sur l'application cartographique de la Ville de Cergy : Cergy côté Nature
Jeu de données… See the full description on the dataset page: https://huggingface.co/datasets/french-open-data/cours-oasis-de-cergy.
