datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fineweb-edu
FineWeb-Edu (Lance Format)
A Lance-formatted version of FineWeb-Edu — over 1.5 billion educational web passages with cleaned text, source metadata, language detection signals, and 384-dim text embeddings — available directly from the Hub at hf://datasets/lance-format/fineweb-edu/data/train.lance.
Key features
Cleaned passage text in the text column with the source url and title carried alongside.
Language detection signals (language, language_probability) for filtered… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/fineweb-edu.LLaVA-OneVision-1.5-Instruct-Data-qwen-formatformat-textformat-jsonformat-jsonllogiqa2_formatted
Dataset Card for "logiqa2_formatted"
More Information needed
format-listjsonfiltering-pretraining-mix-arrow-formatOpenvid-1M
OpenVid Dataset (Lance Format)
Lance format version of the OpenVid dataset with 937,957 high-quality videos stored with inline video blobs, embeddings, and rich metadata.
Why Lance?
Lance is an open-source format designed for multimodal AI data, offering significant advantages over traditional formats for modern AI workloads.
Blazing Fast Random Access: Optimized for fetching scattered rows, making it ideal for random sampling, real-time ML serving, and interactive… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/Openvid-1M.doc-formats-jsonl-1
[doc] formats - jsonl - 1
This dataset contains one jsonl file at the root.
chat_formatted_examplesdoc-formats-parquet-1doc-formats-csv-1
[doc] formats - csv - 1
This dataset contains one csv file at the root:
data.csv
kind,sound
dog,woof
cat,meow
pokemon,pika
human,hello
The YAML section of the README does not contain anything related to loading the data (only the size category metadata):
---
size_categories:
- n<1K
---
glaive-function-calling-v2-formatted
original dataset: glaiveai/glaive-function-calling-v2
{'system_message': 'You are a helpful assistant with access to the following functions. Use them if required -',
'function_description': '{\n "name": "get_random_quote",\n "description": "Get a random quote",\n "parameters": {}\n}',
'conversations': [{'content': 'Hi, can you help me with something?',
'role': 'user'},
{'content': "Of course! I'm here to assist you. What do you need help with?",
'role': 'assistant'}… See the full description on the dataset page: https://huggingface.co/datasets/heegyu/glaive-function-calling-v2-formatted.Krea-2-Turbo-Checkpoint-Format-Benchmark
Krea 2 Turbo ComfyUI Format Fidelity Benchmark
This release is a paired, deterministic comparison of eight Krea 2 Turbo checkpoint formats in ComfyUI: BF16, FP8 Scaled, INT8 ConvRot, MXFP8, NVFP4, INT4 ConvRot W4A4, GGUF Q8_0, and GGUF Q4_K_M. It contains 240 scored 1024×1024 images, saved float32 decoded tensors and final latents, every denoising trajectory, raw metric tables, telemetry, statistical comparisons, and reproduction code.
Main result
BF16 is the… See the full description on the dataset page: https://huggingface.co/datasets/Merserk/Krea-2-Turbo-Checkpoint-Format-Benchmark.text-dataset-tiny-code-script-py-format
USED of tahamajs/medicine_ds_persian for .parquet file
USED of Alijafarixcs2/persian-it-llama2-2k for .parquet file
USED of Abirate/english_quotes for .jsonl file
NEW FILES (05/12/2025)
NEW FILES (12/26/2025)
NEW FILES (02/15/2026)
fineweb-edu-format-topic
FineWeb-Edu w/ Topic and Format Annotations
FineWeb-Edu dataset consists of 1.3T tokens annotated for Topic and Format using wissamantoun/WebOrganizer-TopicClassifier-ModernBERT and wissamantoun/WebOrganizer-FormatClassifier-ModernBERT classifiers.
Similar to WebOrganizer/Corpus-200B but using FineEdu instead of DCLM.
Topic Labels:
Adult
Art & Design
Software Dev.
Crime & Law
Education & Jobs
Hardware
Entertainment
Social Life
Fashion & Beauty
Finance & Business
Food & Dining… See the full description on the dataset page: https://huggingface.co/datasets/wissamantoun/fineweb-edu-format-topic.openvid-lance
OpenVid (Lance Format)
A Lance-formatted version of the OpenVid-1M corpus — 937,957 high-quality clips with inline MP4 bytes, 1024-dim video embeddings, captions, and rich per-clip quality signals — available directly from the Hub at hf://datasets/lance-format/openvid-lance/data/train.lance.
Key features
Inline MP4 bytes in the video_blob column, stored in a side blob file and surfaced as lazy BlobFile handles via take_blobs — metadata scans, search, and filtering… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/openvid-lance.repo_format3
Dataset: repo_format3
This dataset was uploaded from /mnt/yulan_pretrain/mount/data_final_train/repo_format3/no-curriculum/tmp.
coda-lm-llava-format
CODA-LM Dataset Card
CODA-LM is the multi-modal version of the CODA dataset, used in the CODA-LM paper. Both English and Chinese annotations are available. Check detailed usage in our Github repo.
This repo contains the CODA-LM dataset, which has been reorganized in the LLaVA data format.
You are also welcome to check the original CODA-LM data which contains more metadata vanilla annotations.
Usage
from datasets import load_dataset
# name can be selected from… See the full description on the dataset page: https://huggingface.co/datasets/KaiChen1998/coda-lm-llava-format.Rexverse-2M-formattedglaive-function-calling-v2-formatted
Dataset Card for "glaive-function-calling-v2-formatted"
More Information needed
hle-no-img-prompt-completion-formatmath-hendrycks-genesys-formatpico-banana-smolvlm-format-with-rejected-answer
pico-banana-smolvlm-format-with-rejected-answer
Balanced image-level tampering detection dataset in SmolVLM-style format
with chosen/rejected answer pairs, derived from the pico-banana MCQ
pipeline. Suitable for preference learning (e.g. DPO) and RLHF-style training.
Dataset overview
Same as vanloc1808/pico-banana-smolvlm-format, but each example includes a
rejected_answer field: the answer from the counterpart sample (same
edited/original image pair, opposite… See the full description on the dataset page: https://huggingface.co/datasets/vanloc1808/pico-banana-smolvlm-format-with-rejected-answer.gpqa_formatted
Dataset Card for GPQA
Formatted version of original GPQA dataset. This removes most columns and adds single columns options and answer to contain a list of the possible answers and the index of the correct one.
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy… See the full description on the dataset page: https://huggingface.co/datasets/jeggers/gpqa_formatted.vidore_v3_finance_en_mteb_format
Vidore3FinanceEnRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
Retrieve associated pages according to questions.
Task category
t2i
Domains
Academic
Reference
https://huggingface.co/blog/QuentinJG/introducing-vidore-v3
Source datasets:
vidore/vidore_v3_finance_en
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_task("Vidore3FinanceEnRetrieval")
evaluator… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_finance_en_mteb_format.vidore_v3_computer_science_mteb_format
Vidore3ComputerScienceRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
Retrieve associated pages according to questions.
Task category
t2i
Domains
Academic
Reference
https://huggingface.co/blog/QuentinJG/introducing-vidore-v3
Source datasets:
vidore/vidore_v3_computer_science
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_computer_science_mteb_format.vidore_v3_industrial_mteb_format
Vidore3IndustrialRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
Retrieve associated pages according to questions.
Task category
t2i
Domains
Academic
Reference
https://huggingface.co/blog/QuentinJG/introducing-vidore-v3
Source datasets:
vidore/vidore_v3_industrial
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_task("Vidore3IndustrialRetrieval")… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_industrial_mteb_format.droid
