datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
c4-en-html-with-metadatac4-en-html-with-metadata-ppl-cleanFile list:
"c4-en-html_cc-main-2019-18_pq00-000.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-001.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-002.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-003.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-004.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-005.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-006.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-007.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-008.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-009.jsonl.gz"… See the full description on the dataset page: https://huggingface.co/datasets/masoudjs/c4-en-html-with-metadata-ppl-clean.c4-en-html-with-training_metadata_alldataset_cards_with_metadatamodel_cards_with_metadatarag-mini-bioasq-with-metadataThis dataset is an extension of the rag-mini-bioasq dataset.
Its difference resides in the text-corpus part of the aforementioned set where the metadata was added for each passage.
Metadata contains six separate categories, each in a dedicated column:
Year of the publication (publish_year)
Type of the publication (publish_type)
Country of the publication - often correlated with the homeland of the authors (country)
Number of pages (no_pages)
Authors (authors)
Keywords (keywords)
datasets_with_metadata_and_summariesdataset_cards_with_metadata_with_sample_rowsdataset_cards_with_metadata_with_embeddingsmodels_with_metadata_and_summariesdataset_cards_with_metadata_labelledmini-bioasq-with-metadataThis dataset is an extension of the rag-mini-bioasq dataset.
Its difference resides in the text-corpus part of the aforementioned set where the metadata was added for each passage.
Metadata contains six separate categories, each in a dedicated column:
Year of the publication (publish_year)
Type of the publication (publish_type)
Country of the publication - often correlated with the homeland of the authors (country)
Number of pages (no_pages)
Authors (authors)
Keywords (keywords)
subset-0-with-metadatamodel_cards_with_metadata_with_embeddings
Dataset Card for Hugging Face Hub Model Cards with Embeddings
This dataset consists of model cards for models hosted on the Hugging Face Hub. The model cards are created by the community and provide information about the model, its performance, its intended uses, and more.
This dataset is updated on a daily basis and includes publicly available models on the Hugging Face Hub.
This dataset is made available to help support users wanting to work with a large number of Model Cards… See the full description on the dataset page: https://huggingface.co/datasets/librarian-bots/model_cards_with_metadata_with_embeddings.dataset_cards_with_metadatamodel_cards_with_metadatasynthetic-kv-qwen3-8b-with-metadata
Synthetic KV Qwen3 8B — metadata-enhanced 64K
This dataset is an exact key-value retrieval benchmark. The context begins
with a short schema and task description, followed by records in the form
[KEY: VALUE]. Each question asks for the value belonging to one exact key.
The context is intentionally stored once in compact JSONL format. The
questions[i] entry corresponds to answers[i].
models_with_metadata_and_summaries_vllmbiomed-combined-with-metadata
Dataset Card for "biomed-combined-with-metadata"
More Information needed
AF-with-metaData
About
This dataset is an amalgamation of every entry of the DQ Round Trip Problem Selection spreadsheet (https://docs.google.com/spreadsheets/d/1dEWWzjuEXwf9s4II0CixH4sqopc1flIMFx19UjHiyNU/edit?usp=sharing&resourcekey=0-_G7oxmbh7szV5jx-HxhepQ)
with the addition of a Text column containing
Data Fields
link: A hyperlink to the source where the entry was pulled from
'formal statement': in Isabelle
text: The text entry contains the sample with the structure : The… See the full description on the dataset page: https://huggingface.co/datasets/UDACA/AF-with-metaData.meta_data_annual_reports_tokenized_llama3_8b_with_logged_return_matrix_dropped_null_returns
