datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ptb-cmbcobe-firas-cmb-monopole
COBE/FIRAS CMB monopole spectrum
This dataset contains the 43-point COBE/FIRAS cosmic microwave background
monopole spectrum published by LAMBDA. It includes the monopole intensity, its
residual from a 2.725 K blackbody, the published one-sigma uncertainty, and a
model of Galactic emission at the Galactic poles. The columns retain the
source file's own labels, Column 1 through Column 5; their meanings and
units are documented below.
Mission
COBE
Instrument
FIRAS… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/cobe-firas-cmb-monopole.lengthyflanv2m0_fine_tuning_ocr_ptrn_cmbert_io
m0_fine_tuning_ocr_ptrn_cmbert_io
Introduction
This dataset was used to fine-tuned HueyNemud/das22-10-camembert_pretrained for flat NER task using Flat NER approach [M0].
It contains 19th-century Paris trade directories' entries.
Dataset parameters
Approach : M0
Dataset type : noisy (Pero OCR)
Tokenizer : HueyNemud/das22-10-camembert_pretrained
Tagging format : IO
Counts :
Train : 6084
Dev : 676
Test : 1685
Associated fine-tuned model :… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m0_fine_tuning_ocr_ptrn_cmbert_io.CMB-Exam-Grouped
CMB-Exam-Grouped
This dataset contains medical exam questions with shared background context extracted and grouped together.
Features
question_id: Unique identifier for each question
background_id: ID for grouping questions that share the same background (-1 means no shared background)
background: Shared context/background text extracted from questions
question: The actual question (with background removed if applicable)
option: Multiple choice options (A-F)
answer:… See the full description on the dataset page: https://huggingface.co/datasets/fzkuji/CMB-Exam-Grouped.m0_fine_tuning_ocr_cmbert_io
m0_fine_tuning_ocr_cmbert_io
Introduction
This dataset was used to fine-tuned Jean-Baptiste/camembert-ner for flat NER task using Flat NER approach [M0].
It contains 19th-century Paris trade directories' entries.
Dataset parameters
Approach : M0
Dataset type : noisy (Pero OCR)
Tokenizer : Jean-Baptiste/camembert-ner
Tagging format : IO
Counts :
Train : 6084
Dev : 676
Test : 1685
Associated fine-tuned model : nlpso/m0_flat_ner_ocr_cmbert_io… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m0_fine_tuning_ocr_cmbert_io.m2m3_fine_tuning_ref_ptrn_cmbert_io
m2m3_fine_tuning_ref_ptrn_cmbert_io
Introduction
This dataset was used to fine-tuned HueyNemud/das22-10-camembert_pretrained for nested NER task using Independant NER layers approach [M1].
It contains Paris trade directories entries from the 19th century.
Dataset parameters
Approachrd : M2 and M3
Dataset type : ground-truth
Tokenizer : HueyNemud/das22-10-camembert_pretrained
Tagging format : IO
Counts :
Train : 6084
Dev : 676
Test : 1685
Associated… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_fine_tuning_ref_ptrn_cmbert_io.m0_fine_tuning_ref_cmbert_io
m0_fine_tuning_ref_cmbert_io
Introduction
This dataset was used to fine-tuned Jean-Baptiste/camembert-ner for flat NER task using Flat NER approach [M0].
It contains 19th-century Paris trade directories' entries.
Dataset parameters
Approach : M0
Dataset type : ground-truth
Tokenizer : Jean-Baptiste/camembert-ner
Tagging format : IO
Counts :
Train : 6084
Dev : 676
Test : 1685
Associated fine-tuned model : nlpso/m0_flat_ner_ref_cmbert_io
Entity… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m0_fine_tuning_ref_cmbert_io.cobe-firas-cmb-temperature-map
COBE/FIRAS CMB Temperature Map
This dataset contains LAMBDA's 6,067-row CMB Temperature Map FITS table. The source
fields remain PIXEL, GAL_LON, GAL_LAT, WEIGHT, TEMP, TEMP_SIG,
and RESID_TE, in order and with their source values and units unchanged.
How to use
python -m venv .venv && .venv/bin/pip install datasets huggingface_hub pyarrow
from datasets import load_dataset
dataset = load_dataset("astro-legacy-archive/cobe-firas-cmb-temperature-map"… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/cobe-firas-cmb-temperature-map.m1_qualitative_analysis_ref_cmbert_io
m1_qualitative_analysis_ref_cmbert_io
Introduction
This dataset was used to perform qualitative analysis of Jean-Baptiste/camembert-ner on nested NER task using Independant NER layers approach [M1].
It contains Paris trade directories entries from the 19th century.
Dataset parameters
Approach : M1
Dataset type : ground-truth
Tokenizer : Jean-Baptiste/camembert-ner
Tagging format : IO
Counts :
Train : 6084
Dev : 676
Test : 1685
Associated fine-tuned models… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m1_qualitative_analysis_ref_cmbert_io.m1_qualitative_analysis_ocr_cmbert_iob2
m1_qualitative_analysis_ocr_cmbert_iob2
Introduction
This dataset was used to perform qualitative analysis of Jean-Baptiste/camembert-ner on nested NER task using Independant NER layers approach [M1].
It contains Paris trade directories entries from the 19th century.
Dataset parameters
Approach : M1
Dataset type : noisy (Pero OCR)
Tokenizer : Jean-Baptiste/camembert-ner
Tagging format : IOB2
Counts :
Train : 6084
Dev : 676
Test : 1685
Associated fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m1_qualitative_analysis_ocr_cmbert_iob2.m2m3_qualitative_analysis_ref_cmbert_io
m2m3_qualitative_analysis_ref_cmbert_io
Introduction
This dataset was used to perform qualitative analysis of Jean-Baptiste/camembert-ner on nested NER task using Independant NER layers approach [M1].
It contains Paris trade directories entries from the 19th century.
Dataset parameters
Approachrd : M2 and M3
Dataset type : ground-truth
Tokenizer : Jean-Baptiste/camembert-ner
Tagging format : IO
Counts :
Train : 6084
Dev : 676
Test : 1685
Associated… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_qualitative_analysis_ref_cmbert_io.CosmoPaperQACosmoPaperQA is an evaluation dataset, designed to serve as a benchmark for RAG applications with cosmology papers, with an emphasis on testing for fetching, interprative and explanatory skills.
The information sources for the questions and ideal answers are these 5 papers:
Planck Collaboration, Planck 2018 results. VI. Cosmological parameters, Astron.Astrophys. 641 (2020) A6
Villaescusa-Navarro et al., The CAMELS project: Cosmology and Astrophysics with MachinE Learning Simulations… See the full description on the dataset page: https://huggingface.co/datasets/cmbagent/CosmoPaperQA.cre-cmbs-loans
Commercial Real Estate Loans (CMBS, SEC ABS-EE)
2,313 commercial properties securitised across 22 CMBS trusts, assembled from
the SEC's Regulation AB Exhibit 102 asset-level filings.
Each row is one property backing a commercial mortgage: name, address, type,
size, occupancy, largest tenant, the loan against it, the appraisal, and the
property's revenue, expenses and net operating income. LTV and debt yield are
derived from those.
Provenance
Source filings are SEC… See the full description on the dataset page: https://huggingface.co/datasets/catyung/cre-cmbs-loans.m2m3_qualitative_analysis_ocr_ptrn_cmbert_io
m2m3_qualitative_analysis_ocr_ptrn_cmbert_io
Introduction
This dataset was used to perform qualitative analysis of HueyNemud/das22-10-camembert_pretrained on nested NER task using Independant NER layers approach [M1].
It contains Paris trade directories entries from the 19th century.
Dataset parameters
Approachrd : M2 and M3
Dataset type : noisy (Pero OCR)
Tokenizer : HueyNemud/das22-10-camembert_pretrained
Tagging format : IO
Counts :
Train : 6084
Dev : 676… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_qualitative_analysis_ocr_ptrn_cmbert_io.m2m3_qualitative_analysis_ocr_cmbert_io
m2m3_qualitative_analysis_ocr_cmbert_io
Introduction
This dataset was used to perform qualitative analysis of Jean-Baptiste/camembert-ner on nested NER task using Independant NER layers approach [M1].
It contains Paris trade directories entries from the 19th century.
Dataset parameters
Approachrd : M2 and M3
Dataset type : noisy (Pero OCR)
Tokenizer : Jean-Baptiste/camembert-ner
Tagging format : IO
Counts :
Train : 6084
Dev : 676
Test : 1685
Associated… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_qualitative_analysis_ocr_cmbert_io.CMB-flinksql-24801
Dataset Card for Dataset Name
CMB-flinksql-24801 containes 24,801 training samples
Domain Knowledge Collection
Clean and curate Flink SQL–related content from the Internet, including official documentation, tutorials, technical articles, blogs, and other high-quality resources, to build comprehensive domain knowledge.
Synthetic Data Generation via Qwen3-Max Distillation
Distill high-quality FlinkSQL-to-Natural Language (FlinkSQL2NL) and Natural Language-to-FlinkSQL… See the full description on the dataset page: https://huggingface.co/datasets/CMBTech/CMB-flinksql-24801.m1_fine_tuning_ref_cmbert_iob2
m1_fine_tuning_ref_cmbert_iob2
Introduction
This dataset was used to fine-tuned Jean-Baptiste/camembert-ner for nested NER task using Independant NER layers approach [M1].
It contains Paris trade directories entries from the 19th century.
Dataset parameters
Approach : M1
Dataset type : ground-truth
Tokenizer : Jean-Baptiste/camembert-ner
Tagging format : IOB2
Counts :
Train : 6084
Dev : 676
Test : 1685
Associated fine-tuned models :
Level-1 :… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m1_fine_tuning_ref_cmbert_iob2.m2m3_fine_tuning_ref_cmbert_io
m2m3_fine_tuning_ref_cmbert_io
Introduction
This dataset was used to fine-tuned Jean-Baptiste/camembert-ner for nested NER task using Independant NER layers approach [M1].
It contains Paris trade directories entries from the 19th century.
Dataset parameters
Approachrd : M2 and M3
Dataset type : ground-truth
Tokenizer : Jean-Baptiste/camembert-ner
Tagging format : IO
Counts :
Train : 6084
Dev : 676
Test : 1685
Associated fine-tuned models :
M2 :… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_fine_tuning_ref_cmbert_io.m2m3_fine_tuning_ocr_ptrn_cmbert_iob2
m2m3_fine_tuning_ocr_ptrn_cmbert_iob2
Introduction
This dataset was used to fine-tuned HueyNemud/das22-10-camembert_pretrained for nested NER task using Independant NER layers approach [M1].
It contains Paris trade directories entries from the 19th century.
Dataset parameters
Approachrd : M2 and M3
Dataset type : noisy (Pero OCR)
Tokenizer : HueyNemud/das22-10-camembert_pretrained
Tagging format : IOB2
Counts :
Train : 6084
Dev : 676
Test : 1685
Associated… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_fine_tuning_ocr_ptrn_cmbert_iob2.m2m3_qualitative_analysis_ref_cmbert_iob2
m2m3_qualitative_analysis_ref_cmbert_iob2
Introduction
This dataset was used to perform qualitative analysis of Jean-Baptiste/camembert-ner on nested NER task using Independant NER layers approach [M1].
It contains Paris trade directories entries from the 19th century.
Dataset parameters
Approachrd : M2 and M3
Dataset type : ground-truth
Tokenizer : Jean-Baptiste/camembert-ner
Tagging format : IOB2
Counts :
Train : 6084
Dev : 676
Test : 1685
Associated… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_qualitative_analysis_ref_cmbert_iob2.m1_fine_tuning_ocr_cmbert_io
m1_fine_tuning_ocr_cmbert_io
Introduction
This dataset was used to fine-tuned Jean-Baptiste/camembert-ner for nested NER task using Independant NER layers approach [M1].
It contains Paris trade directories entries from the 19th century.
Dataset parameters
Approach : M1
Dataset type : noisy (Pero OCR)
Tokenizer : Jean-Baptiste/camembert-ner
Tagging format : IO
Counts :
Train : 6084
Dev : 676
Test : 1685
Associated fine-tuned models :
Level-1 :… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m1_fine_tuning_ocr_cmbert_io.m2m3_qualitative_analysis_ref_ptrn_cmbert_iob2
m2m3_qualitative_analysis_ref_ptrn_cmbert_iob2
Introduction
This dataset was used to perform qualitative analysis of HueyNemud/das22-10-camembert_pretrained on nested NER task using Independant NER layers approach [M1].
It contains Paris trade directories entries from the 19th century.
Dataset parameters
Approachrd : M2 and M3
Dataset type : ground-truth
Tokenizer : HueyNemud/das22-10-camembert_pretrained
Tagging format : IOB2
Counts :
Train : 6084
Dev : 676… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_qualitative_analysis_ref_ptrn_cmbert_iob2.m1_qualitative_analysis_ref_cmbert_iob2
m1_qualitative_analysis_ref_cmbert_iob2
Introduction
This dataset was used to perform qualitative analysis of Jean-Baptiste/camembert-ner on nested NER task using Independant NER layers approach [M1].
It contains Paris trade directories entries from the 19th century.
Dataset parameters
Approach : M1
Dataset type : ground-truth
Tokenizer : Jean-Baptiste/camembert-ner
Tagging format : IOB2
Counts :
Train : 6084
Dev : 676
Test : 1685
Associated fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m1_qualitative_analysis_ref_cmbert_iob2.m1_fine_tuning_ref_cmbert_io
m1_fine_tuning_ref_cmbert_io
Introduction
This dataset was used to fine-tuned Jean-Baptiste/camembert-ner for nested NER task using Independant NER layers approach [M1].
It contains Paris trade directories entries from the 19th century.
Dataset parameters
Approach : M1
Dataset type : ground-truth
Tokenizer : Jean-Baptiste/camembert-ner
Tagging format : IO
Counts :
Train : 6084
Dev : 676
Test : 1685
Associated fine-tuned models :
Level-1 :… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m1_fine_tuning_ref_cmbert_io.m0_qualitative_analysis_ocr_cmbert_io
m0_qualitative_analysis_ocr_cmbert_io
Introduction
This dataset was used to perform qualitative analysis of Jean-Baptiste/camembert-ner on flat NER task using Flat NER approach [M0].
It contains 19th-century Paris trade directories' entries.
Dataset parameters
Approach : M0
Dataset type : noisy (Pero OCR)
Tokenizer : Jean-Baptiste/camembert-ner
Tagging format : IO
Counts :
Train : 6084
Dev : 676
Test : 1685
Associated fine-tuned model :… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m0_qualitative_analysis_ocr_cmbert_io.m1_fine_tuning_ocr_ptrn_cmbert_iob2
m1_fine_tuning_ocr_ptrn_cmbert_iob2
Introduction
This dataset was used to fine-tuned HueyNemud/das22-10-camembert_pretrained for nested NER task using Independant NER layers approach [M1].
It contains Paris trade directories entries from the 19th century.
Dataset parameters
Approach : M1
Dataset type : noisy (Pero OCR)
Tokenizer : HueyNemud/das22-10-camembert_pretrained
Tagging format : IOB2
Counts :
Train : 6084
Dev : 676
Test : 1685
Associated fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m1_fine_tuning_ocr_ptrn_cmbert_iob2.m0_qualitative_analysis_ocr_ptrn_cmbert_io
m0_qualitative_analysis_ocr_ptrn_cmbert_io
Introduction
This dataset was used to perform qualitative analysis of HueyNemud/das22-10-camembert_pretrained on flat NER task using Flat NER approach [M0].
It contains 19th-century Paris trade directories' entries.
Dataset parameters
Approach : M0
Dataset type : noisy (Pero OCR)
Tokenizer : HueyNemud/das22-10-camembert_pretrained
Tagging format : IO
Counts :
Train : 6084
Dev : 676
Test : 1685
Associated… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m0_qualitative_analysis_ocr_ptrn_cmbert_io.m1_fine_tuning_ocr_cmbert_iob2
m1_fine_tuning_ocr_cmbert_iob2
Introduction
This dataset was used to fine-tuned Jean-Baptiste/camembert-ner for nested NER task using Independant NER layers approach [M1].
It contains Paris trade directories entries from the 19th century.
Dataset parameters
Approach : M1
Dataset type : noisy (Pero OCR)
Tokenizer : Jean-Baptiste/camembert-ner
Tagging format : IOB2
Counts :
Train : 6084
Dev : 676
Test : 1685
Associated fine-tuned models :
Level-1 :… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m1_fine_tuning_ocr_cmbert_iob2.m2m3_fine_tuning_ocr_ptrn_cmbert_io
m2m3_fine_tuning_ocr_ptrn_cmbert_io
Introduction
This dataset was used to fine-tuned HueyNemud/das22-10-camembert_pretrained for nested NER task using Independant NER layers approach [M1].
It contains Paris trade directories entries from the 19th century.
Dataset parameters
Approachrd : M2 and M3
Dataset type : noisy (Pero OCR)
Tokenizer : HueyNemud/das22-10-camembert_pretrained
Tagging format : IO
Counts :
Train : 6084
Dev : 676
Test : 1685
Associated… See the full description on the dataset page: https://huggingface.co/datasets/nlpso/m2m3_fine_tuning_ocr_ptrn_cmbert_io.
