ddr
Datasets
All datasets matching “ddr”DDR-dataset
DDR - Diabetic Retinopathy Detection Dataset
Image: Dataset Samples.
The DDR (Diabetic Retinopathy Detection) dataset is a large-scale collection of retinal fundus images designed for training and evaluating algorithms in diabetic retinopathy (DR) grading and lesion-level segmentation. It provides both image-level DR labels and pixel-level annotations of pathological features, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/ctmedtech/DDR-dataset.super_eurlexSuper-EURLEX dataset containing legal documents from multiple languages.
The datasets are build/scrapped from the EURLEX Website [https://eur-lex.europa.eu/homepage.html]
With one split per language and sector, because the available features (metadata) differs for each
sector. Therefore, each sample contains the content of a full legal document in up to 3 different
formats. Those are raw HTML and cleaned HTML (if the HTML format was available on the EURLEX website
during the scrapping process) and cleaned text.
The cleaned text should be available for each sample and was extracted from HTML or PDF.
'Cleaned' HTML stands here for minor cleaning that was done to preserve to a large extent the necessary
HTML information like table structures while removing unnecessary complexity which was introduced to the
original documents due to actions like writing each sentence into a new object.
Additionally, each sample contains metadata which was scrapped on the fly, this implies the following
2 things. First, not every sector contains the same metadata. Second, most metadata might be
irrelevant for most use cases.
In our minds the most interesting metadata is the celex-id which is used to identify the legal
document at hand, but also contains a lot of information about the document
see [https://eur-lex.europa.eu/content/tools/eur-lex-celex-infographic-A3.pdf] as well as eurovoc-
concepts, which are labels that define the content of the documents.
Eurovoc-Concepts are, for example, only available for the sectors 1, 2, 3, 4, 5, 6, 9, C, and E.
The Naming of most metadata is kept like it was on the eurlex website, except for converting
it to lower case and replacing whitespaces with '_'.math_text
Mathematical Texts (MT)
Mathematical dataset containing mathematical texts, i.e. texts containing LaTeX formulas, based on the AMPS Khan dataset and the ARQMath dataset V1.3. Based on the retrieved LaTeX texts, more mathematically equivalent versions have been generated by applying randomized LaTeX printing with this SymPy fork using Math Mutator (MAMUT). A positive id corresponds to the ARQMath post id of the generated text version, a negative id indicates an AMPS text.
You can… See the full description on the dataset page: https://huggingface.co/datasets/ddrg/math_text.DDR-dataset
DDR - Diabetic Retinopathy Detection Dataset
Image: Dataset Samples.
The DDR (Diabetic Retinopathy Detection) dataset is a large-scale collection of retinal fundus images designed for training and evaluating algorithms in diabetic retinopathy (DR) grading and lesion-level segmentation. It provides both image-level DR labels and pixel-level annotations of pathological features, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/Aljo-na/DDR-dataset.DDR-dataset
DDR - Diabetic Retinopathy Detection Dataset
Image: Dataset Samples.
The DDR (Diabetic Retinopathy Detection) dataset is a large-scale collection of retinal fundus images designed for training and evaluating algorithms in diabetic retinopathy (DR) grading and lesion-level segmentation. It provides both image-level DR labels and pixel-level annotations of pathological features, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/chanchahh/DDR-dataset.math_formula_retrieval
Dataset Card for MFR (Mathematical Formula Retrieval)
This dataset consists of formula pairs, classified as either mathematical equivalent or not.
Dataset Details
Dataset Description
Mathematical dataset based on 71 famous mathematical identities. Each entry consists of two identities (in formula or textual form), together with a label, whether the two versions describe the same mathematical identity. The false pairs are not randomly chosen, but… See the full description on the dataset page: https://huggingface.co/datasets/ddrg/math_formula_retrieval.
