datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
goes-imerg-42hour-testGERIS-Goes19-uruguay-firesgoes-imerg-42hourgoes-xray-flux
GOES Solar X-Ray Flux (1-Minute)
Credit: NASA/SDO
Part of a dataset collection on Hugging Face.
Dataset description
Solar soft X-ray flux from the GOES X-Ray Sensor (XRS), the operational backbone of solar flare monitoring. Updated daily from NOAA SWPC, growing incrementally at 1-minute cadence.
The GOES (Geostationary Operational Environmental Satellite) X-Ray Sensor measures the Sun's soft X-ray irradiance in two wavelength bands: a "short" 0.05-0.4 nm… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/goes-xray-flux.goes-imerg-6hourcs_csfd-movie-reviews
Dataset Card for CSFD movie reviews (Czech)
Dataset Description
The dataset contains user reviews from Czech/Slovak movie databse website https://csfd.cz.
Each review contains text, rating, date, and basic information about the movie (or TV series).
The dataset has in total (train+validation+test) 30,000 reviews. The data is balanced - each rating has approximately the same frequency.
Dataset Features
Each sample contains:
review_id: unique string identifier… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_csfd-movie-reviews.goes_omni_electron_flux_forecasting
GOES–OMNI >2 MeV Electron Flux Forecasting
This dataset combines cross-calibrated NOAA GOES-14/GOES-16 >2 MeV electron
flux with NASA/GSFC OMNI solar-wind and geomagnetic drivers on a uniform
five-minute UTC grid.
It provides two configurations:
ml-ready (default): scaled causal features, validity flags, unscaled
30-minute/6-hour/12-hour targets, and leakage-safe chronological splits.
scientific-master: unscaled source measurements, instrument context,
calibration factors, and… See the full description on the dataset page: https://huggingface.co/datasets/THULab/goes_omni_electron_flux_forecasting.goes-imerg-6hour-testgoes-imerg-6hour-valgoes-omni-electron-flux-forecasting
GOES–OMNI >2 MeV Electron Flux Forecasting
This dataset combines cross-calibrated NOAA GOES-14/GOES-16 >2 MeV electron
flux with NASA/GSFC OMNI solar-wind and geomagnetic drivers on a uniform
five-minute UTC grid.
It provides two configurations:
ml-ready (default): scaled causal features, validity flags, unscaled
30-minute/6-hour/12-hour targets, and leakage-safe chronological splits.
scientific-master: unscaled source measurements, instrument context,
calibration factors, and… See the full description on the dataset page: https://huggingface.co/datasets/snowsadh/goes-omni-electron-flux-forecasting.cs_czech-named-entity-corpus_2.0
Dataset Card for Czech Named Entity Corpus 2.0
Dataset Description
The dataset contains Czech sentences and annotated named entities. Total number of sentences is around 9,000 and total number of entities is around 34,000. (Total means train + validation + test)
Dataset Features
Each sample contains:
text: source sentence
entities: list of selected entities. Each entity contains:
category_id: string identifier of the entity category
category_str:… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_czech-named-entity-corpus_2.0.ultimate-life-ds20-nature-goes-strange
Ultimate Life — Nature Goes Strange
A public, U.S.-first, source-derived visual-grammar reference pack for MiniMax H3-compatible video study.
What is inside
clips/ — 20 silent H.264 MP4 clips and 20 paired .txt captions.
manifest.json — exact source timestamp, output hash, H3 contract, and clip-level restriction.
source.json — source URL, rights basis, preservation checksum, and probe data when available.
candidate-beats.json — selection ledger used to make the… See the full description on the dataset page: https://huggingface.co/datasets/TheMindExpansionNetwork/ultimate-life-ds20-nature-goes-strange.cs_squad-3.0
Dataset Card for Czech Simple Question Answering Dataset 3.0
This a processed and filtered adaptation of an existing dataset. For raw and larger dataset, see Dataset Source section.
Dataset Description
The data contains questions and answers based on Czech wikipeadia articles.
Each question has an answer (or more) and a selected part of the context as the evidence.
A majority of the answers are extractive - i.e. they are present in the context in the exact form. The… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_squad-3.0.sk_csfd-movie-reviews
Dataset Card for CSFD movie reviews (Slovak)
Dataset Description
The dataset contains user reviews from Czech/Slovak movie databse website https://csfd.cz.
Each review contains text, rating, date, and basic information about the movie (or TV series).
The dataset has in total (train+validation+test) 30,000 reviews. The data is balanced - each rating has approximately the same frequency.
Dataset Features
Each sample contains:
review_id: unique string identifier… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/sk_csfd-movie-reviews.cs_facebook-comments
Dataset Card for Czech Facebook comments
Dataset Description
The dataset contains user comments from Facebook. Each comment contains text, sentiment (positive/negative/neutral).
The dataset has in total (train+validation+test) 6,600 reviews. The data is balanced.
Dataset Features
Each sample contains:
comment_id: unique string identifier of the comment.
sentiment_str: string representation of the rating - "pozitivní" / "neutrální" / "negativní"
sentiment_int:… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_facebook-comments.cs_mall-product-reviews
Dataset Card for Mall.cz Product Reviews (Czech)
Dataset Description
The dataset contains user reviews from Czech eshop <mall.cz>
Each review contains text, sentiment (positive/negative/neutral), and automatically-detected language (mostly Czech, occasionaly Slovak) using lingua-py
The dataset has in total (train+validation+test) 30,000 reviews. The data is balanced.
Train set has 8000 positive, 8000 neutral and 8000 negative reviews.
Validation and test set each have… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_mall-product-reviews.cs_czech-court-decisions-ner
Dataset Card for Czech Court Decisions NER
Dataset Description
Czech Court Decisions NER is a dataset of 300 court decisions published by The Supreme Court of the Czech Republic and the Constitutional Court of the Czech Republic.
In the documents, 4 types of named entities are selected.
Dataset Features
Each sample contains:
filename: file name in the original dataset
text: court decision document in plain text
entities: list of selected entities. Each entity… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_czech-court-decisions-ner.GO_ESMFold_PDB
GO PDB Dataset
Github
Simple, Efficient and Scalable Structure-aware Adapter Boosts Protein Language Models
https://github.com/tyang816/SES-Adapter
VenusFactory: A Unified Platform for Protein Engineering Data Retrieval and Language Model Fine-Tuning
https://github.com/ai4protein/VenusFactory
Citation
Please cite our work if you use our dataset.
@article{tan2024ses-adapter,
title={Simple, Efficient, and Scalable Structure-Aware Adapter Boosts Protein Language… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/GO_ESMFold_PDB.solar-flare-goes-datasetThis dataset is intended to be used for training/testing solar flare forecasting models. It contains various data splits (in json format) of
GOES XRS time series (1 min-cadence) for two variables:
L2 flux/bkg ratio
Flare binary history (0=no flare, 1=flare)
Splits labelled as "_24h" correspond to a time series length of 24h, while those labelled as "_12h" to a length of 12h.
goes-imerg-42hour-valgoes-13-14-15goes-cloud-smoke-v0
