datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wavelet-lstm-camels-models
Wavelet-LSTM CAMELS Streamflow Models
A collection of 61,380 pre-trained LSTM models for daily streamflow forecasting across 620 USGS catchments from the CAMELS dataset.
Each catchment has 99 independently trained models:
33 wavelet filters × 3 lead times (1, 3, 5 days) = 99 wavelet-enhanced models
33 matching baseline models (same architecture, no wavelet transform)
Models are designed to be ensembled across wavelets for robust predictions with uncertainty estimates.… See the full description on the dataset page: https://huggingface.co/datasets/johnswyou/wavelet-lstm-camels-models.W_LSTMix_test_datasetp2-etf-x-lstm-extended-resultsLsTSCP_Media_Vol_3LsTSCP_Media_Vol_2LsTSCP_Media_Vol_1RLM-Evals
RLM Evals
Curated evaluation bundle for comparing RLM policies against the benchmark family used in the Recursive Language Models paper.
The repository stores unsampled benchmark subsets. Sampling for a specific experiment should be done downstream with a fixed seed and recorded in the eval manifest.
Subsets
longbench_v2_codeqa: 50 rows. Source: zai-org/LongBench-v2 train filtered to Code Repository Understanding / Code repo QA.
browsecomp_plus: 830 rows. Source:… See the full description on the dataset page: https://huggingface.co/datasets/lsteno/RLM-Evals.lst20LST20 Corpus is a dataset for Thai language processing developed by National Electronics and Computer Technology Center (NECTEC), Thailand.
It offers five layers of linguistic annotation: word boundaries, POS tagging, named entities, clause boundaries, and sentence boundaries.
At a large scale, it consists of 3,164,002 words, 288,020 named entities, 248,181 clauses, and 74,180 sentences, while it is annotated with
16 distinct POS tags. All 3,745 documents are also annotated with one of 15 news genres. Regarding its sheer size, this dataset is
considered large enough for developing joint neural models for NLP.
Manually download at https://aiforthai.in.th/corpus.phpThaiQA_LST20SuperAI Engineer Season 2 , Machima
Machima_ThaiQA_LST20 เป็นชุดข้อมูลที่สกัดหาคำถาม และคำตอบ จากบทความในชุดข้อมูล LST20 โดยสกัดได้คำถาม-ตอบทั้งหมด 7,642 คำถาม มีข้อมูล 4 คอลัมน์ ประกอบด้วย context, question, answer และ status ตามลำดับ
แสดงตัวอย่างดังนี้
context : ด.ต.ประสิทธิ์ ชาหอมชื่นอายุ 55 ปี ผบ.หมู่งาน ป.ตชด. 24 อุดรธานีถูกยิงด้วยอาวุธปืนอาก้าเข้าที่แขนซ้าย 3 นัดหน้าท้อง 1 นัดส.ต.อ.ประเสริฐ ใหญ่สูงเนินอายุ 35 ปี ผบ.หมู่กก. 1 ปส.2 บช.ปส. ถูกยิงเข้าที่แขนขวากระดูกแตกละเอียดร.ต.อ.ชวพล… See the full description on the dataset page: https://huggingface.co/datasets/SuperAI2-Machima/ThaiQA_LST20.Yord_ThaiQA_LST20พี่ยอด และน้อง ๆ ในทีมบ้านมัณิชมา ร่วมกันสร้างชุดข้อมูล คำถาม - คำตอบ จากชุดข้อมูล LST-20
โดยใช้ POS และ NER เพื่อมาสร้างชุดประโยคคำถาม
ได้ข้อมูลคำถาม - ตอบ ทั้งหมดประมาณ 1,000 แถว
ls-test-clean-plathonic-repBEEG-agents
BEEG agents
Dataset with three splits:
train
eval
sft_traces
Dataset_lstmlstm-classifier
clean.py
Dataset Summary
A news media dataset with image text modality, stored in jsonl format.
Preprocessing & Augmentation
Preprocessing: aggressive
Augmentation: mixup cutmix
Splits & Sampling
Split strategy: temporal
Sampling: stratified
Quality & Labeling
Quality filtering: moderate
Labeling: semi auto
Files
clean.py — main artifact of this repository
License
See the license… See the full description on the dataset page: https://huggingface.co/datasets/ethandavi/lstm-classifier.CQuAE
CQuAE: A New French Question-Answering Corpus for Teaching Assistant
CQuAE (Contextualised Question-Answering for Education) is a French question-answering dataset in the domain of secondary education.
It has been designed to facilitate the development of virtual teaching assistants,
with a particular focus on creating and answering complex questions that go beyond simple fact extraction.
CQuAE includes questions, answers, and corresponding source documents (excerpts of textbook or… See the full description on the dataset page: https://huggingface.co/datasets/LsTam/CQuAE.COCO-Facet
COCO-Facet
COCO-Facet is a benchmark for attribute-focused text-to-image retrieval ("Facets" of images). Annotations are derived from MSCOCO 2017, COCO-Stuff, Visual7W, and VisDial.
Code: https://github.com/lst627/COCO-Facet
Contents
Path
Description
benchmark/*.json
11 retrieval subsets (queries and candidate image references)
val2017.zip
MSCOCO val2017 images (5,000 files, ~788 MB)
VisualDialog_val2018.zip
VisDial val2018 images (2,064 files, ~318 MB)… See the full description on the dataset page: https://huggingface.co/datasets/lst627/COCO-Facet.p2-etf-rnn-lstm-resultsls_train_datalstm-asr-test
preprocess.py
Dataset Summary
A memes dataset with text tabular modality, stored in npy sharded format.
Preprocessing & Augmentation
Preprocessing: auto ml
Augmentation: light
Splits & Sampling
Split strategy: kfold 5
Sampling: contrastive
Quality & Labeling
Quality filtering: strict
Labeling: self training
Files
preprocess.py — main artifact of this repository
License
See the… See the full description on the dataset page: https://huggingface.co/datasets/artemyakovlev/lstm-asr-test.lstm_crypto_datasetThis is dataset where we try to put a lot of data into an LSTM and see what we get.
Mdist_Chem_LST_150k_ft_data_from_full_360kMdist_Chem_LST_200k_ft_data_from_full_360kMdist_Chem_LST_Qwen3_14B_full_ft_dataCQuAE_documents
CQuAE base Documents Dataset Card
Overview
The CQuAE dataset is a new French contextualized question-answering corpus focused on the education domain. It provides a structured annotation system that enhances the dataset's applicability for educational contexts and linguistic research. The primary documents used for annotations are detailed below, capturing a diversity of sources and content suitable for developing robust question-answering models.
CQuAE Documents… See the full description on the dataset page: https://huggingface.co/datasets/LsTam/CQuAE_documents.VerseFusion-LSTV
VerSeFusion-LSTV
A re-fused, PIR-canonical version of the VerSe 2019 and VerSe 2020 vertebra
segmentation challenges, with VERIDAH (Möller 2026) label corrections applied
for thoracolumbar transitional vertebrae.
Dataset stats
Total scans: 68
Total patients: 68
Splits: training=21, validation=23, test=24
Source: VerSe 2019 + VerSe 2020 (combined) with VERIDAH corrections
Canonical orientation: PIR (axis 0 = P, axis 1 = I, axis 2 = R)
VERIDAH-corrected subjects: 13… See the full description on the dataset page: https://huggingface.co/datasets/gregoryschwingmdphd/VerseFusion-LSTV.opus_instruction_format
Dataset Description: opus_instruction_format
This dataset is a translation dataset from opus-en-fr data, in the same format as the Stanford Alpaca dataset. The dataset contains a set of instructions for translation tasks, which include the following two reformulations:
"Traduire la ou les phrases suivantes en anglais" (Translate the following sentence(s) into English)
"Traduce the following sentences in english".
The dataset consists of input sentences in either English or French… See the full description on the dataset page: https://huggingface.co/datasets/LsTam/opus_instruction_format.bist-dp-lstm-trading-turkish_financial_news
turkish_financial_news
Turkish financial news corpus with sentiment labels
Dataset Details
Format: json
Size: ~50MB compressed
Language: Turkish (labels), Numeric (data)
License: MIT
Created: 2025-08-27
Dataset Structure
BIST Historical Data
Symbols: BIST 30 index stocks
Timeframes: 1m, 5m, 15m, 60m, 1d
Features: OHLCV + 131 technical indicators
Date Range: 2019-2024
Technical Indicators
Trend: SMA, EMA, MACD, Bollinger Bands… See the full description on the dataset page: https://huggingface.co/datasets/rsmctn/bist-dp-lstm-trading-turkish_financial_news.Mdist_Chem_LST_250k_ft_data_from_full_360kraw_samples_md
Dataset Card for "raw_samples_md"
More Information needed
africa-unsdg-red-list-index-er-rsk-lst
Africa Unsdg Red List Index Er Rsk Lst | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: parquet - Sector: other_unclassified - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-unsdg-red-list-index-er-rsk-lst.
