CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01johnswyou /wavelet-lstm-camels-models Wavelet-LSTM CAMELS Streamflow Models A collection of 61,380 pre-trained LSTM models for daily streamflow forecasting across 620 USGS catchments from the CAMELS dataset. Each catchment has 99 independently trained models: 33 wavelet filters × 3 lead times (1, 3, 5 days) = 99 wavelet-enhanced models 33 matching baseline models (same architecture, no wavelet transform) Models are designed to be ensembled across wavelets for robust predictions with uncertainty estimates.… See the full description on the dataset page: https://huggingface.co/datasets/johnswyou/wavelet-lstm-camels-models.time-series-forecasting10K<n<100K0 likes10k downloads7mo agoHugging Face02shivDwd /W_LSTMix_test_datasettext10M<n<100M0 likes1.2k downloads1y agoHugging Face03P2SAMAPA /p2-etf-x-lstm-extended-results0 likes1.1k downloads10h agoHugging Face04ReigeKwitt /LsTSCP_Media_Vol_3videon<1K0 likes375 downloads11d agoHugging Face05ReigeKwitt /LsTSCP_Media_Vol_2videon<1K0 likes266 downloads18d agoHugging Face06ReigeKwitt /LsTSCP_Media_Vol_1videon<1K0 likes246 downloads15d agoHugging Face07lsteno /RLM-Evals RLM Evals Curated evaluation bundle for comparing RLM policies against the benchmark family used in the Recursive Language Models paper. The repository stores unsampled benchmark subsets. Sampling for a specific experiment should be done downstream with a fixed seed and recorded in the eval manifest. Subsets longbench_v2_codeqa: 50 rows. Source: zai-org/LongBench-v2 train filtered to Code Repository Understanding / Code repo QA. browsecomp_plus: 830 rows. Source:… See the full description on the dataset page: https://huggingface.co/datasets/lsteno/RLM-Evals.tabular1K<n<10K0 likes192 downloads3mo agoHugging Face08lst-nectec /lst20LST20 Corpus is a dataset for Thai language processing developed by National Electronics and Computer Technology Center (NECTEC), Thailand. It offers five layers of linguistic annotation: word boundaries, POS tagging, named entities, clause boundaries, and sentence boundaries. At a large scale, it consists of 3,164,002 words, 288,020 named entities, 248,181 clauses, and 74,180 sentences, while it is annotated with 16 distinct POS tags. All 3,745 documents are also annotated with one of 15 news genres. Regarding its sheer size, this dataset is considered large enough for developing joint neural models for NLP. Manually download at https://aiforthai.in.th/corpus.phptoken-classification10K<n<100K6 likes176 downloads3y agoHugging Face09SuperAI2-Machima /ThaiQA_LST20SuperAI Engineer Season 2 , Machima Machima_ThaiQA_LST20 เป็นชุดข้อมูลที่สกัดหาคำถาม และคำตอบ จากบทความในชุดข้อมูล LST20 โดยสกัดได้คำถาม-ตอบทั้งหมด 7,642 คำถาม มีข้อมูล 4 คอลัมน์ ประกอบด้วย context, question, answer และ status ตามลำดับ แสดงตัวอย่างดังนี้ context : ด.ต.ประสิทธิ์ ชาหอมชื่นอายุ 55 ปี ผบ.หมู่งาน ป.ตชด. 24 อุดรธานีถูกยิงด้วยอาวุธปืนอาก้าเข้าที่แขนซ้าย 3 นัดหน้าท้อง 1 นัดส.ต.อ.ประเสริฐ ใหญ่สูงเนินอายุ 35 ปี ผบ.หมู่กก. 1 ปส.2 บช.ปส. ถูกยิงเข้าที่แขนขวากระดูกแตกละเอียดร.ต.อ.ชวพล… See the full description on the dataset page: https://huggingface.co/datasets/SuperAI2-Machima/ThaiQA_LST20.text1K<n<10K1 likes150 downloads5y agoHugging Face10SuperAI2-Machima /Yord_ThaiQA_LST20พี่ยอด และน้อง ๆ ในทีมบ้านมัณิชมา ร่วมกันสร้างชุดข้อมูล คำถาม - คำตอบ จากชุดข้อมูล LST-20 โดยใช้ POS และ NER เพื่อมาสร้างชุดประโยคคำถาม ได้ข้อมูลคำถาม - ตอบ ทั้งหมดประมาณ 1,000 แถว tabularn<1K1 likes133 downloads5y agoHugging Face11Adel-Moumen /ls-test-clean-plathonic-rep0 likes103 downloads1y agoHugging Face12lsteno /BEEG-agents BEEG agents Dataset with three splits: train eval sft_traces text1K<n<10K0 likes57 downloads4mo agoHugging Face13rakuks /Dataset_lstmgeospatial0 likes54 downloads1y agoHugging Face14ethandavi /lstm-classifier clean.py Dataset Summary A news media dataset with image text modality, stored in jsonl format. Preprocessing & Augmentation Preprocessing: aggressive Augmentation: mixup cutmix Splits & Sampling Split strategy: temporal Sampling: stratified Quality & Labeling Quality filtering: moderate Labeling: semi auto Files clean.py — main artifact of this repository License See the license… See the full description on the dataset page: https://huggingface.co/datasets/ethandavi/lstm-classifier.0 likes54 downloads28d agoHugging Face15LsTam /CQuAE CQuAE: A New French Question-Answering Corpus for Teaching Assistant CQuAE (Contextualised Question-Answering for Education) is a French question-answering dataset in the domain of secondary education. It has been designed to facilitate the development of virtual teaching assistants, with a particular focus on creating and answering complex questions that go beyond simple fact extraction. CQuAE includes questions, answers, and corresponding source documents (excerpts of textbook or… See the full description on the dataset page: https://huggingface.co/datasets/LsTam/CQuAE.text10K<n<100K0 likes51 downloads2y agoHugging Face16lst627 /COCO-Facet COCO-Facet COCO-Facet is a benchmark for attribute-focused text-to-image retrieval ("Facets" of images). Annotations are derived from MSCOCO 2017, COCO-Stuff, Visual7W, and VisDial. Code: https://github.com/lst627/COCO-Facet Contents Path Description benchmark/*.json 11 retrieval subsets (queries and candidate image references) val2017.zip MSCOCO val2017 images (5,000 files, ~788 MB) VisualDialog_val2018.zip VisDial val2018 images (2,064 files, ~318 MB)… See the full description on the dataset page: https://huggingface.co/datasets/lst627/COCO-Facet.text-to-image10K<n<100K0 likes51 downloads4mo agoHugging Face17P2SAMAPA /p2-etf-rnn-lstm-resultstabular1K<n<10K0 likes48 downloads3mo agoHugging Face18TerryXu666 /ls_train_data0 likes42 downloads1y agoHugging Face19artemyakovlev /lstm-asr-test preprocess.py Dataset Summary A memes dataset with text tabular modality, stored in npy sharded format. Preprocessing & Augmentation Preprocessing: auto ml Augmentation: light Splits & Sampling Split strategy: kfold 5 Sampling: contrastive Quality & Labeling Quality filtering: strict Labeling: self training Files preprocess.py — main artifact of this repository License See the… See the full description on the dataset page: https://huggingface.co/datasets/artemyakovlev/lstm-asr-test.0 likes41 downloads28d agoHugging Face20345rf4gt56t4r3e3 /lstm_crypto_datasetThis is dataset where we try to put a lot of data into an LSTM and see what we get. tabular100K<n<1M0 likes31 downloads8mo agoHugging Face21yileitu /Mdist_Chem_LST_150k_ft_data_from_full_360ktext100K<n<1M0 likes30 downloads3mo agoHugging Face22yileitu /Mdist_Chem_LST_200k_ft_data_from_full_360k0 likes30 downloads3mo agoHugging Face23yileitu /Mdist_Chem_LST_Qwen3_14B_full_ft_datatext100K<n<1M0 likes25 downloads3mo agoHugging Face24LsTam /CQuAE_documents CQuAE base Documents Dataset Card Overview The CQuAE dataset is a new French contextualized question-answering corpus focused on the education domain. It provides a structured annotation system that enhances the dataset's applicability for educational contexts and linguistic research. The primary documents used for annotations are detailed below, capturing a diversity of sources and content suitable for developing robust question-answering models. CQuAE Documents… See the full description on the dataset page: https://huggingface.co/datasets/LsTam/CQuAE_documents.text1K<n<10K0 likes23 downloads2y agoHugging Face25gregoryschwingmdphd /VerseFusion-LSTV VerSeFusion-LSTV A re-fused, PIR-canonical version of the VerSe 2019 and VerSe 2020 vertebra segmentation challenges, with VERIDAH (Möller 2026) label corrections applied for thoracolumbar transitional vertebrae. Dataset stats Total scans: 68 Total patients: 68 Splits: training=21, validation=23, test=24 Source: VerSe 2019 + VerSe 2020 (combined) with VERIDAH corrections Canonical orientation: PIR (axis 0 = P, axis 1 = I, axis 2 = R) VERIDAH-corrected subjects: 13… See the full description on the dataset page: https://huggingface.co/datasets/gregoryschwingmdphd/VerseFusion-LSTV.image-segmentationn<1K0 likes22 downloads4mo agoHugging Face26LsTam /opus_instruction_format Dataset Description: opus_instruction_format This dataset is a translation dataset from opus-en-fr data, in the same format as the Stanford Alpaca dataset. The dataset contains a set of instructions for translation tasks, which include the following two reformulations: "Traduire la ou les phrases suivantes en anglais" (Translate the following sentence(s) into English) "Traduce the following sentences in english". The dataset consists of input sentences in either English or French… See the full description on the dataset page: https://huggingface.co/datasets/LsTam/opus_instruction_format.text10K<n<100K0 likes21 downloads3y agoHugging Face27rsmctn /bist-dp-lstm-trading-turkish_financial_news turkish_financial_news Turkish financial news corpus with sentiment labels Dataset Details Format: json Size: ~50MB compressed Language: Turkish (labels), Numeric (data) License: MIT Created: 2025-08-27 Dataset Structure BIST Historical Data Symbols: BIST 30 index stocks Timeframes: 1m, 5m, 15m, 60m, 1d Features: OHLCV + 131 technical indicators Date Range: 2019-2024 Technical Indicators Trend: SMA, EMA, MACD, Bollinger Bands… See the full description on the dataset page: https://huggingface.co/datasets/rsmctn/bist-dp-lstm-trading-turkish_financial_news.tabular-classification10K<n<100K0 likes21 downloads1y agoHugging Face28yileitu /Mdist_Chem_LST_250k_ft_data_from_full_360ktext100K<n<1M0 likes21 downloads3mo agoHugging Face29LsTam /raw_samples_md Dataset Card for "raw_samples_md" More Information needed textn<1K0 likes20 downloads2y agoHugging Face30electricsheepafrica /africa-unsdg-red-list-index-er-rsk-lst Africa Unsdg Red List Index Er Rsk Lst | Africa (Electric Sheep Africa metadata inventory) Size category: 1K<n<10K - Formats: parquet - Sector: other_unclassified - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-unsdg-red-list-index-er-rsk-lst.tabulartabular-classification1K<n<10K0 likes16 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.