datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
allsides_text_proper_truncatedFineWeb-Edu-10B-Shuffled-DOC-Truncatedimdb-truncated
Dataset Card for "imdb-truncated"
More Information needed
FineWeb-Edu-10B-Obfuscation-Truncatedopen-r1-truncated-coding-pythonFineWeb-Edu-10B-Shuffled-SENT-Truncatedopenr1-math-verified-solutions-truncatedprm800k-truncated-prefix-legacy-results
PRM800K truncated-prefix legacy results
Public archive of legacy experiment artifacts produced before complete PRM800K attempts were restored. Scientific interpretation and the corrected rerun are documented in the source repository.
Declared files: 60748
Declared bytes: 187013895604
Verification: exact remote paths, file count, and aggregate size at a pinned revision
Full remote content readback: not performed
mlsum-spanish-truncated-512
Dataset Card for "mlsum-spanish-truncated-512"
More Information needed
wmt19-de-en-truncatedsubmissions-truncated-10kimdb-truncatedmR3-Dataset-100K-EasyToHard-Truncatedimdb-truncatedyelp_reviews_encoded_hidden_outputs_truncatedtruncated_dyck3tldr-17-ChatML-tokenized-truncatedtruncated_dyck3_full_widthfic_top_likes_with_all_summaries_base_qwen_truncatedorm-v0-truncated-binarymovie_rationales_truncated2026-05-11_twist-lerobot-truncated-return-home-exp-truncated-return-home-20260520-095140This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aloha",
"total_episodes": 45,
"total_frames": 52398,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 200,
"video_files_size_in_mb": 200,
"fps": 50,
"splits": {
"train": "0:45"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lyl472324464/2026-05-11_twist-lerobot-truncated-return-home-exp-truncated-return-home-20260520-095140.pile-LlamaTokenizerFast-32k-truncated-toy
LADM: Long-context Training Data Selection with Attention-based Dependency Measurement for LLMs
📖 Paper • 🤗 HF Repo
🔍 Table of Contents
🌐 Overview
📚 Preparation
⏳ Data Selection
📈 Training
📝 Citation
🌐 Overview
Long-context modeling has drawn more and more attention in the area of Large Language Models (LLMs). Continual training with long-context data becomes the de-facto method to equip LLMs with the ability to process long inputs. However… See the full description on the dataset page: https://huggingface.co/datasets/UltraRonin/pile-LlamaTokenizerFast-32k-truncated-toy.training_data_truncatedself-talk_gpt3.5_gpt4o_prefpairs_truncated2048_cutto1turnsopen-thoughts-4-2samples-math-qwen3-235b-a22b-truncated-test
open-thoughts-4-2samples-math-qwen3-235b-a22b-truncated-test
Test dataset with 2 samples. Sample 0 has a complete response; sample 1 is truncated to ~500 chars.
2026-05-12_twist-lerobot-truncated-return-home-exp-truncated-return-home-20260520-095140This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aloha",
"total_episodes": 57,
"total_frames": 70749,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 200,
"video_files_size_in_mb": 200,
"fps": 50,
"splits": {
"train": "0:57"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lyl472324464/2026-05-12_twist-lerobot-truncated-return-home-exp-truncated-return-home-20260520-095140.OpenHermes-2.5-1k-longest-truncatedThis is a dataset that was created from HuggingFaceH4/OpenHermes-2.5-1k-longest.
The purpose is to be able to use in axolotl config by adding:
datasets:
- path: Mihaiii/OpenHermes-2.5-1k-longest-truncated
type: alpaca
I eliminated all "glaive-code-assist" rows + some others.
See the OpenHermes-2.5-1k-longest-truncated.ipynb notebook for details on how the dataset was constructed.
zkml-github-repos-truncatedThis dataset is a truncated version of this one but where the format is compatible with MLX-lora,
using {"text": "This is an example for the model."}, and where each entry has been truncated, following some code logic (i.e., following classes, functions etc) to ensure
each entry is smaller than 2048 tokens.
alpacaeval_mt_dev_8maxturns_truncated3500
