datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Grounded_3D_LLM_with_Referent_Tokens_Dataset
Grounded 3D-LLM Dataset
For detailed information and resources, please visit the following links:
Paper
Arxiv
Project Website
Dataset Access
Code
We are in the process of releasing our data incrementally:
Processed ScanNet200 PCD(~7G):
Each .npyfile represents a N*12 array with the following structure:
coordinates, color, normals, segments, labels = (
points[:, :3],
points[:, 3:6],
points[:, 6:9],
points[:, 9]… See the full description on the dataset page: https://huggingface.co/datasets/ShuaiYang03/Grounded_3D_LLM_with_Referent_Tokens_Dataset.vision-token-compression-bench
OPTIC-Bench
Optical Text In-Context Benchmark: how reliably do LLMs consume text
delivered as rendered images versus plain text tokens?
In summary, the evaluation reported here finds that optical text compression
is effective only within a narrow and specific envelope. Delivering content
as rendered images genuinely reduces input tokens, by thirteen to
fifty-four per cent depending on the model and the language, but only when
the document is long, the rendering is dense and the… See the full description on the dataset page: https://huggingface.co/datasets/translorentz/vision-token-compression-bench.omnimcp_cyber_token_revocation_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_cyber_token_revocation_teaser.m23k-tokenized
m1: Unleash the Potential of Test-Time Scaling for Medical Reasoning in Large Language Models
A simple test-time scaling strategy, with minimal fine-tuning, can unlock strong medical reasoning within large language models.
⚡ Introduction
Hi! Welcome to the huggingface repository for m1 (https://github.com/UCSC-VLAA/m1)!
m1 is a medical LLM designed to enhance reasoning through efficient test-time scaling. It enables lightweight models to match or exceed the performance of… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/m23k-tokenized.prompts_under_512_tokens
Under 512 Tokens Prompts Dataset
Created by Aipresso LIMITED, London, UK
⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use
Dataset Overview
Specialized collection of short-form English prompts (under 512 tokens), perfect for training models with context length constraints or faster iteration cycles.
📊 Dataset Statistics
Metric
Value
Total Files
200
Rows Per File
10,000
Total Rows
2,000,000
Token Range
1 to 511 tokens… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/prompts_under_512_tokens.m1k-tokenized
m1: Unleash the Potential of Test-Time Scaling for Medical Reasoning in Large Language Models
A simple test-time scaling strategy, with minimal fine-tuning, can unlock strong medical reasoning within large language models.
⚡ Introduction
Hi! Welcome to the huggingface repository for m1 (Github, Paper)!
m1 is a medical LLM designed to enhance reasoning through efficient test-time scaling. It enables lightweight models to match or exceed the performance of much larger… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/m1k-tokenized.medium_512_1k_tokens_prompts
Medium 512-1K Tokens Prompts Dataset
Created by Aipresso LIMITED, London, UK
⚠️ By using this dataset you agree to our Terms of Use.
Overview
703 high-quality English prompts whose length lies between 512 and 1 000 tokens.Every prompt has been de-duplicated, cleaned and token-counted with the GPT-2 tokenizer.
Statistics
Rows
Token range
File size
Format
703
512 – 1 000
2.9 MB
CSV
Use-cases
Medium-context language-model fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/medium_512_1k_tokens_prompts.long_over_1k_tokens_prompts
Long Over 1K Tokens Prompts Dataset
Created by Aipresso LIMITED, London, UK
⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use
Dataset Overview
Specialized collection of long-form English prompts (≥ 1 000 tokens) for training advanced models that require extensive context and complex reasoning.
📊 Dataset Statistics
Metric
Value
Total Rows
289
Token Range
1 001 – 10 000 tokens
File Size
≈ 3.7 MB
Format
Single CSV file
Target… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/long_over_1k_tokens_prompts.pashto-warmup-tokens
Pashto Warmup Tokens Dataset
This dataset contains a curated, deduplicated collection of high-quality, contextually accurate Pashto linguistic examples. It maps structural language tasks directly to the most critical vocabulary tokens in Pashto, providing a reliable corpus for token warmup, instruction tuning, evaluation, and post-OCR text correction workflows.
Dataset Summary
The initial release consists of 4,087 verified entries targeting high-frequency and… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-warmup-tokens.
