datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ConverSeg
ConverSeg: Conversational Image Segmentation
ConverSeg is a benchmark for grounding abstract, intent-driven concepts into pixel-accurate masks. Unlike standard referring expression datasets, ConverSeg focuses on physical reasoning, affordances, and safety.
Dataset Structure
The dataset contains two splits:
sam_seeded: 1,194 samples generated via SAM2 + VLM verification.
human_annotated: 493 samples with human-drawn masks (initialized from COCO).
Licensing &… See the full description on the dataset page: https://huggingface.co/datasets/aadarsh99/ConverSeg.sessionsAAD-Across-All-Domains-Testclaude-gemini-reasoning
Claude-gemini-reasoning
[!CAUTION]
Aproximately 22% of this dataset had its reasoning truncated. I will try and fix this soon, but for now I would just remove the reasoning from those examples to have some non-reasoning examples. Thanks to idkwhattoputherenow for pointing out the issue.
This dataset contains 13,627 highly curated, deduplicated questions requiring deep multi-step reasoning, complex logic, advanced mathematics, physics, and coding.
The data has been meticulously… See the full description on the dataset page: https://huggingface.co/datasets/Aadeshisdoingsomething/claude-gemini-reasoning.sona-corpus
THE SONA CORPUS — Noisy-to-Clean Hindi–English Parallel Dataset
A clean, bilingual dataset card you can read at a glance and use immediately.
Curated by: Aditya (AADIMIND)
Languages: Hindi, English
Total examples: 581312 (INPUT: 256 TOKEN• TARGET: 256 TOKEN)
Tasks: Text cleaning, GEC, OCR post-processing, Seq2Seq fine-tuning
License: MIT
Source: Hindi Wikipedia (HiWiki) processed into noisy–clean pairs
Repo: https://huggingface.co/datasets/AADIMIND/sona-corpus… See the full description on the dataset page: https://huggingface.co/datasets/AADIMIND/sona-corpus.aaditya__Llama3-OpenBioLLM-70B-details
Dataset Card for Evaluation run of aaditya/Llama3-OpenBioLLM-70B
Dataset automatically created during the evaluation run of model aaditya/Llama3-OpenBioLLM-70B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/aaditya__Llama3-OpenBioLLM-70B-details.telco-5G-data-faultsSynthetic test dataset for 5G data service faults in the core and RAN network domains. It is used to train the telcoLLM to simulate assistance model for network operations.
dummy_requestsncertsprimate_datasettest_resultsIndiaLaw-19kreverse_answersxpathaadsVedika_V1encrypted_conversationaa_de_ssskill-extraction-finetune-datasetGPT_4o_mini_Fine_tune
