datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
triveni-raw
📦 Pretraining Corpus
📊 Dataset Overview
This dataset combines data from two major sources—Vaani and Flickr30k—to support multilingual and multimodal model pretraining.
Source
Languages
Samples per Language
Total Samples
Vaani
Hindi, English, Hinglish
30,195
90,585
Flickr30k
Hindi, English, Hinglish
31,014
93,042
Total
—
—
183,627
📁 Dataset Sources
🗣️ Vaani Dataset
License: CC-BY-4.0
Description:
VAANI is an… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/triveni-raw.Triveni
📦 Pretraining Corpus
📊 Dataset Overview
This dataset combines data from two major sources—Vaani and Flickr30k—to support multilingual and multimodal model pretraining.
Source
Languages
Samples per Language
Total Samples
Vaani
Hindi, English, Hinglish
30,195
90,585
Flickr30k
Hindi, English, Hinglish
31,014
93,042
Total
—
—
183,627
📁 Dataset Sources
🗣️ Vaani Dataset
License: CC-BY-4.0
Description:
VAANI is an… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/Triveni.latentsig-med-triage-router
LatentSig Medical Triage Router Dataset
1,000 verified medical triage tool-call samples — 500 English + 500 Hinglish — for fine-tuning Small Language Models (SLMs) as structured medical triage routers.
Overview
This dataset trains SLMs (1B–3B parameters) to act as reliable structured tool-callers for clinical medical triage. Given a patient symptom description, the model must:
Select the correct tool from 7 available medical tools
Output a valid JSON tool call… See the full description on the dataset page: https://huggingface.co/datasets/fhai50032/latentsig-med-triage-router.kenya-bee-health-qa-image-triples
Kenya Bee Health Training Data
This folder is a starter database for BeeCare Anywhere / Gemma Apiary. It is intentionally small, transparent, and license-aware: use it to prove the Q/A/image-triple pipeline, then expand it with Kenyan field data before trusting model behavior in production.
Important Model Note
google/gemma-2b is a text-to-text, decoder-only model. It cannot directly read pictures. Use these image triples with a vision-capable model path, for… See the full description on the dataset page: https://huggingface.co/datasets/yahelr1/kenya-bee-health-qa-image-triples.photography-triangle-essence
The Photography Triangle / Треугольник фотографии
Reader's Guide for Language Models / Руководство по чтению для языковых моделей
Version 1.4.2 — May 2026
This repository contains the authorial essence (a reader's guide for language models) accompanying the book The Photography Triangle / "Треугольник фотографии" by Alexander Zabara (Paris, 2025), together with the full text of the book in both language editions.
Данный репозиторий содержит авторскую эссенцию… See the full description on the dataset page: https://huggingface.co/datasets/photographytriangle/photography-triangle-essence.babel-briefings
Babel Briefings News Headlines Dataset README
Break Free from the Language Barrier
Version: 1 - Date: 30 Oct 2023
Collected and Prepared by Felix Leeb (Max Planck Institute for Intelligent Systems, Tübingen, Germany)
License: Babel Briefings Headlines Dataset © 2023 by Felix Leeb is licensed under CC BY-NC-SA 4.0
Check out our paper on arxiv.
This dataset contains 4,719,199 news headlines across 30 different languages collected between 8 August 2020 and 29 November 2021. The… See the full description on the dataset page: https://huggingface.co/datasets/Tribhuvand/babel-briefings.tricad-code
TriView2CAD-Code
Dimensioned orthographic engineering drawings -> executable CadQuery code.
200,000 samples of prefabricated bridge piers (160,000 train / 40,000 test), each a 1475x1475 three-view drawing
(front / top / side) with every dimension annotated, paired with a CadQuery program
that rebuilds the part exactly.
input one PNG holding the front, top and side views, fully dimensioned
output CadQuery (Python) source; executing it yields the corresponding solid… See the full description on the dataset page: https://huggingface.co/datasets/AI4Manufacturing/tricad-code.
