datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
diamond-price-predictor-logs2
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/manojdec25/diamond-price-predictor-logs2.anveshana
Dataset Card for Anveshana
Dataset Details
Dataset Description
we embarked on a comprehensive benchmarking study to explore and evaluate current state-of-the-art models for Cross-Lingual Information Retrieval (CLIR) from English to Sanskrit. Our primary objective is to assess the effectiveness of these models in accurately retrieving Sanskrit documents based on English queries. To achieve this, we meticulously assembled a robust dataset, focusing on the… See the full description on the dataset page: https://huggingface.co/datasets/manojbalaji1/anveshana.salesdataOpen_Platypus_Orcagarage-bAInd/Open-Platypus -> in Orca
customer-support-tickets
Featuring Labeled Customer Emails and Support Responses
🔧 Synthetic IT Ticket Generator — Custom Dataset
Create a dataset tailored to your own queues & priorities (no PII).
👉 Generate custom data
Define your queues, priorities, language
Need an on-prem AI to auto-classify tickets?→ Open Ticket AI
There are 2 Versions of the dataset, the new version has more tickets, but only languages english and german. So please look at both files, to find what best fits your needs.… See the full description on the dataset page: https://huggingface.co/datasets/manojroyal23/customer-support-tickets.OrcaOrca Dataset
1. orca_1m_gpt4.csv - ~1M Orca Data generated by using GPT-4
2. orca_3.5m_gpt3.5.csv - ~3.5M Orca Data generated by using GPT-3.5
SanskritDatasettrainingroman-nepali-gemma-finalmath-word-problemsnyc-taxi-2025-duckdb
Taxi-Revenue-and-Surge-Analytics
Every taxi ride tells a data story. From raw trip logs to executive insight, analyzed 9.3M NYC Yellow Taxi trips by cleaning messy source data, building tested dimensional models, and uncovering revenue trends, surge patterns, and top-performing zones through an interactive dashboard
ecommerce_qna
Ecommerce Customer Query Dataset
Roman Nepali Dataset Generated for Customer Support Question Answer
ift-nepali-v5roman-nepali-alpacaecommerce_ragmga_breed_ng_manokgemini3goemostrength_weaknessagrichattelugu_Colloqual_dataset.csvSuperKartDatasetML
