datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
trivia_qa
Dataset Card for "trivia_qa"
Dataset Summary
TriviaqQA is a reading comprehension dataset containing over 650K
question-answer-evidence triples. TriviaqQA includes 95K question-answer
pairs authored by trivia enthusiasts and independently gathered evidence
documents, six per question on average, that provide high quality distant
supervision for answering the questions.
Supported Tasks and Leaderboards
More Information Needed
Languages… See the full description on the dataset page: https://huggingface.co/datasets/mandarjoshi/trivia_qa.clinical-trials-protocolstrinhquang2003trinhhuy2001trinhkhoi2005trinhminh2005trinhphuong2000trinhphat2003trinhthu2002trinhthihong1993potter-plant-identificationtrinhngoc2005trinhdung2001trinhthuhien99trinhbao2002trinity-dataset-v3trinhthanhha1996trinhducbao77trinhngocmaivnsoft-trigger-verifiedtrinity-dataset-v3VC3_KFtripitaka-mbu
Multi-File CSV Dataset
คำอธิบาย
พระไตรปิฎกและอรรถกถาไทยฉบับมหามกุฏราชวิทยาลัย จำนวน ๙๑ เล่ม
ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์
01/010001.csv: เล่ม 1 หน้า 1
01/010002.csv: เล่ม 1 หน้า 2
...
02/020001.csv: เล่ม 2 หน้า 1
คำอธิบายของแต่ละเล่ม
เล่ม ๑ (863 หน้า): พระวินัยปิฎก มหาวิภังค์ เล่ม ๑ ภาค ๑
เล่ม ๒ (664 หน้า): พระวินัยปิฎก มหาวิภังค์ เล่ม ๑ ภาค ๒
เล่ม ๓: พระวินัยปิฎก มหาวิภังค์ เล่ม ๑ ภาค ๓
เล่ม ๔: พระวินัยปิฎก มหาวิภังค์ เล่ม ๒
เล่ม ๕: พระวินัยปิฎก… See the full description on the dataset page: https://huggingface.co/datasets/uisp/tripitaka-mbu.TheBioCollection
TheBioCollection
TheBioCollection is a 52.6B-token pretraining-scale corpus for biology that transforms heterogeneous biological resources into LLM training-friendly data. It is built through a construction pipeline that collects resources across biological domains, refines them through deduplication, entity tagging and augmentation, enriches them with tool-computed biological properties, and render them as instruction-form data with programmatically verifiable answers. The… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/TheBioCollection.based_triviaqatripadvisor-hotel-reviews
Dataset Card for "tripadvisor-hotel-reviews"
Dataset Summary
Hotels play a crucial role in traveling and with the increased access to information new pathways of selecting the best ones emerged.
With this dataset, consisting of 20k reviews crawled from Tripadvisor, you can explore what makes a great hotel and maybe even use this model in your travels!
Citations on a scale from 1 to 5.
Languages
english
Citation Information
If you use this dataset in… See the full description on the dataset page: https://huggingface.co/datasets/argilla/tripadvisor-hotel-reviews.trip-plus-databasedvla-place7obj-roll-250hz-events-250fps-delta-trimThis dataset was created using LeRobot.
Dataset Description
7-object place (MuJoCo / robosuite, 250 Hz, event + RGB)
14,083 successful demonstrations of a Panda arm picking an object off the table and dropping it
into a receptacle (bowl / plate / tray), across 7 objects:
5 grocery objects (10,083 episodes): can, apple, avocado, potato, lemon
2 round textured objects (4,000 episodes): tennis ball, baseball
This dataset is the union of… See the full description on the dataset page: https://huggingface.co/datasets/mickeykang/dvla-place7obj-roll-250hz-events-250fps-delta-trim.tripadvisor_hotel_reviewstrigger_dataset
