CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01simpleG2023 /chinese-clean-energy-battery-open-intelligence 🔬 Chinese Clean Energy, Battery Chemistry & Smart Grid Open Intelligence Dataset Curated open intelligence dataset tracking authentic Chinese scientific breakthroughs in Solid-State Battery chemistry, Perovskite Solar cells, Ultra-High Voltage (UHV) power grids, and industrial decarbonization. [!IMPORTANT] Data Completeness & Research Authenticity Notice: Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-clean-energy-battery-open-intelligence.tabulartext-retrieval1K<n<10K0 likes354 downloads3h agoHugging Face02BatsResearch /planetarium Dataset Card for Planetarium🪐 Planetarium🪐 is a dataset and benchmark for assessing LLMs in translating natural language descriptions of planning problems into PDDL. We developed a robust method for comparing PDDL problem descriptions using graph isomorphism. Dataset Details This dataset is a set of pairs of planning problems in PDDL and natural language descriptions from the Blocks World and Gripper domains. The task is to take descriptions of various initial and goal… See the full description on the dataset page: https://huggingface.co/datasets/BatsResearch/planetarium.tabulartranslation100K<n<1M18 likes227 downloads2y agoHugging Face03costadev00 /openai-terra-batch-wiki-brazil-1000-partial-20260724-01 OpenAI Terra Batch — Wikipédia PT-BR (run parcial) Checkpoint publicável de uma execução real e interrompida do fluxo document_task_matrix. A execução planejou gerar uma matriz de 1.000 documentos da Wikipédia em português por 25 tasks canônicas usando a Responses API Batch e o modelo gpt-5.6-terra. Este repositório não representa a conclusão dos 25.000 pares planejados. Ele contém somente os 1.282 candidatos aceitos após a reconciliação offline de todos os resultados Batch já… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/openai-terra-batch-wiki-brazil-1000-partial-20260724-01.texttext-generation1K<n<10K0 likes143 downloads2mo agoHugging Face04batterydata /battery-device-data-qa Battery Device QA Data Battery device records, including anode, cathode, and electrolyte. Examples of the question answering evaluation dataset: {'question': 'What is the cathode?', 'answer': 'Al foil', 'context': 'The blended slurry was then cast onto a clean current collector (Al foil for the cathode and Cu foil for the anode) and dried at 90 °C under vacuum overnight.', 'start index': 645} {'question': 'What is the anode?', 'answer': 'Cu foil', 'context': 'The blended slurry was… See the full description on the dataset page: https://huggingface.co/datasets/batterydata/battery-device-data-qa.tabularquestion-answeringn<1K7 likes121 downloads3y agoHugging Face05batuhanaktas /kids-multilingual-benchmark TinyAya v2 — Multilingual Benchmark for Children's AI Companions 2,312 child–AI conversational prompts across 23 languages, evaluated against four models with five-judge LLM-as-judge validation. 📄 Companion article: see HF Articles by @batuhanaktas. 💻 Code: https://github.com/aktasbatuhan/cohere-tiny-aya-for-kids Dataset summary This dataset contains: benchmark/items.jsonl — 2,312 benchmark items in 23 languages. Each item is a structured prompt designed to mimic… See the full description on the dataset page: https://huggingface.co/datasets/batuhanaktas/kids-multilingual-benchmark.texttext-generation10K<n<100K1 likes87 downloads5mo agoHugging Face06lmarena-ai /Llama-3-70b-battlesChatbot Arena user conversations between Llama-3-70b VS GPT-4-1025 or Llama-3-70b VS Claude-3-Opus with user preference votes. Single turn. Excludes ties. Used in Llama Data Analysis blog post and "VibeCheck: Discover and Quantify Qualitative Differences in Large Language Models" (Paper, Code). Citation @article{dunlap_vibecheck, title={VibeCheck: Discover and Quantify Qualitative Differences in Large Language Models}, author={Lisa Dunlap and Krishna Mandal and Trevor… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/Llama-3-70b-battles.textquestion-answering1K<n<10K3 likes71 downloads2y agoHugging Face07batuhanozkose /Rehber-CoT-Science 🧬 Rehber-CoT-Science: Turkish Scientific Reasoning Dataset Turkish Scientific Computational Reasoning (Chain-of-Thought) Dataset Multi-step scientific problem-solving dataset with verifiable Python code and detailed explanations Dataset • Author 📌 Changelog Eski sürümlere erişim: Branch menüsünden v1 seçebilirsiniz. Version Date Changes v2.0 24.12.2025 ✨ Yeni explained_answer alanı eklendi, Statistics domain eklendi, 712 örneğe genişletildi… See the full description on the dataset page: https://huggingface.co/datasets/batuhanozkose/Rehber-CoT-Science.textquestion-answering1K<n<10K4 likes59 downloads9mo agoHugging Face08costadev00 /smoke-openai-terra-batch-brasil-25-20260724-01 Smoke OpenAI Terra Batch — Brasil × 25 tasks Run real de validação do fluxo matricial document_task_matrix, executada sobre um único documento da Wikipédia em português com o título Brasil. Cada uma das 25 tasks canônicas recebeu exatamente um slot inicial. Resultado status: completed documentos: 1 pares planejados: 25 exemplos aceitos: 25 pares pulados: 0 pares esgotados: 0 resultados reais do backend: 27 retries com nova chamada: 2 backend: openai_api… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/smoke-openai-terra-batch-brasil-25-20260724-01.texttext-generationn<1K0 likes54 downloads2mo agoHugging Face09bryandts /instruction-dataset-indo-java-sunda-bali-gayo-batak-alas-minang-betawitexttext-generation100K<n<1M0 likes25 downloads2y agoHugging Face10batuhanozkose /Rehber-Bench-Mini Rehber-Bench-Mini Rehber-Bench-Mini, batuhanozkose/Rehber-CoT-Science veri setinden secilmis 50 soruluk kucuk ama dengeli bir Turkce bilimsel reasoning benchmarkidir. Tasarim Toplam soru sayisi: 50 Difficulty dagilimi: {'easy': 16, 'hard': 17, 'medium': 17} Domain dagilimi: {'Biology': 12, 'Chemistry': 8, 'Engineering': 5, 'Math': 5, 'Physics': 14, 'Science': 3, 'Statistics': 2, 'Computer Science': 1} Selection politikasi: deterministik seed: 42 domain hedefleri sabit… See the full description on the dataset page: https://huggingface.co/datasets/batuhanozkose/Rehber-Bench-Mini.textquestion-answeringn<1K1 likes24 downloads7mo agoHugging Face11BrainDelay /BatVenom BatVenom: Dual-Personality Roleplay Dataset 🦇🕷️ This dataset contains over 200+ hand-crafted and AI-assisted roleplay scenarios designed to fine-tune Large Language Models (LLMs) into the "BatVenom" persona—a hybrid of Batman (Bruce Wayne) and the Venom Symbiote. 📊 Dataset Structure The data is provided in the Alpaca/LLaMA-Factory format: instruction: The context or setup of the scene. input: The specific user prompt or dialogue. output: The formatted response showing… See the full description on the dataset page: https://huggingface.co/datasets/BrainDelay/BatVenom.texttext-generationn<1K0 likes21 downloads7mo agoHugging Face12dotiendat711 /real-estate-batdongsan.com.vn Bộ dữ liệu tin đăng căn hộ Việt Nam Tóm tắt Bộ dữ liệu này gồm các bản ghi tin đăng căn hộ tại Việt Nam, được export từ tầng hiển thị của backend bất động sản. Mỗi dòng tương ứng với một tin đăng/property post, bao gồm tiêu đề, mô tả, thuộc tính có cấu trúc, vị trí hành chính, giá, diện tích, ảnh, định danh nguồn và thông tin tiện ích xung quanh. Bộ dữ liệu phù hợp cho các bài toán tìm kiếm bất động sản, truy hồi ngữ nghĩa, retrieval-augmented generation (RAG)… See the full description on the dataset page: https://huggingface.co/datasets/dotiendat711/real-estate-batdongsan.com.vn.imagetext-retrieval1K<n<10K2 likes21 downloads4mo agoHugging Face13BatSilver /NLP-to-Semantic-Query_Benchmark_Dataset NLP-to-Semantic-Query Benchmark Dataset Overview This dataset is designed for evaluating AI agents and LLM systems that translate natural language analytical questions into structured semantic queries. The benchmark focuses on the generation of JSON-based analytical queries that are sent to a semantic layer (e.g. Cube.js) to retrieve analytical results from databases. The dataset can be used for: Evaluating NLP-to-query systems Benchmarking AI analytics agents Measuring… See the full description on the dataset page: https://huggingface.co/datasets/BatSilver/NLP-to-Semantic-Query_Benchmark_Dataset.texttext-generationn<1K0 likes18 downloads4mo agoHugging Face14milerssliu /battery-device-data-qa Battery Device QA Data Battery device records, including anode, cathode, and electrolyte. Examples of the question answering evaluation dataset: {'question': 'What is the cathode?', 'answer': 'Al foil', 'context': 'The blended slurry was then cast onto a clean current collector (Al foil for the cathode and Cu foil for the anode) and dried at 90 °C under vacuum overnight.', 'start index': 645} {'question': 'What is the anode?', 'answer': 'Cu foil', 'context': 'The blended slurry was… See the full description on the dataset page: https://huggingface.co/datasets/milerssliu/battery-device-data-qa.tabularquestion-answeringn<1K0 likes3 downloads6mo agoHugging Face15BattleTag /GeneratedBy_GPT4ogated Dataset Card for generatedBy GPT4o This dataset is generated by chatgpt-4o with documents, logs about cybersecurity area. Chatgpt is used to create prompt and response for training and testing based on provided content. Dataset Card Contact hychen3637@gmail.com textquestion-answering1K<n<10K0 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.