datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Assignment3aviation-gate-assignment-arrival-bank-coherence-risk-v0.1What this repo is for
Detect when arrivals land
but cannot park.
Flags
dense arrival bank with low gate availability
slow gate turns creating holding
remote stand use rising without bank pressure
gate holds as a downstream delay trigger
supply-chain-analysis-assignment
Supply Chain Disruption & Recovery Analysis
🎥 Presentation Video
📊 Project Overview
This project explores a dataset of 100,000 supply chain disruption events. The goal is to identify key factors influencing financial loss and recovery time.
Key Insights from EDA:
Costliest Disruption: Cyber Attacks result in the highest average revenue loss.
Production Impact: There is a strong correlation (0.76) between disruption severity and production impact.… See the full description on the dataset page: https://huggingface.co/datasets/IdoTreibatch/supply-chain-analysis-assignment.nlph-assignment1-deid-training-datanlp-assignment-news-data
Dataset Structure
Data Instances
{
"headline": "string",
"label": "string"
}
Data Fields
The data fields are:
headline: a string feature.
label: a classification label, with possible values including positive, negative and neutral.
Data Splits
The nlp-assignment-news-data dataset has 3 splits: train, validation, and test.
How to use it
from datasets import load_dataset
# This download train, validation and test sets.
ds =… See the full description on the dataset page: https://huggingface.co/datasets/hamza-student-123/nlp-assignment-news-data.Assignment_1_EDA
Bitcoin (BTC) Price Action & Technical Indicators Analysis
Video Presentation
Your browser does not support the video tag.
Project Overview
This research analyzes the "Multi-Model Trading Data" dataset, which consists of Bitcoin (BTC) historical trading data.
this data set haves 7.26K rows and 18 columns.
The Goal: To investigate the direct relationship between Bitcoin’s Price Movements and key technical indicators (Volume ,RSI, MACD, and Stoch RSI)… See the full description on the dataset page: https://huggingface.co/datasets/Yoel125/Assignment_1_EDA.assignment1OR-assignment-summariesqa_assignmentAssignmentDatasettraining_data_assignmentAssignmentemotion_detectAssignment_1B_Task_5
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/JosephFeig/Assignment_1B_Task_5.training_data_assignment_pretrainedEDA_Assignment
🎓 Student Dropout Prediction Dataset — EDA Assignment
By Tomer Bash | Data Science Course — Assignment #1
📹 Presentation Video
Presentation Video link - https://youtu.be/KyafBx9W7Qg
📌 Dataset Overview
Property
Details
Source
Kaggle
Rows
4,424 students
Features
35 columns
Target Variable
Target — Graduate, Enrolled, Dropout
Task Type
Multi-class Classification
The dataset contains demographic, financial, academic, and… See the full description on the dataset page: https://huggingface.co/datasets/Bashifu/EDA_Assignment.take_home_assignmentAssignment-1B-Datasetfinal.assignment
Exploratory Data Analysis (EDA) and Embeddings Evaluation
This dataset contains synthetic venue recommendations for London, generated to support a semantic recommendation system. Each record represents a recommended venue with metadata such as venue type, area, price level, and a natural language explanation.
Dataset Overview
Number of user queries: 10,000
Total recommended venues: 30,000
Unique venues: 513
Columns:
venue_name
venue_type
area
price_level
why… See the full description on the dataset page: https://huggingface.co/datasets/galsolomon9/final.assignment.eng-ai-assignment1b-task2Assignmentstourism-assignmentassignment-1-evaluation-resultseng-ai-assignment1b-kagglenoaa-wind-assignment2
NOAA Wind – 10-Day Slice (Assignment 2)
Station: 8724580Product: windFull query period: 20210601 → 20210630Chosen 10-day block (CSV): 2021-06-01 → 2021-06-10
This dataset contains hourly wind observations retrieved from the NOAA CO-OPS API.
API
Base: https://api.tidesandcurrents.noaa.gov/api/prod/datagetter
Parameters: station=8724580, product=wind, interval=h, units=english, time_zone=lst_ldt, format=json.
Files
data/wind_8724580_2021-06-01_2021-06-10.csv —… See the full description on the dataset page: https://huggingface.co/datasets/azizsi/noaa-wind-assignment2.Data_For_Assignment
Base Model Blind Spots: CohereLabs/tiny-aya-base
Model Tested
Model Name: CohereLabs/tiny-aya-base
This is a 3.35 billion parameter pretrained base model published just 12 days ago. It is an open-weights research release optimized for strong multilingual representation across 70+ languages. Because it is a base model designed for downstream instruction tuning, it predicts the next token based on raw web data rather than acting as a conversational assistant.… See the full description on the dataset page: https://huggingface.co/datasets/Soukaina588956468/Data_For_Assignment.
