J0nasW/paperswithcode
A cleaned dataset from paperswithcode.com Last dataset update: July 2023 This is a cleaned up dataset optained from paperswithcode.com through their API service. It represents a set of around 56K carefully categorized papers into 3K tasks and 16 areas. The papers contain arXiv and NIPS IDs as well as title, abstract and other meta information. It can be used for training text classifiers that concentrate on the use of specific AI and ML methods and frameworks.… See the full description on the dataset page: https://huggingface.co/datasets/J0nasW/paperswithcode.
2105
1taskID,pwc_id,title,description,areaID26fd143b2-e3dc-457a-a199-b8d671750e3c,data-poisoning,Data Poisoning,"**Data Poisoning** is an adversarial attack that tries to manipulate the training dataset in order to control the prediction behavior of a trained model such that the model will label malicious examples into a desired classes (e.g., labeling spam e-mails as safe).
3
4
5<span class=""description-source"">Source: [Explaining Vulnerabilities to Adversarial Machine Learning through Visual Analytics ](https://arxiv.org/abs/1907.07296)</span>",adversarial69f76c114-12a8-47fa-aa1f-8a680ec819e2,model-posioning,Model Posioning,,adversarial75fddb458-0a2d-435b-bcb0-322be02b3a48,dnn-testing,DNN Testing,Testing the reliability of DNNs.,adversarial88a2df849-20d2-47cd-ad63-3d4a698aafb3,provable-adversarial-defense,Provable Adversarial Defense,,adversarial9b9e36f78-ce57-4421-a42b-e8e26d39c0a8,backdoor-defense-for-data-free-distillation,Backdoor Defense for Data-Free Distillation with Poisoned Teachers,Defend against backdoor attack from poisoned teachers.,adversarial10111f3a46-923c-4580-9e22-58a0e076d3f0,phishing-website-detection,Phishing Website Detection,,adversarial11cbfecba9-553e-458c-a23b-25249ff518a9,website-fingerprinting-defense,Website Fingerprinting Defense,,adversarial129b2678f3-f3ed-44dc-a55c-78fd5d4f5fd0,backdoor-attack,Backdoor Attack,"Backdoor attacks inject maliciously constructed data into a training set so that, at test time, the trained model misclassifies inputs patched with a backdoor trigger as an adversarially-desired target class.",adversarial136d013445-7084-4675-892c-0f0c0829cc9d,adversarial-defense,Adversarial Defense,"Competitions with currently unpublished results:
14
15- [TrojAI](https://pages.nist.gov/trojai/)",adversarial16c445a68f-d8f7-49df-b1d3-c1e6322ed571,real-world-adversarial-attack,Real-World Adversarial Attack,Adversarial attacks that are presented in the real world,adversarial17fce9fa98-756e-4701-8003-159adc7c6a10,inference-attack,Inference Attack,,adversarial18b14f408c-6b59-4dee-935c-2ce10c601198,optimize-the-trajectory-of-uav-which-plays-a,Optimize the trajectory of UAV which plays a BS in communication system,,adversarial1936bcd0a3-4fec-4413-9cd7-ca7d78ec550b,adversarial-attack,Adversarial Attack,"An **Adversarial Attack** is a technique to find a perturbation that changes the prediction of a machine learning model. The perturbation can be very small and imperceptible to human eyes.20 21 22<span class=""description-source"">Source: [Recurrent Attention Model with Log-Polar Mapping is Robust against Adversarial Attacks ](https://arxiv.org/abs/2002.05388)</span>",adversarial23248fc705-85fc-4d29-878b-800a22f84a92,adversarial-text,Adversarial Text,"Adversarial Text refers to a specialised text sequence that is designed specifically to influence the prediction of a language model. Generally, Adversarial Text attack are carried out on Large Language Models (LLMs). Research on understanding different adversarial approaches can help us build effective defense mechanisms to detect malicious text input and build robust language models.",adversarial24c6924411-81d5-487d-a0f7-63b6706459af,design-synthesis,Design Synthesis,,adversarial25f243de31-4f70-4a28-b60e-21da65356e64,adversarial-robustness,Adversarial Robustness,Adversarial Robustness evaluates the vulnerabilities of machine learning models under various types of adversarial attacks.,adversarial26113c6061-e6bb-484c-ad69-4d97b985d3c4,exposure-fairness,Exposure Fairness,,adversarial27599853aa-d278-44b1-b358-a2e8f3640866,website-fingerprinting-attacks,Website Fingerprinting Attacks,,adversarial28c383316c-821c-420b-81b7-b0b327a6b6d2,model-extraction,Model extraction,"Model extraction attacks, aka model stealing attacks, are used to extract the parameters from the target model. Ideally, the adversary will be able to steal and replicate a model that will have a very similar performance to the target model.",adversarial296ba6acaa-d106-40bc-8cdc-f911d90fb8fa,audio-declipping,Audio declipping,Audio declipping is the task of estimating the original audio signal given its clipped measurements.,audio300e3930a8-378c-4074-999c-17dc325298bc,voice-conversion,Voice Conversion,"**Voice Conversion** is a technology that modifies the speech of a source speaker and makes their speech sound like that of another target speaker without changing the linguistic information.
31
32
33<span class=""description-source"">Source: [Joint training framework for text-to-speech and voice conversion using multi-source Tacotron and WaveNet ](https://arxiv.org/abs/1903.12389)</span>",audio344165f4be-9672-49ed-8d31-ed8ba985ab8e,chord-recognition,Chord Recognition,,audio35d34e2b02-c8f6-4047-97f3-c13bc7cfc868,bandwidth-extension,Bandwidth Extension,Bandwidth extension is the task of expanding the bandwidth of a signal in a way that approximates the original or desired higher bandwidth signal.,audio363bb03e25-85da-49dd-aa3e-58450bd88551,synthetic-speech-detection,Synthetic Speech Detection,Detect fake synthetic speech generated using machine learning,audio3719f79f8d-29e7-40a8-a5ae-02e9b792b259,audio-tagging,Audio Tagging,"Audio tagging is a task to predict the tags of audio clips. Audio tagging tasks include music tagging, acoustic scene classification, audio event classification, etc.",audio38f6ed6635-2302-47c5-afb8-e47ecfa21b93,directional-hearing,Directional Hearing,Extremely low-latency audio source separation from a known direction of arrival.,audio392a7caa61-07d0-4bc9-862c-342e602d956d,bird-audio-detection,Bird Audio Detection,,audio40561a4f04-5199-498c-ae64-bbd553563218,shooter-localization,Shooter Localization,Shooter localization based on videos.,audio419af95e17-c86e-4384-aaa7-33a0d03e133a,speaker-orientation,Speaker Orientation,Direction of Voice or speaker orientation of the person with respect to the target device.,audio42a195b1a7-342e-4610-bbb5-4317e7179566,audio-generation,Audio Generation,"Audio generation (synthesis) is the task of generating raw audio such as speech.
43
44<span style=""color:grey; opacity: 0.6"">( Image credit: [MelNet](https://arxiv.org/pdf/1906.01083v1.pdf) )</span>",audio45cfe1fce7-ac96-468a-b13c-f7024e60a0fa,sound-event-detection,Sound Event Detection,"**Sound Event Detection** (SED) is the task of recognizing the sound events and their respective temporal start and end time in a recording. Sound events in real life do not always occur in isolation, but tend to considerably overlap with each other. Recognizing such overlapping sound events is referred as polyphonic SED.46 47 48<span class=""description-source"">Source: [A report on sound event detection with different binaural features ](https://arxiv.org/abs/1710.02997)</span>",audio495feddbfc-e04c-44b7-b374-f2b783dad8f5,audio-signal-processing,Audio Signal Processing,"This is a general task that covers transforming audio inputs into audio outputs, not limited to existing PaperWithCode categories of Source Separation, Denoising, Classification, Recognition, etc.",audio5066606443-b911-4e5a-b3ec-38c0162f287a,direction-of-arrival-estimation,Direction of Arrival Estimation,Estimating the direction-of-arrival (DOA) of a sound source from multi-channel recordings.,audio514a332fbd-e4c6-47fe-9523-a1861ad61b09,soundscape-evaluation,Soundscape evaluation,Evaluation of soundscape in accordance to ISO/TS 12913-2,audio526026d958-26f0-4b30-bb24-2cbc6c8924ae,room-impulse-response,Room Impulse Response (RIR),"**Room Impulse Response (RIR)** is an audio signal processing task that involves capturing and analyzing the acoustic characteristics of a room or an environment. The goal is to measure and model the way sound waves interact with the space, including reflections, reverberation, and echoes.",audio53ee04f35b-a512-4d22-9a85-fa30ab6879b8,vowel-classification,Vowel Classification,,audio5488c2e3a6-7ce0-41ec-a132-8fce7643112a,audio-fingerprint,Audio Fingerprint,,audio558df66d85-353e-4f5f-a028-84249de34876,pitch-control,Pitch control,,audio56f733071e-35e2-450f-839d-ffe6adced489,underwater-acoustic-classification,Underwater Acoustic Classification,Classification of underwater acoustic data,audio57a4930f8b-c5b5-49a8-b0b5-063b8d10b41c,active-speaker-localization,Active Speaker Localization,"Active Speaker Localization (ASL) is the process of spatially localizing an active speaker (talker) in an environment using either audio, vision or both.",audio584ca24c2f-1abe-4606-b72f-4135b0baefd3,acoustic-scene-classification,Acoustic Scene Classification,"The goal of acoustic scene classification is to classify a test recording into one of the provided predefined classes that characterizes the environment in which it was recorded.
59
60Source: [DCASE 2019](http://dcase.community/challenge2019/task-acoustic-scene-classification)
61Source: [DCASE 2018](https://dcase.community/challenge2018/task-acoustic-scene-classification)",audio6248356260-c73f-481d-819f-1cb9ccd2a4d4,timbre-interpolation,Timbre Interpolation,,audio63c13c7126-bd53-455c-aeaa-9e66569928d5,audio-denoising,Audio Denoising,,audio64db9465ba-8c3a-4462-84f1-a430feb25215,sound-event-localization-and-detection,Sound Event Localization and Detection,"Given multichannel audio input, a sound event detection and localization (SELD) system outputs a temporal activation track for each of the target sound classes, along with one or more corresponding spatial trajectories when the track indicates activity. This results in a spatio-temporal characterization of the acoustic scene that can be used in a wide range of machine cognition tasks, such as inference on the type of environment, self-localization, navigation without visual input or with occluded targets, tracking of specific types of sound sources, smart-home applications, scene visualization systems, and audio surveillance, among others.",audio657ea55ca2-f7fe-4b1b-ae2f-ca23aa3242ed,audio-multiple-target-classification,Audio Multiple Target Classification,,audio66e95aaace-7fbb-4d0e-bf07-cc1ea8c2c5e8,bird-species-classification-with-audio-visual,Bird Species Classification With Audio-Visual Data,,audio678708ffc6-dab4-47aa-bfc3-d999bac28bc3,audio-captioning,Audio captioning,,audio682843e13c-9847-43bf-abb5-703a3f21f259,audio-inpainting,Audio inpainting,Filling in holes in audio data,audio6967d41471-ddbf-48c2-8071-3a75ce916379,bird-classification,Bird Classification,,audio70829bc069-f39c-45ff-854e-f17ff108abcd,music-compression,Music Compression,,audio7169524fb8-047d-4a6c-a2b4-b512767874e7,gunshot-detection,Gunshot Detection,,audio728f67593e-1263-4a59-9067-c22270bc4c55,language-identification,Language Identification,Language identification is the task of determining the language of a text.,audio73fd43566c-7133-4956-bd1c-a91e9636276e,instrument-recognition,Instrument Recognition,,audio74e202d49a-33ee-4d68-9509-994b58fca9fd,audio-super-resolution,Audio Super-Resolution,AUDIO SUPER-RESOLUTION or speech bandwidth extension (Upsampling Ratio = 2),audio75680b1bec-9873-42c2-802d-183552d2f438,target-speaker-extraction,Target Speaker Extraction,Extract the dialogue content of the specified target in a multi-person dialogue.,audio76754c67b9-747b-45ab-b324-67ac73231e91,audio-classification,Audio Classification,"**Audio Classification** is a machine learning task that involves identifying and tagging audio signals into different classes or categories. The goal of audio classification is to enable machines to automatically recognize and distinguish between different types of audio, such as music, speech, and environmental sounds.",audio77aefa41ac-5fce-4299-8939-342fce8d2511,environmental-sound-classification,Environmental Sound Classification,Classification of Environmental Sounds. Most often sounds found in Urban environments. Task related to noise monitoring.,audio7850dd1159-d7f2-4d93-aa50-6cd8334ff5dd,fake-voice-detection,fake voice detection,,audio79d1c97289-6799-4b81-894a-17d16fbd3b5d,audio-source-separation,Audio Source Separation,"**Audio Source Separation** is the process of separating a mixture (e.g. a pop band recording) into isolated sounds from individual sources (e.g. just the lead vocals).
80
81
82<span class=""description-source"">Source: [Model selection for deep audio source separation via clustering analysis ](https://arxiv.org/abs/1910.12626)</span>",audio830af1b5a6-e530-44b2-b05e-071299a35e4e,audio-signal-recognition,Audio Signal Recognition,,audio84bad4a297-3f72-4d3e-b3e1-4ed5c9c7a791,text-to-music-generation,Text-to-Music Generation,,audio856a894212-022e-40e2-acfe-36b1c864776c,audio-effects-modeling,Audio Effects Modeling,"Modeling of audio effects such as reverberation, compression, distortion, etc.",audio868846f94c-b4c1-43aa-b2bb-21bcdcf36f01,few-shot-audio-classification,Few-Shot Audio Classification,Few-shot classification for audio signals. Presents a unique challenge compared to other few-shot domains as we deal with temporal dependencies as well,audio877e2c45e2-fcb5-4e23-83c1-2f8235c3a041,sound-classification,Sound Classification,,audio88f240f9ae-77e5-401d-8b2b-0c437bd67844,audio-visual-synchronization,Audio-Visual Synchronization,,audio89b475e807-f49a-45be-91bd-d1235118435c,zero-shot-multi-speaker-tts,Zero-Shot Multi-Speaker TTS,,audio904cf0d042-a585-45f0-b4ab-e5c79f1948c1,music-generation,Music Generation,"**Music Generation** is the task of generating music or music-like sounds from a model or algorithm. The goal is to produce a sequence of notes or sound events that are similar to existing music in some way, such as having the same style, genre, or mood.",audio919a32a5e5-503c-42ef-8b7c-8190ac8d585d,streaming-target-sound-extraction,Streaming Target Sound Extraction,"This task is a variant of the [Target Sound Extraction](https://paperswithcode.com/task/target-sound-extraction) task, with the constraint of causal streaming inference. Aiming for an algorithmic latency of less than 20 ms, at each time step, streaming audio models operate on an input audio chunk of length less than 20 ms. The causal constraint means that the model only has the knowledge of past chunks and no future chunks.",audio92999913fe-91a0-4917-b9a9-74d4a43c4165,real-time-directional-hearing,Real-time Directional Hearing,Directional hearing models that also support real-time on-device inference,audio93587ca093-54b0-487b-b699-444e593f7127,voice-anti-spoofing,Voice Anti-spoofing,Discriminate genuine speech and spoofing attacks,audio94844b362b-5213-4a85-b6b7-128252fbb22c,audio-dequantization,Audio Dequantization,Audio Dequantization is a process of estimating the original signal from its quantized counterpart.,audio951b199a3c-ad07-4fb9-8716-26bcd700a84a,single-label-target-sound-extraction,Single-Label Target Sound Extraction,"Single-Label Target Sound Extraction is the task of extracting a given class of sounds from an audio mixture. The audio mixture may contain background noise with a relatively low amplitude compared to the foreground mixture components. The choice of the sound class is provided as input to the model in form of a string, integer, or a one-hot encoding of the sound class.",audio964ec78a95-37bc-4f77-8c10-47b70e47eb39,target-sound-extraction,Target Sound Extraction,"Target Sound Extraction is the task of extracting a sound corresponding to a given class from an audio mixture. The audio mixture may contain background noise with a relatively low amplitude compared to the foreground mixture components. The choice of the sound class is provided as input to the model in form of a string, integer, or a one-hot encoding of the sound class.",audio978d54025e-e889-4323-9c14-9a5b18ecef29,acoustic-novelty-detection,Acoustic Novelty Detection,"Detect novel events given acoustic signals, either in domestic or outdoor environments.",audio986cb0994c-edc0-4134-815f-828476da0906,self-supervised-sound-classification,Self-Supervised Sound Classification,,audio9997dddbfd-b0de-489d-8acb-e265993711b1,inference-optimization,Inference Optimization,,audio1006a0fc811-cd11-41ea-8483-b43384b313b0,fault-localization,Fault localization,,computer-code101b7123449-480d-483f-983b-e2bbdc3a49b9,nmt,NMT,"Neural machine translation is an approach to machine translation that uses an artificial neural network to predict the likelihood of a sequence of words, typically modeling entire sentences in a single integrated model.",computer-code10275054394-b88e-47b8-b1b0-b8284f044022,code-generation,Code Generation,"**Code Generation** is an important field to predict explicit code or program structure from multimodal data sources such as incomplete code, programs in another programming language, natural language descriptions or execution examples. Code Generation tools can assist the development of automatic programming tools to improve programming productivity.
103
104
105<span class=""description-source"">Source: [Deep Learning for Source Code Modeling and Generation ](https://arxiv.org/abs/2002.05442)</span>
106
107Image source: [Measuring Coding Challenge Competence With APPS](https://paperswithcode.com/paper/measuring-coding-challenge-competence-with)",computer-code108c545b464-3b82-47d9-9cda-451e1004cfc4,compiler-optimization,Compiler Optimization,Machine learning guided compiler optimization,computer-code10919a10edb-2e12-4e19-af0e-8cd08184f746,write-computer-programs-from-specifications,Write Computer Programs From Specifications,,computer-code1105e076123-b1ea-4d08-b4e7-c5a715c0020e,api-sequence-recommendation,API Sequence Recommendation,,computer-code1113214205e-7197-4fe6-930e-b937df6f53ac,code-summarization,Source Code Summarization,"**Code Summarization** is a task that tries to comprehend code and automatically generate descriptions directly from the source code.
112
113
114<span class=""description-source"">Source: [Improving Automatic Source Code Summarization via Deep Reinforcement Learning ](https://arxiv.org/abs/1811.07234)</span>",computer-code11558848877-3c4b-46d9-85b2-228c41f09e5d,contextual-embedding-for-source-code,Contextual Embedding for Source Code,,computer-code11678d3301a-61ff-42ab-a2b8-ad3da3d3dac1,formalize-foundations-of-universal-algebra-in,Formalize foundations of universal algebra in dependent type theory,,computer-code117d20340c9-b513-47f0-bdc6-17e13f1be88e,sentinel-1-sar-processing,Sentinel-1 SAR processing,,computer-code11884bf8080-b8a2-4427-b499-f12b592461d5,enumerative-search,Enumerative Search,,computer-code119f47512f2-93b6-4870-ad7b-c127aa56b764,type-prediction,Type prediction,,computer-code120f01eb769-bfc3-44bc-a962-cccbc6810df1,code-translation,Code Translation,"Code translation is the process of converting code written in one programming language to another programming language while maintaining the same functionality. This process is also known as code conversion, source-to-source translation, or transpilation. Code translation is often performed when developers want to take advantage of new programming languages, improve code performance, or maintain legacy systems. Some common examples include translating code from Python to Java, or from JavaScript to TypeScript.",computer-code121faee1947-6132-4107-9d9b-0c4de77d7e7c,value-prediction,Value prediction,,computer-code12259fd5ddc-9794-4ff7-b331-9e122c81b6ba,program-repair,Program Repair,Task of teaching ML models to modify an existing program to fix a bug in a given code.,computer-code1239c9d0413-4d42-40a7-9297-bea6aa9981ad,variable-misuse,Variable misuse,,computer-code124e4be7a9c-4edc-473b-9e28-c8dc1d737593,nature-inspired-optimization-algorithm,Nature-Inspired Optimization Algorithm,,computer-code125e9c67180-bdf1-41e3-bcd7-5b7d71650315,git-commit-message-generation,Git Commit Message Generation,,computer-code126d9a0706a-14a3-4532-9eb9-03b414cee8fd,tiling-deployment,Tiling & Deployment,Data tiling over 3 memory hierarchy levels and deployment on microcontroller.,computer-code127d7ef25b8-2e06-4e5d-a5e6-6adeb99f8319,single-image-portrait-relighting,Single-Image Portrait Relighting,,computer-code128b161dcff-a1d9-4a07-b7cb-08d83395acf7,code-classification,Code Classification,,computer-code12998419e53-e9ec-46ab-ad8d-e25255867da1,sparse-subspace-based-clustering,Sparse subspace-based clustering,,computer-code1307d08fd6c-b509-4f27-8d16-386deda40095,learning-to-execute,Learning to Execute,,computer-code1311d664fdd-7531-4096-9924-84505629695c,exception-type,Exception type,,computer-code132d3b06eb8-117c-42d5-acb2-63c6d6bf1769,neural-network-simulation,Neural Network simulation,Simulation of abstract or biophysical neural networks in silico,computer-code13310932eaa-8f78-467d-ab3e-7ee88e657bee,edit-script-generation,Edit script generation,"Generating edit scripts by comparing 2 different files or strings to convert one to another. this script will contain instruction like insert, delete and substitute.",computer-code134bfd8624a-a1aa-4a9e-a909-36203d2f157c,low-rank-compression,Low-rank compression,,computer-code135e12aad32-cad2-4d13-9a10-858777de85b0,chart-question-answering,Chart Question Answering,Question Answering task on charts images,computer-code136c069c9b2-2a5d-499f-9a12-159aa6395942,swapped-operands,Swapped operands,,computer-code137ce8d93c0-a4da-407f-8716-bab66d36b832,function-docstring-mismatch,Function-docstring mismatch,,computer-code13806837c7e-0f2a-4d9a-97d7-a96e5c1f1656,webcam-rgb-image-classification,Webcam (RGB) image classification,,computer-code139ab6bc3b7-9161-4dcd-842f-87f01f903c0a,codesearchnet-java,CodeSearchNet - Java,,computer-code140931d6c91-f3a1-4161-9430-d358dbca7631,code-documentation-generation,Code Documentation Generation,"Code Documentation Generation is a supervised task where a code function is the input to the model, and the model generates the documentation for this function.
141
142Description from: [CodeTrans: Towards Cracking the Language of Silicone's Code Through Self-Supervised Deep Learning and High Performance Computing](https://arxiv.org/pdf/2104.02443.pdf)",computer-code143a0031fff-0de2-43db-9f82-2aae2fa3841e,editcompletion,EditCompletion,"Given a code snippet that is partially edited, the goal is to predict a completion of the edit for the rest of the snippet.",computer-code14497f19c78-ae0e-4c10-a29c-9337648c6693,motion-style-transfer,Motion Style Transfer,,computer-code14509c72d3d-e4d7-489c-b389-ca859c74c2fa,file-difference,File difference,"Generate edit script comparing 2 strings or files, which contains instruction of insert, delete and substitute to convert first string to the second.",computer-code14698ab73fa-eb91-4fcc-b5cd-0133ca4d0658,video-defect-classification,Video Defect Classification,"Quick-View (QV) Inspection is one commonly-used technology. However, it is quite labor-intensive to find defects from a huge number of QV videos. To tackle this problem, we propose a video defect classification task, which is to predict the categories of pipe
147defects in a short QV video.",computer-code1482989d89f-6353-44be-9d79-04c30cb7a830,spectral-efficiency-analysis-of-uplink,Spectral Efficiency Analysis of Uplink-Downlink Decoupled Access in C-V2X Networks,Code for Spectral Efficiency Analysis of Uplink-Downlink Decoupled Access in C-V2X Networks,computer-code149ae0cbbd0-4a0a-439c-a5c4-10b3185e2213,sql-to-text,SQL-to-Text,"<span style=""color:grey; opacity: 0.6"">( Image credit: [SQL-to-Text Generation with Graph-to-Sequence Model](https://arxiv.org/pdf/1809.05255v2.pdf) )</span>",computer-code1503c3c8aa2-ad5c-4c0b-9167-957092eaa55c,program-synthesis,Program Synthesis,"Program synthesis is the process of automatically generating a program or code snippet that satisfies a given specification or set of requirements. This can include generating code from a formal specification, a natural language description, or example inputs and outputs. The primary goal of program synthesis is to minimize human intervention in the coding process, reduce errors, and improve productivity.
151
152Program synthesis often involves the use of advanced algorithms, artificial intelligence, and machine learning techniques to search the space of possible programs that meet the given constraints. This process can be guided by a variety of techniques, such as constraint solving, symbolic execution, and genetic algorithms.",computer-code153964bd246-ee2b-4d83-a5ff-f15c5c3f0ea2,log-parsing,Log Parsing,"**Log Parsing** is the task of transforming unstructured log data into a structured format that can be used to train machine learning algorithms. The structured log data is then used to identify patterns, trends, and anomalies, which can support decision-making and improve system performance, security, and reliability. The log parsing process involves the extraction of relevant information from log files, the conversion of this information into a standardized format, and the storage of the structured data in a database or other data repository.",computer-code154a19f918a-9489-49b0-a094-17ca6de850bf,program-induction,Program induction,Generating program code for domain-specific tasks,computer-code155c7790180-feb0-43f5-b48c-f58f53a5cd29,programming-error-detection,Programming Error Detection,,computer-code15663bcab46-013f-4c52-94f1-a189175799f2,text-to-sql,Text-To-SQL,"**Text-to-SQL** is a task in natural language processing (NLP) where the goal is to automatically generate SQL queries from natural language text. The task involves converting the text input into a structured representation and then using this representation to generate a semantically correct SQL query that can be executed on a database.
157
158<span style=""color:grey; opacity: 0.6"">( Image credit: [SyntaxSQLNet](https://arxiv.org/pdf/1810.05237v2.pdf) )</span>",computer-code159046d9b48-a71d-48f1-a1a7-5d881c7d74e8,sql-synthesis,SQL Synthesis,,computer-code160f3f0fadc-f3dd-41c6-b6af-a83c786e8fd6,text-to-code-generation,Text-to-Code Generation,"**Text-to-Code Generation** is a task where we can generate code based on the natural language description.
161
162Source: [Text-to-code Generation with TensorFlow, 🤗 & MBPP](https://www.kaggle.com/code/rhtsingh/text-to-code-generation-with-tensorflow-mbpp)",computer-code1636bfb15cd-9e35-4237-a344-aa82d37b302b,paraphrase-generation,Paraphrase Generation,"Paraphrase Generation involves transforming a natural language sentence to a new sentence, that has the same semantic meaning but a different syntactic or lexical surface form.",computer-code16406092d2d-2132-486b-b29a-dd6c1b1940b2,sql-chatbots,Sql Chatbots,,computer-code1650389d435-5263-44bf-9007-de7bbdde59b9,code-search,Code Search,"The goal of **Code Search** is to retrieve code fragments from a large code corpus that most closely match a developer’s intent, which is expressed in natural language.
166
167
168<span class=""description-source"">Source: [When Deep Learning Met Code Search ](https://arxiv.org/abs/1905.03813)</span>",computer-code1695fafca91-f1e5-4963-a28f-e42e29bb7b76,wrong-binary-operator,Wrong binary operator,,computer-code170204a46eb-75a6-4ac4-84d5-cdd0d1ec71a1,code-comment-generation,Code Comment Generation,,computer-code171fc88936a-738b-47d7-8cdd-227f95f86c68,annotated-code-search,Annotated Code Search,Annotated code search is the retrieval of code snippets paired with brief descriptions of their intent using natural language queries.,computer-code1720e91d5bd-a029-4c3e-8ef7-208168a52e26,manufacturing-quality-control,Manufacturing Quality Control,AI for Quality control in manufacturing processes.,computer-vision1731b3e56c7-7bad-4419-a408-8b0d268a5fdb,fine-grained-image-recognition,Fine-Grained Image Recognition,,computer-vision174a259573b-8d0d-43a0-bd45-430123eb1a07,camouflage-segmentation,Camouflage Segmentation,,computer-vision175976463fe-5f4e-43b6-a1af-4c5884e107a3,layout-design,Layout Design,,computer-vision176ab6a11d2-2972-41f0-8c68-b46e2ddcaa63,video-propagation,Video Propagation,Propagating information in processed frames to unprocessed frames,computer-vision177bf563485-7d8a-4dc6-84b5-41e11dcbcd97,few-shot-temporal-action-localization,Few Shot Temporal Action Localization,Detect Action using few labeled samples,computer-vision178c2c8912d-119e-43ee-a464-9acb2d87509b,one-shot-object-detection,One-Shot Object Detection,"<span style=""color:grey; opacity: 0.6"">( Image credit: [Siamese Mask R-CNN
179](https://github.com/bethgelab/siamese-mask-rcnn) )</span>",computer-vision1809958c075-b8e0-4938-baa5-72f9ebb2dac5,color-manipulation,color manipulation,,computer-vision1815024747a-5530-43eb-9b6c-27592f836d80,single-class-few-shot-image-synthesis,Single class few-shot image synthesis,"The goal of single class few-shot image synthesis task is to learn a generative model that can generate
182samples with visual attributes from as few as two or more input images images belonging to the same class.",computer-vision183c208b637-f319-447b-9355-a42f680ca884,sperm-morphology-classification,Sperm Morphology Classification,Multi-class classification of sperm head morphology.,computer-vision184f99bf7fa-d464-4abe-9c78-f777bed1ccab,salt-and-pepper-noise-removal,Salt-And-Pepper Noise Removal,"Salt-and-pepper noise is a form of noise sometimes seen on images. It is also known as impulse noise. This noise can be caused by sharp and sudden disturbances in the image signal. It presents itself as sparsely occurring white and black pixels.
185
186<span style=""color:grey; opacity: 0.6"">( Image credit: [NAMF](https://arxiv.org/pdf/1910.07787v1.pdf) )</span>",computer-vision187c0e9a69a-fee6-4524-800f-104ea7bad6d7,style-transfer,Style Transfer,"**Style Transfer** is a technique in computer vision and graphics that involves generating a new image by combining the content of one image with the style of another image. The goal of style transfer is to create an image that preserves the content of the original image while applying the visual style of another image.
188
189<span style=""color:grey; opacity: 0.6"">( Image credit: [A Neural Algorithm of Artistic Style](https://arxiv.org/pdf/1508.06576v2.pdf) )</span>",computer-vision190a9173b27-50ae-49ba-97a8-c3dd064a3269,intensity-image-denoising,intensity image denoising,,computer-vision191b83ddd70-0cd9-4397-8788-996b5e27f738,table-recognition,Table Recognition,,computer-vision192cdd45765-8ad4-49ca-9e6c-f940a9bc91a6,interest-point-detection,Interest Point Detection,,computer-vision193f7d273ea-b979-4da0-aa7a-055cf5d2d820,lossy-compression-artifact-reduction,Lossy-Compression Artifact Reduction,,computer-vision1940ccc3c6e-95f2-4e5b-abf8-2a19cd3e2b70,text-based-image-editing,Text-based Image Editing,,computer-vision19523799022-93d0-4bde-9012-0ae14d607691,3d-face-animation,3D Face Animation,Image: [Cudeiro et al](https://arxiv.org/pdf/1905.03079v1.pdf),computer-vision196cfadd6e5-3bdd-42b8-8f28-c48cf68209d5,physiological-computing,Physiological Computing,,computer-vision197295509b1-485a-41a3-8034-874a8ab97592,unsupervised-object-localization,Unsupervised Object Localization,,computer-vision198199246ce-afca-48a2-9d89-089d1e1d66d6,traffic-sign-recognition,Traffic Sign Recognition,"Traffic sign recognition is the task of recognising traffic signs in an image or video.
199
200<span style=""color:grey; opacity: 0.6"">( Image credit: [Novel Deep Learning Model for Traffic Sign Detection Using Capsule
201Networks ](https://arxiv.org/pdf/1805.04424v1.pdf) )</span>",computer-vision202df31dc86-76af-4e21-a47c-97c54fd94113,facial-recognition-and-modelling,Facial Recognition and Modelling,Facial tasks in machine learning operate based on images or video frames (or other datasets) focussed on human faces.,computer-vision203fd6a4d2d-2c82-4e31-8c5f-68466c1f51a0,video-compressive-sensing,Video Compressive Sensing,,computer-vision204a0af1805-bc94-457d-b671-7c67fded3feb,semantic-part-detection,Semantic Part Detection,,computer-vision205f42db817-aa22-495b-bf69-93d873e3e0d9,unsupervised-3d-semantic-segmentation,Unsupervised 3D Semantic Segmentation,Unsupervised 3D Semantic Segmentation,computer-vision206ac67f6cd-349e-4874-9965-826851c2f914,future-hand-prediction,Future Hand Prediction,,computer-vision20792663016-88d8-4095-baa1-5b2147142277,matching-disparate-images,Matching Disparate Images,,computer-vision208528e9c82-8046-4835-942b-208ce5d888db,3d-feature-matching,3D Feature Matching,Image: [Choy et al](https://paperswithcode.com/paper/fully-convolutional-geometric-features),computer-vision2098126ae3d-69dd-4da7-bedc-8b0ccfd0f86b,video-editing,Video Editing,,computer-vision2108899ac03-ebcf-423c-925f-21bd5802fb75,spectral-reconstruction,Spectral Reconstruction,,computer-vision211ae9cd9a5-5b66-44ba-b10a-7318c8c8c40c,depth-image-estimation,Depth Image Estimation,,computer-vision212d58fc61f-40bc-4c00-9c91-3853117c907f,multi-hypotheses-3d-human-pose-estimation,Multi-Hypotheses 3D Human Pose Estimation,,computer-vision2130be19bff-e7fb-419d-8c06-f9d0e56ae6d7,svbrdf-estimation,SVBRDF Estimation,SVBRDF Estimation,computer-vision21444776105-7a6d-43f5-b363-dbca3660afa2,document-enhancement,Document Enhancement,,computer-vision215945ada03-cc7a-4f03-824a-2a3032c0d932,robust-face-alignment,Robust Face Alignment,"Robust face alignment is the task of face alignment in unconstrained (non-artificial) conditions.
216
217<span style=""color:grey; opacity: 0.6"">( Image credit: [Deep Alignment Network](https://github.com/MarekKowalski/DeepAlignmentNetwork) )</span>",computer-vision218e0531d71-eeca-4a97-be2e-6194dabb2737,rice-grain-disease-detection,Rice Grain Disease Detection,,computer-vision219980e490b-c80e-4ca0-912a-27b368725e23,offline-surgical-phase-recognition,Offline surgical phase recognition,"Offline surgical phase recognition: the first 40 videos to train, the last 40 videos to test.",computer-vision22095cd9793-d1d5-42e2-913a-4cdc9c4bbf49,explainable-models,Explainable Models,,computer-vision221a1af98e8-122d-4a57-bee3-799300dc7e0f,kinship-face-generation,Kinship face generation,Kinship face generation,computer-vision22248098bc6-f28c-453b-982e-a5582b4fed5e,sensor-modeling,Sensor Modeling,"<span style=""color:grey; opacity: 0.6"">( Image credit: [LiDAR Sensor modeling and Data augmentation with GANs for Autonomous driving](https://arxiv.org/abs/1905.07290) )</span>",computer-vision223371344a1-3fcc-400e-85b6-1a034b261248,face-quality-assessement,Face Quality Assessement,Estimate the usability of a given face image for recognition,computer-vision2244e0c0c5d-a920-477e-a5cb-c1bb3adb6b82,ifc-entity-classification,IFC Entity Classification,,computer-vision22569107063-fb16-4fe8-a948-e0899d25a565,fashion-compatibility-learning,Fashion Compatibility Learning,,computer-vision226a7309764-d1c7-4316-a05a-3300e44b360d,robust-face-recognition,Robust Face Recognition,"Robust face recognition is the task of performing recognition in an unconstrained environment, where there is variation of view-point, scale, pose, illumination and expression of the face images.
227
228<span style=""color:grey; opacity: 0.6"">( Image credit: [MeGlass dataset](https://github.com/cleardusk/MeGlass) )</span>",computer-vision22966b392ce-d32b-42b3-b883-8bf4d5f0126b,color-image-denoising,Color Image Denoising,,computer-vision230a43db77c-38a0-48a9-8f2d-e5fbd6f75d44,adversarial-attack-detection,Adversarial Attack Detection,The detection of adversarial attacks.,computer-vision231b8f48755-6116-4204-8a26-83fe84141f68,dense-object-detection,Dense Object Detection,,computer-vision2321c867704-6058-4853-8731-7d22169b170e,video-style-transfer,Video Style Transfer,,computer-vision233877be219-cec3-41af-91c7-4e009a583454,audio-visual-video-captioning,Audio-Visual Video Captioning,,computer-vision23422a42437-c945-4d1a-9bb8-eea0af95149b,weakly-supervised-panoptic-segmentation,Weakly-supervised panoptic segmentation,,computer-vision235e444f779-29fe-4691-ab4d-00e575e40fda,satellite-image-classification,Satellite Image Classification,"Satellite image classification is the most significant technique used in remote sensing for the computerized study and pattern recognition of satellite information, which is based on diversity structures of the image that involve rigorous validation of the training samples depending on the used classification algorithm.",computer-vision2365fca84f4-032f-41f7-b7cc-dae30d7cb4c0,camera-calibration,Camera Calibration,"Camera calibration involves estimating camera parameters(including camera intrinsics and extrinsics) to infer geometric features from captured sequences, which is crucial for computer vision and robotics. Driven by different architectures of the neural network,
237the researchers have developed two main paradigms for learning-based camera calibration and its applications. One is Regression-based Calibration,Reconstruction-based Calibration is another.",computer-vision23819691350-7c8f-4b13-a0af-a76d6db3da27,face-presentation-attack-detection,Face Presentation Attack Detection,,computer-vision2395130d8dd-f687-41db-b9a8-490fd4c87442,shadow-removal,Shadow Removal,Remove shadow from background,computer-vision240a71b5dab-267d-4c86-9f40-85e64403f5ed,prostate-zones-segmentation,Prostate Zones Segmentation,,computer-vision2414a2ac0d8-b946-40c9-a450-d146ffd98500,template-matching,Template Matching,,computer-vision242da62d84f-4fd9-4c9c-aa3e-4c22f16704a1,zero-shot-action-recognition,Zero-Shot Action Recognition,,computer-vision243f77daf68-708c-4c47-84b8-a29c5188a0b9,motion-prediction,motion prediction,,computer-vision244b8cf8c5c-e1ca-4a6c-8347-dbe4857fd7fe,referring-image-matting-refmatte-rw100,Referring Image Matting (RefMatte-RW100),"Expression-based referring image matting on natural images and manually labelled annotations, i.e., RefMatte-RW100, taking the image and a flowery expression as the input.",computer-vision2455e6d2df1-a69f-440e-ae76-81972a889783,video-visual-relation-detection,Video Visual Relation Detection,"**Video Visual Relation Detection (VidVRD)** aims to detect instances of visual relations of interest in a video, where a visual relation instance is represented by a relation triplet <subject, predicate, object> with the trajectories of the subject and object. As compared to still images, videos provide a more natural set of features for detecting visual relations, such as the dynamic relations like “A-follow-B” and “A-towards-B”, and temporally changing relations like “A-chase-B” followed by “A-hold-B”. Yet, VidVRD is technically more challenging than ImgVRD due to the difficulties in accurate object tracking and diverse relation appearances in the video domain.
246
247<span class=""description-source"">Source: [ImageNet-VidVRD Video Visual Relation Dataset](https://xdshang.github.io/docs/imagenet-vidvrd.html)</span>",computer-vision248fee08922-362d-41d6-8cb6-675db54ba472,person-re-identification,Person Re-Identification,"**Person Re-Identification** is a computer vision task in which the goal is to match a person's identity across different cameras or locations in a video or image sequence. It involves detecting and tracking a person and then using features such as appearance, body shape, and clothing to match their identity in different frames. The goal is to associate the same person across multiple non-overlapping camera views in a robust and efficient manner.",computer-vision249aac03e3f-e00a-4918-bde6-f89cebf37dc1,mutual-gaze,Mutual Gaze,Detect if two people are looking at each other,computer-vision25033d9acbe-f338-4007-b382-4a205a2f1f06,affordance-recognition,Affordance Recognition,Affordance recognition from Human-Object Interaction,computer-vision2514a3f8369-ff14-42ad-a1fb-074f77bd9f39,scene-parsing,Scene Parsing,"Scene parsing is to segment and parse an image into different image regions associated with semantic categories, such as sky, road, person, and bed. [MIT Description](http://sceneparsing.csail.mit.edu/#:~:text=Scene%20parsing%20is%20to%20segment,the%20algorithms%20of%20scene%20parsing.)",computer-vision252054df625-1176-4676-ac27-7a0e9158bedd,3d-shape-retrieval,3D Shape Classification,Image: [Sun et al](https://arxiv.org/pdf/1804.04610v1.pdf),computer-vision253b8bc11cb-f902-4727-8e03-c93f3cf6198d,semi-supervised-learning-for-image-captioning,Semi Supervised Learning for Image Captioning,,computer-vision254850fe1ca-6a05-4dbf-84c9-810dae4eb6ca,3d-character-animation-from-a-single-photo,3D Character Animation From A Single Photo,Image: [Weng et al](https://arxiv.org/pdf/1812.02246v1.pdf),computer-vision255b9ef01c9-01af-4413-9380-79cf4c8bc5f3,fine-grained-image-inpainting,Fine-Grained Image Inpainting,,computer-vision25646d2d68f-146d-40f3-8535-96b516753f90,kinematic-based-workflow-recognition,Kinematic Based Workflow Recognition,,computer-vision257723dc4a2-4bf5-4308-919c-c2a5630befe4,3d-point-cloud-reconstruction,3D Point Cloud Reconstruction,Encoding and reconstruction of 3D point clouds.,computer-vision258cebe8ba3-21a5-426d-8af0-c87e50a73c91,3d-instance-segmentation-1,3D Instance Segmentation,Image: [OccuSeg](https://arxiv.org/pdf/2003.06537v3.pdf),computer-vision2596b09bb39-64ff-4185-92ea-12d33274014a,visual-question-answering,Visual Question Answering (VQA),"**Visual Question Answering (VQA)** is a task in computer vision that involves answering questions about an image. The goal of VQA is to teach machines to understand the content of an image and answer questions about it in natural language.
260
261Image Source: [visualqa.org](https://visualqa.org/)",computer-vision262ec37baec-50ef-4cad-8a65-1ca06f664a7e,bbbc021-nsc-accuracy,BBBC021 NSC Accuracy,"BBBC021 is a dataset of fully imaged human cells. Cells are treated with one of 113 small molecules at 8 concentrations, and fluorescent images are captured staining for nucleus, actin and microtubules. The phenotypic profiling problem is presented, where the goal is to extract features containing meaningful information about the cellular phenotype exhibited. Each of 103 unique compound concentration treatment is labeled with a mechanism-of-action (MOA). The MOA is predicted for each unique treatment (averaging features over all treatment examples) by matching the MOA of the closest point excluding points of the same compound. The dataset and more information can be found at https://bbbc.broadinstitute.org/BBBC021.",computer-vision263416172ad-ac7b-4507-b893-25857b06fc14,unsupervised-video-summarization,Unsupervised Video Summarization,"**Unsupervised video summarization** approaches overcome the need for ground-truth data (whose production requires time-demanding and laborious manual annotation procedures), based on learning mechanisms that require only an adequately large collection of original videos for their training. Specifically, the training is based on heuristic rules, like the sparsity, the representativeness, and the diversity of the utilized input features/characteristics.",computer-vision26474804cb0-bfea-4681-9f38-2b47599e887e,age-and-gender-estimation,Age and Gender Estimation,Age and gender estimation is a dual-task of identifying the age via regression analysis and classification of gender of a person.,computer-vision265a36d6d88-4d0f-469d-a780-ae938b6eb8b8,hyperspectral-image-segmentation,Hyperspectral Image Segmentation,,computer-vision266bc7706ad-b0d8-4d35-ae22-857ff82b660c,font-style-transfer,Font Style Transfer,**Font style transfer** is the task of converting text written in one font into text written in another font while preserving the meaning of the original text. It is used to change the appearance of text while keeping its content intact.,computer-vision267ba03b4b0-e787-4958-8cb3-b133bea8a5a3,3d-dense-shape-correspondence,3D Dense Shape Correspondence,"Finding a meaningful correspondence between two or more shapes is one of the most fundamental shape analysis tasks. The problem can be generally stated as: given input shapes S1,S2,...,SN, find a meaningful relation (or mapping) between their elements. Under different contexts, the problem has also been referred to as registration, alignment, or simply, matching. Shape correspondence is a key algorithmic component in tasks such as 3D scan alignment and space-time reconstruction, as well as an indispensable prerequisite in diverse applications including attribute transfer, shape interpolation, and statistical modeling.",computer-vision268bec27b92-e0cd-462d-a186-7023850e20bc,box-supervised-instance-segmentation,Box-supervised Instance Segmentation,This task aims to achieve instance segmentation with weakly bounding box annotations.,computer-vision269c3dc83df-723e-4689-9505-969450bb0650,scene-change-detection,Scene Change Detection,"Scene change detection (SCD) refers to the task of localizing changes and identifying change-categories given two scenes. A scene can be either an RGB (+D) image or a 3D reconstruction (point cloud). If the scene is an image, SCD is a form of pixel-level prediction because each pixel in the image is classified according to a category. On the other hand, if the scene is point cloud, SCD is a form of point-level prediction because each point in the cloud is classified according to a category.
270
271Some example benchmarks for this task are VL-CMU-CD, PCD, and CD2014. Recently, more complicated benchmarks such as ChangeSim, HDMap, and Mallscape are released.
272
273Models are usually evaluated with the Mean Intersection-Over-Union (Mean IoU), Pixel Accuracy, or F1 metrics.",computer-vision274216f4430-39c9-4d66-ace1-7dce794e6a74,shape-from-texture,Shape from Texture,,computer-vision275394d8587-f5ce-4277-8768-24f864c323b7,object-slam,Object SLAM,SLAM (Simultaneous Localisation and Mapping) at the level of object,computer-vision27646969927-f59f-4bc0-b744-57c5d8ee6f4b,3d-canonical-hand-pose-estimation,3D Canonical Hand Pose Estimation,Image: [Lin et al](https://arxiv.org/pdf/2006.01320v1.pdf),computer-vision277dcf514ab-af7c-4fc3-adf6-8c3321bfa817,vnla,VNLA,Find objects in photorealistic environments by requesting and executing language subgoals.,computer-vision278e93ca207-f3ed-4500-aba3-732dc5f16298,multiple-action-detection,Multiple Action Detection,,computer-vision27943cb2468-492d-41b5-a6db-f40ac3e91bd4,pedestrian-attribute-recognition,Pedestrian Attribute Recognition,"Pedestrian attribution recognition is the task of recognizing pedestrian features - such as whether they are talking on a phone, whether they have a backpack, and so on.
280
281<span style=""color:grey; opacity: 0.6"">( Image credit: [HydraPlus-Net: Attentive Deep Features for Pedestrian Analysis](https://arxiv.org/pdf/1709.09930v1.pdf) )</span>",computer-vision2828b38205a-c5d3-4ba8-a319-a7da2ada5034,multi-human-parsing,Multi-Human Parsing,"Multi-human parsing is the task of parsing multiple humans in crowded scenes.
283
284<span style=""color:grey; opacity: 0.6"">( Image credit: [Multi-Human Parsing](https://github.com/ZhaoJ9014/Multi-Human-Parsing) )</span>",computer-vision28557311aa0-b8bb-4f60-b386-9e1bb968afdc,fine-grained-action-detection,Fine-Grained Action Detection,,computer-vision286b4dc1162-eb1c-41a5-9953-90b9095e3f8e,multi-label-image-retrieval,Multi-Label Image Retrieval,,computer-vision28725ed0d1b-2ef5-4655-bc49-15b5406c85c4,talking-head-generation,Talking Head Generation,"Talking head generation is the task of generating a talking face from a set of images of a person.
288
289<span style=""color:grey; opacity: 0.6"">( Image credit: [Few-Shot Adversarial Learning of Realistic Neural Talking Head Models](https://arxiv.org/pdf/1905.08233v2.pdf) )</span>",computer-vision290d718cff9-ea8e-4d66-9b82-6a4ad11fc8a6,overlapped-100-5,Overlapped 100-5,,computer-vision29163c2e631-ddb3-45b4-8ff5-fee100a5679f,text-guided-image-editing,text-guided-image-editing,Editing images using text prompts.,computer-vision292edce8046-b226-4d1a-8d0f-ea07edb930b6,car-pose-estimation,Car Pose Estimation,,computer-vision293799d0614-5ffb-4f9a-93d4-8febb9205427,aerial-video-saliency-prediction,Aerial Video Saliency Prediction,,computer-vision29494fb6801-595d-4589-aa0b-dfa63acd7b83,optical-character-recognition,Optical Character Recognition (OCR),"**Optical Character Recognition** or **Optical Character Reader** (OCR) is the electronic or mechanical conversion of images of typed, handwritten or printed text into machine-encoded text, whether from a scanned document, a photo of a document, a scene-photo (for example the text on signs and billboards in a landscape photo, license plates in cars...) or from subtitle text superimposed on an image (for example: from a television broadcast)",computer-vision295ade08720-9421-41a3-a61b-08c167eb8718,finger-vein-recognition,Finger Vein Recognition,,computer-vision296e1405af4-b588-47ef-b210-638a4eeecfb5,cross-domain-activity-recognition,Cross-Domain Activity Recognition,,computer-vision29738a38adc-3757-4804-9a45-87dd838d5c7e,sketch-based-image-retrieval,Sketch-Based Image Retrieval,,computer-vision298cd1d3e5b-8245-40da-9597-48918d6630cb,localization-in-video-forgery,Localization In Video Forgery,,computer-vision2993b4b65b4-77bf-49c0-a895-da75e039bb65,text-spotting,Text Spotting,"Text Spotting is the combination of Scene Text Detection and Scene Text Recognition in an end-to-end manner.
300It is the ability to read natural text in the wild.",computer-vision301fa33d39c-ef4d-4664-b99d-d00d50e2544a,temporal-metadata-manipulation-detection,Temporal Metadata Manipulation Detection,Detecting when the timestamp of an outdoor photograph has been manipulated,computer-vision30212db39be-a63b-4116-a7ef-d5a11cd1efef,video-kinematic-segmentation-base-workflow,"Video, Kinematic & Segmentation Base Workflow Recognition",,computer-vision30318140740-8671-4c59-a0f3-0a97ae2d7b88,intelligent-surveillance,Intelligent Surveillance,,computer-vision30472599c6e-cfa5-4377-9908-01a49837b3cc,occlusion-estimation,Occlusion Estimation,,computer-vision30521da2130-a239-4897-8812-b0db2f4dad2a,out-of-distribution-detection,Out-of-Distribution Detection,Detect out-of-distribution or anomalous examples.,computer-vision3067f1e78cb-63b4-4a98-83d8-e9b1eeb4f157,lung-nodule-3d-detection,Lung Nodule 3D Detection,,computer-vision307f1805e1a-b2fb-4120-a1a2-c774df0460e2,3d-semantic-instance-segmentation,3D Semantic Instance Segmentation,Image: [3D-SIS](https://github.com/Sekunde/3D-SIS),computer-vision30805fa3d5f-1672-46c9-9e40-df31d6dd8b77,data-ablation,Data Ablation,"Data Ablation is the study of change in data, and its effects in the performance of Neural Networks.",computer-vision3095abfc39b-2b83-46db-9d5c-2eb01777727f,stereo-depth-estimation,Stereo Depth Estimation,,computer-vision3100689431a-6c93-4a2d-b089-9ba047b63a5d,image-comprehension,Image Comprehension,,computer-vision311b922190a-9548-427a-ad58-cc159c616762,blood-cell-count,Blood Cell Count,,computer-vision3129dcc8bd7-5de0-4df2-8008-790ca9b1d8be,face-to-face-translation,Face to Face Translation,"Given a video of a person speaking in a source language, generate a video of the same person speaking in a target language.",computer-vision313e3041a77-1b9d-4e39-b1df-308b9af05b2e,whole-slide-images,whole slide images,,computer-vision314a5a5d818-1ca8-4fbe-b0c9-0d5c17f5697b,depth-image-upsampling,Depth Image Upsampling,,computer-vision3157fb4867c-5654-4b6a-a044-a6e7b12d93d5,video-classification,Video Classification,"**Video Classification** is the task of producing a label that is relevant to the video given its frames. A good video level classifier is one that not only provides accurate frame labels, but also best describes the entire video given the features and the annotations of the various frames in the video. For example, a video might contain a tree in some frame, but the label that is central to the video might be something else (e.g., “hiking”). The granularity of the labels that are needed to describe the frames and the video depends on the task. Typical tasks include assigning one or more global labels to the video, and assigning one or more labels for each frame inside the video.
316
317
318<span class=""description-source"">Source: [Efficient Large Scale Video Classification ](https://arxiv.org/abs/1505.06250)</span>",computer-vision319a3f83dec-1202-483e-86c4-03d64d34e64d,pose-prediction,Pose Prediction,Pose prediction is to predict future poses given a window of previous poses.,computer-vision320b3a58287-bffe-42c6-adb9-b2b1ab253f7e,body-detection,Body Detection,Detection of the persons or the characters defined in the dataset.,computer-vision32180cda782-ce65-4995-a695-52e3b141272a,medical-image-retrieval,Medical Image Retrieval,,computer-vision322abccca0c-7b88-4453-bb56-5a85c2bfb98e,reconstruction,Reconstruction,,computer-vision3232c99156f-e5d9-49c3-a52f-14b985930d16,soil-moisture-estimation,Soil moisture estimation,,computer-vision3243fa5f31d-23e6-4dac-ba54-6c25b94319b7,calving-front-delineation-in-synthetic,Calving Front Delineation In Synthetic Aperture Radar Imagery,"Delineating the calving front of a marine-terminating glacier in synthetic aperture radar (SAR) imagery. This can, for example, be done through Semantic Segmentation.",computer-vision3259f8957c2-332c-41d6-a11a-fb406b48d016,overlapping-pose-estimation,Overlapping Pose Estimation,Pose estimation with overlapping poses.,computer-vision32623841ae0-647a-40dd-ba67-c0bfe55ed6cd,video-anomaly-detection,Video Anomaly Detection,,computer-vision32751dc6e90-ade2-4a35-aed6-38c8f94ada97,image-quality-estimation,Image Quality Estimation,,computer-vision328b7e61bfc-f316-41e8-a655-520995c582fb,hand-gesture-recognition,Hand Gesture Recognition,,computer-vision3298235a86d-e484-4bf6-b245-315191355afd,overlapped-5-3,Overlapped 5-3,,computer-vision3309855a142-dbd6-46e2-b31e-3586efd0a78b,interactive-video-object-segmentation,Interactive Video Object Segmentation,"The interactive scenario assumes the user gives iterative refinement inputs to the algorithm, in our case in the form of a scribble, to segment the objects of interest. Methods have to produce a segmentation mask for that object in all the frames of a video sequence taking into account all the user interactions.",computer-vision3311e18c425-8400-47c8-a0fb-f92a328573a1,space-time-video-super-resolution,Space-time Video Super-resolution,,computer-vision3321d78dd7f-d924-4380-b35b-c8ebd0764c06,image-manipulation-detection,Image Manipulation Detection,"The task of detecting images or image parts that have been tampered or manipulated (sometimes also referred to as doctored). This typically encompasses image splicing, copy-move, or image inpainting.",computer-vision333d593bc67-825b-4e3e-a455-87aa36a7fafa,few-shot-action-recognition,Few Shot Action Recognition,"Few-shot (FS) action recognition is a challenging com-
334puter vision problem, where the task is to classify an unlabelled query video into one of the action categories in the support set having limited samples per action class.",computer-vision33566b7d07b-d656-4a37-b68e-86359fb32791,camera-absolute-pose-regression,camera absolute pose regression,,computer-vision33648c623af-7b5c-4cd2-b6b7-aaef41bcda41,procedure-learning,Procedure Learning,"Given a set of videos of the same task, the goal is to identify the key-steps required to perform the task.",computer-vision337f4b5487f-955b-48b4-b832-d370826627f8,coos-7-accuracy,COOS-7 Accuracy,"COOS-7 contains 132,209 single-cell images of mouse cells, where the task is to predict protein subcellular localization. Images are spread over 1 training set and 4 testing sets, where each single-cell image contains a protein and nucleus fluorescent channels. COOS-7 provides a classification setting where four test datasets have increasing degrees of covariate shift: some images are random subsets of the training data, while others are from experiments reproduced months later and imaged by different instruments. While most classifiers perform well on test datasets similar to the training dataset, all classifiers failed to generalize their performance to datasets with greater covariate shifts. Read more at https://www.alexluresearch.com/publication/coos/.",computer-vision338cc478b80-09db-429c-939d-5ff3b8caf3e4,geometric-matching,Geometric Matching,,computer-vision339a7a2c2ae-1c8f-4519-8416-cb6aae068ae3,incomplete-multi-view-clustering,Incomplete multi-view clustering,,computer-vision340c0bbf2cc-fb41-4425-9ca9-1c4ea8051457,semi-supervised-human-pose-estimation,Semi-Supervised Human Pose Estimation,Semi-supervised human pose estimation aims to leverage the unlabelled data along with labeled data to improve the model performance.,computer-vision341bf9b27b6-74bb-4bf6-93cd-63213f37ca2c,training-free-3d-point-cloud-classification,Training-free 3D Point Cloud Classification,Evaluation on target datasets for 3D Point Cloud Classification without any training,computer-vision342a570d232-1b9a-46c2-994f-8d9d427f11ce,multiple-people-tracking,Multiple People Tracking,,computer-vision3439a20a044-61b5-439b-8698-0ac869dd889d,monocular-3d-object-detection,Monocular 3D Object Detection,Monocular 3D Object Detection is the task to draw 3D bounding box around objects in a single 2D RGB image. It is localization task but without any extra information like depth or other sensors or multiple-images.,computer-vision3442cb07cda-08ba-40d4-80ef-a7622b7043a9,sample-probing,Sample Probing,,computer-vision345ba15895c-0e50-4474-bca8-9b0fe814bd12,cloud-removal,Cloud Removal,"The majority of all optical observations collected via spaceborne satellites are affected by haze or clouds. Consequently, persistent cloud coverage affects the remote sensing practitioner's capabilities of a continuous and seamless monitoring of our planet. **Cloud removal** is the task of reconstructing cloud-covered information while preserving originally cloud-free details.
346
347Image Source: [URL](https://patrickTUM.github.io/cloud_removal/)",computer-vision348e0260aec-3eb0-4f16-8e62-9e2f951beb84,speaker-specific-lip-to-speech-synthesis,Speaker-Specific Lip to Speech Synthesis,"How accurately can we infer an individual’s speech style and content from his/her lip movements? [1]
349
350In this task, the model is trained on a specific speaker, or a very limited set of speakers.
351
352[1] Learning Individual Speaking Styles for Accurate Lip to Speech Synthesis, CVPR 2020.",computer-vision3539d80a0a3-02bf-405b-ae61-fbdfa7685f57,bokeh-effect-rendering,Bokeh Effect Rendering,,computer-vision354c342fcfd-d872-4b64-a6ee-b3bf3b41dd3c,video,Video,,computer-vision355f992ca31-07ee-4dd3-9941-fb8ef61a6dbe,evolving-domain-generalization,Evolving Domain Generalization,,computer-vision3565bdfbed8-ee91-4893-b1ba-0b8d9c2be409,infrared-image-super-resolution,Infrared image super-resolution,Aims at upsampling the IR image and create the high resolution image with help of a low resolution image.,computer-vision35717856ed4-9415-4ae0-ba36-b9b392b8c7b4,referring-expression-generation,Referring expression generation,Generate referring expressions,computer-vision3582ef55c58-dcdc-4d40-9db9-e5528a8d8827,image-super-resolution,Image Super-Resolution,"**Image Super-Resolution** is a machine learning task where the goal is to increase the resolution of an image, often by a factor of 4x or more, while maintaining its content and details as much as possible. The end result is a high-resolution version of the original image. This task can be used for various applications such as improving image quality, enhancing visual detail, and increasing the accuracy of computer vision algorithms.",computer-vision359222235c9-08e8-407f-8394-a9abd16d98de,video-deinterlacing,Video Deinterlacing,,computer-vision3606b786155-01ba-45d4-96ab-31f77673b222,unsupervised-semantic-segmentation-with,Unsupervised Semantic Segmentation with Language-image Pre-training,A segmentation task which does not utilise any human-level supervision for semantic segmentation except for a backbone which is initialised with features pre-trained with image-level labels.,computer-vision36154532c49-e0de-4067-9914-1f7fb4e0bc82,point-cloud-classification-dataset,Point cloud classification dataset,,computer-vision3626ba03727-0f83-4559-8812-0feefacbac76,head-detection,Head Detection,,computer-vision36324967556-a98f-4bb4-94bb-3887175a3315,classifier-calibration,Classifier calibration,Confidence calibration – the problem of predicting probability estimates representative of the true correctness likelihood – is important for classification models in many applications. The two common calibration metrics are Expected Calibration Error (ECE) and Maximum Calibration Error (MCE).,computer-vision36460932ef7-e075-471b-aff0-49b5d8d53473,pso-convnets-dynamics-1,PSO-ConvNets Dynamics 1,Incorporating distilled Cucker-Smale elements into PSO algorithm using KNN and intertwine training with SGD,computer-vision3658266034c-658d-4856-8032-03036ff92a1b,trajectory-prediction,Trajectory Prediction,"**Trajectory Prediction** is the problem of predicting the short-term (1-3 seconds) and long-term (3-5 seconds) spatial coordinates of various road-agents such as cars, buses, pedestrians, rickshaws, and animals, etc. These road-agents have different dynamic behaviors that may correspond to aggressive or conservative driving styles.
366
367
368<span class=""description-source"">Source: [Forecasting Trajectory and Behavior of Road-Agents Using Spectral Clustering in Graph-LSTMs ](https://arxiv.org/abs/1912.01118)</span>",computer-vision369703ffa05-82d0-41ec-917b-aa206eac02a8,scene-labeling,Scene Labeling,,computer-vision3701f047631-e1ce-487d-914b-83daaf78811d,transparent-objects,Transparent objects,,computer-vision3719e4441b4-89de-4375-a58d-0fedfd312e8f,continuous-object-recognition,Continuous Object Recognition,"Continuous object recognition is the task of performing object recognition on a data stream and learning continuously, trying to mitigate issues such as catastrophic forgetting.
372
373<span style=""color:grey; opacity: 0.6"">( Image credit: [CORe50 dataset](https://vlomonaco.github.io/core50/) )</span>",computer-vision3747da35209-aa26-4a5e-8110-ad6438905257,3d-inpainting,3D Inpainting,"**3D Inpainting** is the removal of unwanted objects
375from a 3D scene, such that the replaced region is visually
376plausible and consistent with its context.",computer-vision37715e74df3-be84-4fee-a091-7334b22f0786,image-to-gps-verification,Image-To-Gps Verification,"The image-to-GPS verification task asks whether a given image is taken at a claimed GPS location.
378
379<span style=""color:grey; opacity: 0.6"">( Image credit: [Image-to-GPS Verification Through A Bottom-Up Pattern Matching Network](https://arxiv.org/pdf/1811.07288v1.pdf) )</span>",computer-vision380cd34d626-1b22-4929-ba43-b76815733792,hand-detection,Hand Detection,"As an important subject in the field of computer vision, hand detection plays an important role in many tasks such as human-computer interaction, automatic driving, virtual reality and so on.",computer-vision381b15957f8-1c8c-42ce-a65a-2450d28fa8ca,vision-language-navigation,Vision-Language Navigation,"Vision-language navigation (VLN) is the task of navigating an embodied agent to carry out natural language instructions inside real 3D environments.
382
383<span style=""color:grey; opacity: 0.6"">( Image credit: [Learning to Navigate Unseen Environments:
384Back Translation with Environmental Dropout](https://arxiv.org/pdf/1904.04195v1.pdf) )</span>",computer-vision38544150d39-d078-4511-b94d-49bc08ffd190,self-supervised-image-classification,Self-Supervised Image Classification,"This is the task of image classification using representations learnt with self-supervised learning. Self-supervised methods generally involve a pretext task that is solved to learn a good representation and a loss function to learn with. One example of a loss function is an autoencoder based loss where the goal is reconstruction of an image pixel-by-pixel. A more popular recent example is a contrastive loss, which measure the similarity of sample pairs in a representation space, and where there can be a varying target instead of a fixed target to reconstruct (as in the case of autoencoders).
386
387A common evaluation protocol is to train a linear classifier on top of (frozen) representations learnt by self-supervised methods. The leaderboards for the linear evaluation protocol can be found below. In practice, it is more common to fine-tune features on a downstream task. An alternative evaluation protocol therefore uses semi-supervised learning and finetunes on a % of the labels. The leaderboards for the finetuning protocol can be accessed [here](https://paperswithcode.com/task/semi-supervised-image-classification).
388
389You may want to read some blog posts before reading the papers and checking the leaderboards:
390
391- [Contrastive Self-Supervised Learning](https://ankeshanand.com/blog/2020/01/26/contrative-self-supervised-learning.html) - Ankesh Anand
392- [The Illustrated Self-Supervised Learning](https://amitness.com/2020/02/illustrated-self-supervised-learning/) - Amit Chaudhary
393- [Self-supervised learning and computer vision](https://www.fast.ai/2020/01/13/self_supervised/) - Jeremy Howard
394- [Self-Supervised Representation Learning](https://lilianweng.github.io/lil-log/2019/11/10/self-supervised-learning.html) - Lilian Weng
395
396There is also Yann LeCun's talk at AAAI-20 which you can watch [here](https://vimeo.com/390347111) (35:00+).
397
398<span style=""color:grey; opacity: 0.6"">( Image credit: [A Simple Framework for Contrastive Learning of Visual Representations](https://arxiv.org/pdf/2002.05709v1.pdf) )</span>",computer-vision399f6b579a6-1084-45f3-89e0-77703c7e83b5,webpage-object-detection,Webpage Object Detection,Detect Web Element for various classes from candidate web elements obtained from DOM tree (No need for Bounding Box Regression),computer-vision4008ac147eb-cfb9-44a5-8451-0ddc57c30616,contrastive-learning,Contrastive Learning,"**Contrastive Learning** is a deep learning technique for unsupervised representation learning. The goal is to learn a representation of data such that similar instances are close together in the representation space, while dissimilar instances are far apart.
401
402It has been shown to be effective in various computer vision and natural language processing tasks, including image retrieval, zero-shot learning, and cross-modal retrieval. In these tasks, the learned representations can be used as features for downstream tasks such as classification and clustering.
403
404<span style=""color:grey; opacity: 0.6"">(Image credit: [Schroff et al. 2015](https://arxiv.org/abs/1503.03832))</span>",computer-vision4057a2c0795-026d-4cbd-8022-56a30e2921b8,vehicle-key-point-and-orientation-estimation,Vehicle Key-Point and Orientation Estimation,,computer-vision4065f7b96ee-e452-4602-8a59-3fc6ac13c9e8,action-triplet-detection,Action Triplet Detection,"Detecting and localizing bounding boxes of tools and anatomies. Then prediction their relationship as action triplet <instrument, verb, target>",computer-vision407e7a93535-cfde-47e7-8a9c-769071d6715e,self-driving-cars,Self-Driving Cars,"Self-driving cars : the task of making a car that can drive itself without human guidance.
408
409<span style=""color:grey; opacity: 0.6"">( Image credit: [Learning a Driving Simulator](https://github.com/commaai/research) )</span>",computer-vision410eeaad7a2-7616-4ac0-b7fd-c9be0574516f,gesture-to-gesture-translation,Gesture-to-Gesture Translation,,computer-vision4119f03f48f-c645-4423-a800-4c96d28ea951,3d-hand-pose-estimation,3D Hand Pose Estimation,Image: [Zimmerman et l](https://arxiv.xsrg/pdf/1705.01389v3.pdf),computer-vision41209dd3238-9b34-41b0-ad46-cce3b4805874,texture-synthesis,Texture Synthesis,"The fundamental goal of example-based **Texture Synthesis** is to generate a texture, usually larger than the input, that faithfully captures all the visual characteristics of the exemplar, yet is neither identical to it, nor exhibits obvious unnatural looking artifacts.413 414 415<span class=""description-source"">Source: [Non-Stationary Texture Synthesis by Adversarial Expansion ](https://arxiv.org/abs/1805.04487)</span>",computer-vision4161c158e2e-bdd7-4420-9d63-82a81ff050ec,dehazing,Dehazing,,computer-vision417e3392a2c-667f-4150-a2cb-445eac963ed5,3d-point-cloud-reinforcement-learning,3D Point Cloud Reinforcement Learning,Reinforcement learning / robot learning from 3D point clouds,computer-vision418499b9dac-abf4-45df-9b46-7749e7be4cc5,action-spotting,Action Spotting,,computer-vision41944dce5d7-5783-4ec7-8724-1a62ffeb58c0,image-editing,Image Editing,,computer-vision4206046d910-07a6-419c-a4f6-ee34b635cb25,thoracic-disease-classification,Thoracic Disease Classification,,computer-vision421960b7cfb-0f71-4b49-982e-77aca1b228f0,point-cloud-completion,Point Cloud Completion,,computer-vision4228641b703-0514-42d1-bce6-8f75de09c0b9,depiction-invariant-object-recognition,Depiction Invariant Object Recognition,"Depiction invariant object recognition is the task of recognising objects irrespective of how they are visually depicted (line drawing, realistic shaded drawing, photograph etc.).
423
424<span style=""color:grey; opacity: 0.6"">( Image credit: [SwiDeN](https://arxiv.org/pdf/1607.08764v1.pdf) )</span>",computer-vision42597e34d22-ed15-44eb-ad66-7212eaa29674,iris-recognition,Iris Recognition,,computer-vision42603055b5d-357f-4a10-973b-f2bbe5b37314,pedestrian-detection,Pedestrian Detection,"Pedestrian detection is the task of detecting pedestrians from a camera.
427
428Further state-of-the-art results (e.g. on the KITTI dataset) can be found at [3D Object Detection](https://paperswithcode.com/task/object-detection).
429
430<span style=""color:grey; opacity: 0.6"">( Image credit: [High-level Semantic Feature Detection: A New Perspective for Pedestrian Detection](https://github.com/liuwei16/CSP) )</span>",computer-vision431c96f28e3-0c7d-44ba-9998-b99317274c22,co-saliency-detection,Co-Salient Object Detection,"**Co-Salient Object Detection** is a computational problem that aims at highlighting the common and salient foreground regions (or objects) in an image group. Please also refer to the online benchmark: http://dpfan.net/cosod3k/
432
433
434
435
436<span style=""color:grey; opacity: 0.6"">( Image credit: [Taking a Deeper Look at Co-Salient Object Detection, CVPR2020](https://openaccess.thecvf.com/content_CVPR_2020/papers/Fan_Taking_a_Deeper_Look_at_Co-Salient_Object_Detection_CVPR_2020_paper.pdf) )</span>",computer-vision437f6d7576e-7e15-4064-9ed0-3a6a0097dd61,deep-attention,Deep Attention,,computer-vision43873c6438e-a779-4eeb-b17e-2dcb8e8d9941,steganalysis,Steganalysis,Detect the usage of Steganography,computer-vision4392508f20c-b3e9-4680-a263-1fdd2e8d6b53,motion-detection-in-non-stationary-scenes,Motion Detection In Non-Stationary Scenes,,computer-vision440dd1b1b27-7257-4eb3-988c-896413c053d1,document-image-classification,Document Image Classification,"Document image classification is the task of classifying documents based on images of their contents.
441
442<span style=""color:grey; opacity: 0.6"">( Image credit: [Real-Time Document Image Classification using Deep CNN and Extreme Learning Machines](https://arxiv.org/pdf/1711.05862v1.pdf) )</span>",computer-vision4433fa98e1f-f9f6-45fd-a754-0844f08bffe8,image-imputation,Image Imputation,"Image imputation is the task of creating plausible images from low-resolution images or images with missing data.
444
445<span style=""color:grey; opacity: 0.6"">( Image credit: [NASA](https://www.jpl.nasa.gov/edu/news/2019/4/19/how-scientists-captured-the-first-image-of-a-black-hole/) )</span>",computer-vision446558be24c-25af-4630-86a1-8caa57fdbe19,safety-perception-recognition,Safety Perception Recognition,City safety perception recognition,computer-vision447e7864627-809c-41f7-9551-63dba7b6a8ca,3d-object-classification,3D Object Classification,"3D Object Classification is the task of predicting the class of a 3D object point cloud. It is a voxel level prediction where each voxel is classified into a category. The popular benchmark for this task is the ModelNet dataset. The models for this task are usually evaluated with the Classification Accuracy metric.
448
449Image: [Sedaghat et al](https://arxiv.org/pdf/1604.03351v2.pdf)",computer-vision450b56206a6-8243-4c90-8408-7149ccaef62e,interspecies-facial-keypoint-transfer,Interspecies Facial Keypoint Transfer,Find cross-domain semantic correspondence between faces of different species,computer-vision4517652506d-6ca9-4769-adb5-475675bc00da,classify-3d-point-clouds,Classify 3D Point Clouds,,computer-vision45239ca221a-f35f-4312-9045-1f6f39ca1816,image-fusion,Image Fusion,,computer-vision4533f3d93e2-e9a7-4751-87fd-b7c93345d151,junction-detection,Junction Detection,,computer-vision45440283550-eba0-4132-9396-573f4cd8c637,aerial-video-semantic-segmentation,Aerial Video Semantic Segmentation,,computer-vision455e5add8c5-8a34-4d68-baf0-a18ad241732f,panoptic-segmentation,Panoptic Segmentation,"**Panoptic Segmentation** is a computer vision task that combines semantic segmentation and instance segmentation to provide a comprehensive understanding of the scene. The goal of panoptic segmentation is to segment the image into semantically meaningful parts or regions, while also detecting and distinguishing individual instances of objects within those regions.
456
457<span style=""color:grey; opacity: 0.6"">( Image credit: [Detectron2](https://github.com/facebookresearch/detectron2) )</span>",computer-vision458ceb4fce9-e64f-43b3-8c32-dc32d6b4d591,seeing-beyond-the-visible,Seeing Beyond the Visible,"The objective of this challenge is to automate the process of estimating the soil parameters, specifically, potassium (KKK), phosphorus pentoxide (P2O5P_2O_5P2O5), magnesium (MgMgMg) and pHpHpH, through extracting them from the airborne hyperspectral images captured over agricultural areas in Poland (the exact locations are not revealed). To make the solution applicable in real-life use cases, all the parameters should be estimated as precisely as possible.",computer-vision4596966e2f2-af07-47e3-a44b-b12011c93367,defocus-blur-detection,Defocus Blur Detection,,computer-vision460c786c34b-ec6b-4989-9cf5-86105753a3c0,3d-surface-generation,3D Surface Generation,Image: [AtlasNet](https://arxiv.org/pdf/1802.05384v3.pdf),computer-vision461b5470826-fe4d-43f9-9f64-265ae7eef21a,3d-room-layouts-from-a-single-rgb-panorama,3D Room Layouts From A Single RGB Panorama,Image: [Zou et al](https://arxiv.org/pdf/1803.08999v1.pdf),computer-vision462b1dc544c-0786-409a-8cfc-77c5548f0c5c,pulmonary-arteryvein-classification,Pulmonary Artery–Vein Classification,,computer-vision4636b5cb2a0-edfc-44eb-a4bd-d0add6980790,heterogeneous-face-recognition,Heterogeneous Face Recognition,"Heterogeneous face recognition is the task of matching face images acquired from different sources (i.e., different sensors or different wavelengths) for identification or verification.
464
465<span style=""color:grey; opacity: 0.6"">( Image credit: [Pose Agnostic Cross-spectral Hallucination via Disentangling Independent Factors](https://arxiv.org/pdf/1909.04365v1.pdf) )</span>",computer-vision46612cf4514-7616-43f1-95f3-cc3bfa127a24,mobile-periocular-recognition,Mobile Periocular Recognition,"Periocular recognition is the task of recognising a person based on their eyes (periocular).
467
468<span style=""color:grey; opacity: 0.6"">( Image credit: [Heterogeneity Aware Deep Embedding for Mobile Periocular Recognition](https://arxiv.org/pdf/1811.00846v1.pdf) )</span>",computer-vision4695c59921c-b7ae-4d73-ba3e-58a9c7fbf7a7,class-incremental-learning,Class Incremental Learning,Incremental learning of a sequence of tasks when the task-ID is not available at test time.,computer-vision470d0750c62-498b-4beb-9257-5e0d40df9209,event-data-classification,Event data classification,,computer-vision471c0689bfe-5ffc-4837-9828-5e3706e1b9b5,zero-shot-segmentation,Zero Shot Segmentation,,computer-vision4725229ed4d-6203-4266-aca8-6df0e8c59ed8,animated-gif-generation,Animated GIF Generation,,computer-vision473545dda68-45d0-43e6-b31f-f0e97dc045b5,single-shot-hdr-reconstruction,Single-shot HDR Reconstruction,"SVE-based HDR imaging, also known as single-shot HDR imaging, algorithms capture a scene with pixel-wise varying exposures in a single image and then computationally synthesize an HDR image, which benefits from the multiple exposures of the single image.",computer-vision474ffbaaf3c-3ae6-486e-b333-9d755763f531,zero-shot-transfer-image-classification,Zero-Shot Transfer Image Classification,,computer-vision475f0ec4f70-9844-469d-8fe8-294584090c15,self-knowledge-distillation,Self-Knowledge Distillation,,computer-vision4767cb187be-8404-407d-acd9-4e80337a9bd1,colorization,Colorization,"**Colorization** is the process of adding plausible color information to monochrome photographs or videos. Colorization is a highly undetermined problem, requiring mapping a real-valued luminance image to a three-dimensional color-valued one, that has not a unique solution.
477
478
479<span class=""description-source"">Source: [ChromaGAN: An Adversarial Approach for Picture Colorization ](https://arxiv.org/abs/1907.09837)</span>",computer-vision480c39b7095-04ba-4414-a1c1-4445a437aa5d,stereo-matching,Stereo Matching Hand,,computer-vision481a3165a70-eeb5-4932-b0d8-38a28d48d0c8,sketch,Sketch,,computer-vision4824d7f83ea-88ad-4470-82c3-b42f9923b8c6,point-cloud-reconstruction,Point cloud reconstruction,"This task aims to solve inherent problems in raw point clouds: sparsity, noise, and irregularity.",computer-vision483425adff8-97cf-4db6-8a8b-fa5d419a2bd4,sports-analytics,Sports Analytics,,computer-vision484161ec6cc-0f5d-4ea2-8916-cf7e625e95cc,lung-nodule-3d-classification,Lung Nodule 3D Classification,,computer-vision485b47da698-6ddc-409c-81e5-95c17755c7d6,underwater-image-restoration,Underwater Image Restoration,Underwater image restoration aims to rectify the distorted colors and present the true colors of the underwater scene.,computer-vision4869842db39-3e69-4dab-8351-d15cbabfc1c5,partially-view-aligned-multi-view-learning,Partially View-aligned Multi-view Learning,"In multi-view learning, Partially View-aligned Problem (PVP) refers to the case when only a portion of data is aligned, thus leading to data inconsistency.",computer-vision48737adb025-133f-4dc3-b783-5c300fc918b6,scene-text-recognition,Scene Text Recognition,See [Scene Text Detection](https://paperswithcode.com/task/scene-text-detection) for leaderboards in this task.,computer-vision488e774a091-7dec-4609-a405-75993420ec3b,face-image-quality,Face Image Quality,,computer-vision489aa7a113b-1d5b-43f5-bfd2-24439adcb386,demosaicking,Demosaicking,"Most modern digital cameras acquire color images by measuring only one color channel per pixel, red, green, or blue, according to a specific pattern called the Bayer pattern. **Demosaicking** is the processing step that reconstruct a full color image given these incomplete measurements.
490
491
492<span class=""description-source"">Source: [Revisiting Non Local Sparse Models for Image Restoration ](https://arxiv.org/abs/1912.02456)</span>",computer-vision4933a645d1b-7a5e-4ac1-aca2-97cf2f671728,handwritten-word-generation,Handwritten Word Generation,,computer-vision494a772458d-b740-43fd-98c6-fdf8239a9ff1,physical-attribute-prediction,Physical Attribute Prediction,,computer-vision495e4199210-f7a5-418e-a7c9-44eef6787ef9,face-recognition,Face Recognition,"**Facial Recognition** is the task of making a positive identification of a face in a photo or video image against a pre-existing database of faces. It begins with detection - distinguishing human faces from other objects in the image - and then works on identification of those detected faces.
496
497The state of the art tables for this task are contained mainly in the consistent parts of the task : the face verification and face identification tasks.
498
499<span style=""color:grey; opacity: 0.6"">( Image credit: [Face Verification](https://shuftipro.com/face-verification) )</span>",computer-vision50073076f7f-3696-4c1b-b6de-e052a6122dcd,2d-object-detection,2D Object Detection,,computer-vision501ffa05148-d8d9-412a-8dc0-6d9ce0dae722,hand-gesture-recognition-1,Hand-Gesture Recognition,,computer-vision50262519edd-dd37-4fa4-9bbd-c77df86b7cef,saliency-prediction,Saliency Prediction,A saliency map is a model that predicts eye fixations on a visual scene.,computer-vision503598266d6-9ad8-4c1b-80d0-889d181543d8,video-individual-counting,Video Individual Counting,,computer-vision5042c84af61-c3ed-465d-a56b-5a687b1865be,road-scene-understanding,road scene understanding,,computer-vision505decdb3d9-e999-4a4b-aa9f-5e8879d978ab,pose-tracking,Pose Tracking,"**Pose Tracking** is the task of estimating multi-person human poses in videos and assigning unique instance IDs for each keypoint across frames. Accurate estimation of human keypoint-trajectories is useful for human action recognition, human interaction understanding, motion capture and animation.
506
507
508<span class=""description-source"">Source: [LightTrack: A Generic Framework for Online Top-Down Human Pose Tracking ](https://arxiv.org/abs/1905.02822)</span>",computer-vision509efdfa4d9-2923-46e0-8601-d15811fb1c93,physical-video-anomaly-detection,Physical Video Anomaly Detection,Detecting if an entire short clip of a physical or mechanical process features an anomalous motion,computer-vision5100ff35f8d-1c2a-44ea-8f5f-bfff6349fa9c,short-term-object-interaction-anticipation,Short-term Object Interaction Anticipation,,computer-vision5114f84d5f5-df2b-4cf7-90d6-2a872ca9f035,human-object-interaction-motion-tracking,Human-Object-interaction motion tracking,,computer-vision512e95b2a2d-60f2-4dd7-9312-de3c1d9b3d4d,sketch-recognition,Sketch Recognition,,computer-vision5135f514594-7e9d-4485-831f-5c6ba687cc2c,image-smoothing,image smoothing,,computer-vision5149a0551f8-2f53-4f80-bcc0-270f7f2ff2e2,image-dehazing,Image Dehazing,"<span style=""color:grey; opacity: 0.6"">( Image credit: [Densely Connected Pyramid Dehazing Network](https://github.com/hezhangsprinter/DCPDN) )</span>",computer-vision515ac57b90e-a6a7-435f-b9b7-cdb5b0109620,action-quality-assessment,Action Quality Assessment,Assessing/analyzing/quantifying how well an action was performed.,computer-vision5165da5be03-bbff-41cf-ad3c-8de291b38bb6,multi-oriented-scene-text-detection,Multi-Oriented Scene Text Detection,,computer-vision517e0df1e19-6873-4df5-a154-94791463fe90,hand,Hand,,computer-vision51819b3ae32-7764-42b8-8beb-69da638c7bf2,referring-image-matting-keyword-based,Referring Image Matting (Keyword-based),"Keyword-based referring image matting, taking an image and a keyword word as the input.",computer-vision5190121e17f-8041-4c49-bb4a-a8e453694d99,single-object-discovery,Single-object discovery,,computer-vision52092e29716-2c1e-46a8-874a-28385251812f,deblurring,Deblurring,"**Deblurring** is a computer vision task that involves removing the blurring artifacts from images or videos to restore the original, sharp content. Blurring can be caused by various factors such as camera shake, fast motion, and out-of-focus objects, and can result in a loss of detail and quality in the captured images. The goal of deblurring is to produce a clear, high-quality image that accurately represents the original scene.
521
522<span style=""color:grey; opacity: 0.6"">( Image credit: [Deblurring Face Images using Uncertainty Guided Multi-Stream Semantic Networks](https://arxiv.org/pdf/1907.13106v1.pdf) )</span>",computer-vision5236dfbd461-c177-4d48-bbe5-0faa7ea2ba87,unsupervised-long-term-person-re,Unsupervised Long Term Person Re-Identification,"Long-term Person Re-Identification(Clothes-Changing Person Re-ID) is a computer vision task in which the goal is to match a person's identity across different cameras, clothes, and locations in a video or image sequence. It involves detecting and tracking a person and then using features such as appearance, and body shape to match their identity in different frames. The goal is to associate the same person across multiple non-overlapping camera views in a robust and efficient manner.",computer-vision524d3ccf96f-264a-415d-b2d0-b3508eac2f7c,language-based-temporal-localization,Language-Based Temporal Localization,,computer-vision52507477428-009f-4e13-a02d-b6e054b8a053,skeleton-based-action-recognition,Skeleton Based Action Recognition,"**Skeleton-based Action Recognition** is a computer vision task that involves recognizing human actions from a sequence of 3D skeletal joint data captured from sensors such as Microsoft Kinect, Intel RealSense, and wearable devices. The goal of skeleton-based action recognition is to develop algorithms that can understand and classify human actions from skeleton data, which can be used in various applications such as human-computer interaction, sports analysis, and surveillance.
526
527<span style=""color:grey; opacity: 0.6"">( Image credit: [View Adaptive Neural Networks for High
528Performance Skeleton-based Human Action
529Recognition](https://arxiv.org/pdf/1804.07453v3.pdf) )</span>",computer-vision5301cd36ca0-926c-4c58-9908-ba28af6466a5,pose-estimation,Pose Estimation,"**Pose Estimation** is a computer vision task where the goal is to detect the position and orientation of a person or an object. Usually, this is done by predicting the location of specific keypoints like hands, head, elbows, etc. in case of Human Pose Estimation.
531
532A common benchmark for this task is [MPII Human Pose](https://paperswithcode.com/sota/pose-estimation-on-mpii-human-pose)
533
534<span style=""color:grey; opacity: 0.6"">( Image credit: [Real-time 2D Multi-Person Pose Estimation on CPU: Lightweight OpenPose](https://github.com/Daniil-Osokin/lightweight-human-pose-estimation.pytorch) )</span>",computer-vision5357d084b81-6bdb-481f-84b6-40c26318b2b6,overlapped-10-1,Overlapped 10-1,,computer-vision5362cbd4db6-d55e-4468-94b1-8faaf16817f4,facial-attribute-classification,Facial Attribute Classification,"Facial attribute classification is the task of classifying various attributes of a facial image - e.g. whether someone has a beard, is wearing a hat, and so on.
537
538<span style=""color:grey; opacity: 0.6"">( Image credit: [Multi-task Learning of Cascaded CNN for Facial Attribute Classification
539](https://arxiv.org/pdf/1805.01290v1.pdf) )</span>",computer-vision540f5ddeb93-3361-4cf9-b635-02d0663bfd24,material-classification,Material Classification,,computer-vision54141397ff1-b9a9-4ff3-bee8-b1fe4bffbdab,natural-image-orientation-angle-detection,Natural Image Orientation Angle Detection,"Image orientation angle detection is a pretty challenging task for a machine because the machine has to learn the features of an image in such a way so that it can detect the arbitrary angle by which the image is rotated. Though there are some modern cameras with features involving inertial sensors that can correct image orientation in steps of 90 degrees, those features are seldom used. In this paper, we propose a method to detect the orientation angle of a digitally captured image where the image may have been captured by a camera at a tilted angle (between 0\degree to 359\degree).",computer-vision542f9e9de16-f37e-4ad9-8a20-594a41a3d417,motion-detection,Motion Detection,"**Motion Detection** is a process to detect the presence of any moving entity in an area of interest. Motion Detection is of great importance due to its application in various areas such as surveillance and security, smart homes, and health monitoring.543 544 545<span class=""description-source"">Source: [Different Approaches for Human Activity Recognition– A Survey ](https://arxiv.org/abs/1906.05074)</span>",computer-vision54616b9901d-5e3a-42ae-8f8b-8e401955b48c,image-matting,Image Matting,"**Image Matting** is the process of accurately estimating the foreground object in images and videos. It is a very important technique in image and video editing applications, particularly in film production for creating visual effects. In case of image segmentation, we segment the image into foreground and background by labeling the pixels. Image segmentation generates a binary image, in which a pixel either belongs to foreground or background. However, Image Matting is different from the image segmentation, wherein some pixels may belong to foreground as well as background, such pixels are called partial or mixed pixels. In order to fully separate the foreground from the background in an image, accurate estimation of the alpha values for partial or mixed pixels is necessary.
547
548
549<span class=""description-source"">Source: [Automatic Trimap Generation for Image Matting ](https://arxiv.org/abs/1707.00333)</span>
550
551<span class=""description-source"">Image Source: [Real-Time High-Resolution Background Matting](https://arxiv.org/pdf/2012.07810v1.pdf)</span>",computer-vision552b3101a43-dc62-489b-bcf7-082e13dcc193,semi-supervised-video-object-segmentation,Semi-Supervised Video Object Segmentation,The semi-supervised scenario assumes the user inputs a full mask of the object(s) of interest in the first frame of a video sequence. Methods have to produce the segmentation mask for that object(s) in the subsequent frames.,computer-vision553b6d2e61d-7800-4940-91be-1e233bc53e99,scene-segmentation,Scene Segmentation,"Scene segmentation is the task of splitting a scene into its various object components.
554
555Image adapted from [Temporally coherent 4D reconstruction of complex dynamic scenes](https://paperswithcode.com/paper/temporally-coherent-4d-reconstruction-of2).",computer-vision556cb72fd30-befb-4c26-aead-354ad7da525f,fine-grained-visual-recognition,Fine-Grained Visual Recognition,,computer-vision5575aea1b66-b51a-4129-b773-7549932a1d8e,object-discovery,Object Discovery,"**Object Discovery** is the task of identifying previously unseen objects.558 559 560<span class=""description-source"">Source: [Unsupervised Object Discovery and Segmentation of RGBD-images ](https://arxiv.org/abs/1710.06929)</span>",computer-vision561acfd1896-5bc7-4c35-8d28-b5bed2b2a8b1,multimodal-forgery-detection,Multimodal Forgery Detection,**Multimodal Forgery Detection** task is a deep forgery detection method which uses both video and audio.,computer-vision5626eaf348f-170f-4fa0-8dd1-cd683a8f91a4,blind-image-quality-assessment,Blind Image Quality Assessment,,computer-vision563b6f95aeb-9089-46fd-9e6c-66d73684daa0,frame-duplication-detection,Frame Duplication Detection,,computer-vision564d20cf4e1-9548-4ab9-b590-fb12ba937fc9,simultaneous-localization-and-mapping,Simultaneous Localization and Mapping,"Simultaneous localization and mapping (SLAM) is the task of constructing or updating a map of an unknown environment while simultaneously keeping track of an agent's location within it.
565
566<span style=""color:grey; opacity: 0.6"">( Image credit: [ORB-SLAM2](https://arxiv.org/pdf/1610.06475v2.pdf) )</span>",computer-vision56733f37971-c887-49e7-b5ae-1314d3a3de01,video-grounding,Video Grounding,"**Video grounding** is the task of linking spoken language descriptions to specific video segments. In video grounding, the model is given a video and a natural language description, such as a sentence or a caption, and its goal is to identify the specific segment of the video that corresponds to the description. This can involve tasks such as localizing the objects or actions mentioned in the description within the video, or associating a specific time interval with the description.",computer-vision568a99223d5-5ada-477d-b61f-0a86db4dfb23,point-cloud-registration,Point Cloud Registration,"**Point Cloud Registration** is a fundamental problem in 3D computer vision and photogrammetry. Given several sets of points in different coordinate systems, the aim of registration is to find the transformation that best aligns all of them into a common coordinate system. Point Cloud Registration plays a significant role in many vision applications such as 3D model reconstruction, cultural heritage management, landslide monitoring and solar energy analysis.569 570 571<span class=""description-source"">Source: [Iterative Global Similarity Points : A robust coarse-to-fine integration solution for pairwise 3D point cloud registration ](https://arxiv.org/abs/1808.03899)</span>",computer-vision572e03d10a2-8421-4c63-9198-46ba4f245d0b,3d-face-reconstruction,3D Face Reconstruction,"**3D Face Reconstruction** is a computer vision task that involves creating a 3D model of a human face from a 2D image or a set of images. The goal of 3D face reconstruction is to reconstruct a digital 3D representation of a person's face, which can be used for various applications such as animation, virtual reality, and biometric identification.
573
574<span style=""color:grey; opacity: 0.6"">( Image credit: [3DDFA_V2](https://github.com/cleardusk/3DDFA_V2) )</span>",computer-vision575278f9c04-eca5-4374-9804-57a1381cefa3,3d-car-instance-understanding,3D Car Instance Understanding,"3D Car Instance Understanding is the task of estimating properties (e.g.translation, rotation and shape) of a moving or parked vehicle on the road.
576
577<span style=""color:grey; opacity: 0.6"">( Image credit: [Occlusion-Net](http://openaccess.thecvf.com/content_CVPR_2019/papers/Reddy_Occlusion-Net_2D3D_Occluded_Keypoint_Localization_Using_Graph_Networks_CVPR_2019_paper.pdf) )</span>",computer-vision57808d7f4fa-d560-4cf4-8e51-9c7790f4c989,compositional-zero-shot-learning,Compositional Zero-Shot Learning,"**Compositional Zero-Shot Learning (CZSL)** is a computer vision task in which the goal is to recognize unseen compositions fromed from seen state and object during training. The key challenge in CZSL is the inherent entanglement between the state and object within the context of an image. Some example benchmarks for this task are MIT-states, UT-Zappos, and C-GQA. Models are usually evaluated with the Accuracy for both seen and unseen compositions, as well as their Harmonic Mean(HM).
579
580<span style=""color:grey; opacity: 0.6"">( Image credit: [Heosuab](https://hellopotatoworld.tistory.com/24) )</span>",computer-vision5819da8b6ea-ffb9-46be-a5c6-4f0506e2ff86,video-to-shop,Video-to-Shop,,computer-vision582d0c51986-42fa-472e-b5c1-0875b18f76a5,3d-semantic-scene-completion,3D Semantic Scene Completion,"This task was introduced in ""Semantic Scene Completion from a Single Depth Image"" (https://arxiv.org/abs/1611.08974) at CVPR 2017 . The target is to infer the dense 3D voxelized semantic scene from an incompleted 3D input (e.g. point cloud, depth map) and an optional RGB image. A recent summary can be found in the paper ""3D Semantic Scene Completion: a Survey"" (https://arxiv.org/abs/2103.07466), published at IJCV 2021.",computer-vision58327334d01-0264-46ee-96d2-5acb384e72e7,multispectral-object-detection,Multispectral Object Detection,,computer-vision584ab0c58a8-c6c4-456b-8cd9-c439aa964d64,video-based-workflow-recognition,Video Based Workflow Recognition,,computer-vision58524bbae38-a471-4616-8978-e55514f77962,deception-detection-in-videos,Deception Detection In Videos,,computer-vision5869d52b4e7-f6a6-473f-8def-f40481eca435,dense-captioning,Dense Captioning,,computer-vision58784418cfa-fd0d-48eb-8ca1-6432df553744,amodal-panoptic-segmentation,Amodal Panoptic Segmentation,The goal of this task is to simultaneously predict the pixel-wise semantic segmentation labels of the visible regions of stuff classes and the instance segmentation labels of both the visible and occluded regions of thing classes.,computer-vision588d7510111-0f3e-42df-a405-cb15bd12395e,3d-shape-modeling,3D Shape Modeling,Image: [Gkioxari et al](https://arxiv.org/pdf/1906.02739v2.pdf),computer-vision5893e1fcbf6-ba45-44e4-bc5c-fa884f4eb2f9,visual-grounding,Visual Grounding,"Visual Grounding (VG) aims to locate the most relevant object or region in an image, based on a natural language query. The query can be a phrase, a sentence, or even a multi-round dialogue. There are three main challenges in
590VG:
591
592* What is the main focus in a query?
593* How to understand an image?
594* How to locate an object?",computer-vision59541cd7aae-743f-4d7b-a350-9bf4f95ee71a,referring-expression,Referring Expression,"Referring expressions places a bounding box around
596the instance corresponding to the provided description and
597image.",computer-vision5983520dd70-da4d-4f03-9c50-f28944eae6e4,document-layout-analysis,Document Layout Analysis,"""**Document Layout Analysis** is performed to determine physical structure of a document, that is, to determine document components. These document components can consist of single connected components-regions [...] of
599pixels that are adjacent to form single regions [...] , or group
600of text lines. A text line is a group of characters, symbols,
601and words that are adjacent, “relatively close” to each other
602and through which a straight line can be drawn (usually with
603horizontal or vertical orientation)."" L. O'Gorman, ""The document spectrum for page layout analysis,"" in IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 15, no. 11, pp. 1162-1173, Nov. 1993.
604
605Image credit: [PubLayNet: largest dataset ever for document layout analysis](https://arxiv.org/pdf/1908.07836v1.pdf)",computer-vision606c4464b8c-d9dc-4cb2-894b-33ce25325286,multi-object-colocalization,Multi-object colocalization,,computer-vision607a8001a41-cc88-4992-8227-ee6630670bde,feature-compression,Feature Compression,"Compress data for machine interpretability to perform downstream tasks, rather than for human perception.",computer-vision60810f5532f-56a5-4f56-9481-f3683ff74dfa,visual-prompting,Visual Prompting,"Visual Prompting is the task of streamlining computer vision processes by harnessing the power of prompts,
609inspired by the breakthroughs of text prompting in NLP. This innovative approach involves using a few visual
610prompts to swiftly convert an unlabeled dataset into a deployed model, significantly reducing development time
611for both individual projects and enterprise solutions.",computer-vision612eb786d0f-96f2-4d13-bd26-3ec85a2db925,unsupervised-text-recognition,Unsupervised Text Recognition,Decompose a text into the letters / tokens that are used to write it.,computer-vision613637b089c-17a2-4ada-8768-084549df3bb4,3d-pose-estimation,3D Pose Estimation,"Image credit: [GSNet: Joint Vehicle Pose and Shape Reconstruction with Geometrical and Scene-aware Supervision
614, ECCV'20](https://www.ecva.net/papers/eccv_2020/papers_ECCV/papers/123600511.pdf)",computer-vision61598fe3ade-d8e0-4fb9-bf51-156a22279e33,video-super-resolution,Video Super-Resolution,"**Video Super-Resolution** is a computer vision task that aims to increase the resolution of a video sequence, typically from lower to higher resolutions. The goal is to generate high-resolution video frames from low-resolution input, improving the overall quality of the video.
616
617<span style=""color:grey; opacity: 0.6"">( Image credit: [Detail-revealing Deep Video Super-Resolution](https://github.com/jiangsutx/SPMC_VideoSR) )</span>",computer-vision618b2df5834-32c9-4b52-9fa1-c76218a3b6bb,point-cloud-super-resolution,Point Cloud Super Resolution,"Point cloud super-resolution is a fundamental problem
619for 3D reconstruction and 3D data understanding. It takes
620a low-resolution (LR) point cloud as input and generates
621a high-resolution (HR) point cloud with rich details",computer-vision6225af47daa-447a-4599-b679-2dfee0ce9de5,image-variation,Image-Variation,"Given an image, generate variations of the image",computer-vision62379fbbe33-564a-47de-8f68-eda7f9762683,gaze-target-estimation,Gaze Target Estimation,Gaze Target Estimation refers to predicting the image 2D gaze location of a person in the image.,computer-vision624e5c5ee2c-f094-43d6-bb74-c5f20c2df8b4,robust-bev-detection,Robust BEV Detection,,computer-vision6251e2b958e-105f-4148-af00-7d902f1a2b9e,aesthetic-image-captioning,Aesthetic Image Captioning,,computer-vision626cea8ced2-eeb2-4104-9a31-7d7822826a5c,video-denoising,Video Denoising,,computer-vision6277756bf4d-a336-4f16-b54f-440210a9f86b,online-surgical-phase-recognition,Online surgical phase recognition,"Online surgical phase recognition: the first 40 videos to train, the last 40 videos to test.",computer-vision628c6e9c36e-2650-4e43-9107-1f1d0321c0ba,rain-removal,Rain Removal,,computer-vision62909ba3e51-f38b-4fdb-bde7-29dfc1dce88e,image-outpainting,Image Outpainting,"Predicting the visual context of an image beyond its boundary.
630
631Image credit: [NUWA-Infinity: Autoregressive over Autoregressive Generation for Infinite Visual Synthesis](https://paperswithcode.com/paper/nuwa-infinity-autoregressive-over?from=n35)",computer-vision632ca422a33-20fe-44db-8ac0-4d0ba224e23b,object-tracking,Object Tracking,"**Object tracking** is the task of taking an initial set of object detections, creating a unique ID for each of the initial detections, and then tracking each of the objects as they move around frames in a video, maintaining the ID assignment. State-of-the-art methods involve fusing data from RGB and event-based cameras to produce more reliable object tracking. CNN-based models using only RGB images as input are also effective. The most popular benchmark is OTB. There are several evaluation metrics specific to object tracking, including HOTA, MOTA, IDF1, and Track-mAP.
633
634<span style=""color:grey; opacity: 0.6"">( Image credit: [Towards-Realtime-MOT
635](https://github.com/Zhongdao/Towards-Realtime-MOT) )</span>",computer-vision6363623574e-0f5f-465d-a740-581031eff08e,curved-text-detection,Curved Text Detection,,computer-vision637ade41a3e-476f-41b3-a0e2-d473c183d05e,action-understanding,Action Understanding,,computer-vision638943f4181-fc0d-4625-9476-c05cb976c266,multi-label-zero-shot-learning,Multi-label zero-shot learning,,computer-vision639da01d7f8-ed03-4b12-9d6f-46be32f773f3,text-to-video-editing,Text-to-Video Editing,,computer-vision640125ef8c4-66ab-41f2-adf7-097842dff356,generalized-few-shot-classification,Generalized Few-Shot Classification,,computer-vision64140653d8b-3420-4f63-ac76-f078845bf002,3d-point-cloud-classification,3D Point Cloud Classification,Image: [Qi et al](https://arxiv.org/pdf/1612.00593v2.pdf),computer-vision6425c588a74-c816-4110-8cbc-d9b092d5c853,photo-to-caricature-translation,Photo-To-Caricature Translation,"Photo-to-caricature translation is the task of adapting a photo to a cartoon or sketch.
643
644<span style=""color:grey; opacity: 0.6"">( Image credit: [WarpGAN](https://arxiv.org/pdf/1811.10100v3.pdf) )</span>",computer-vision6453c308c7a-cf5d-4119-9e5f-50bec1928475,open-vocabulary-panoptic-segmentation,Open Vocabulary Panoptic Segmentation,,computer-vision646676bc096-5276-45d8-8e16-878ac1a97108,camera-shot-boundary-detection,Camera shot boundary detection,"The objective of camera shot boundary detection is to find the transitions between the camera shots in a video and classify the type of camera transition. This task is introduced in SoccerNet-v2, where 3 types of transitions are considered (abrupt, logo, smooth).",computer-vision647ee8cea23-ab9b-4056-bcca-039f690e1531,jpeg-compression-artifact-reduction,Jpeg Compression Artifact Reduction,,computer-vision648e81a962c-c02a-48b9-b02f-f6fe719ba006,3d-shape-reconstruction-from-a-single-2d,3D Shape Reconstruction From A Single 2D Image,Image: [Liao et al](https://arxiv.org/pdf/1811.12016v1.pdf),computer-vision64909351cc1-4387-4f84-bd91-14253d53495c,3d-face-modeling,3D Face Modelling,,computer-vision650dc0a49dd-c5f4-4974-9c29-bb87664aa1a4,steering-control,Steering Control,,computer-vision651873e90df-2572-45b6-acc4-4bd06d0e17de,event-based-optical-flow,Event-based Optical Flow,,computer-vision6526c63cc64-77e0-43d4-98b5-2625bb5f3572,robust-3d-semantic-segmentation,Robust 3D Semantic Segmentation,3D Semantic Segmentation under Out-of-Distribution Scenarios,computer-vision653d6e4e812-1394-4759-9cbc-9c3c70c1f981,kiss-detection,Kiss Detection,,computer-vision6543b9f01d0-e9d6-4bd2-acf2-95d290ee9a7b,photo-retouching,Photo Retouching,,computer-vision65515d6a799-7eec-4257-b588-326b89dbbbca,handwriting-verification,Handwriting Verification,The goal of handwriting verification is to find a measure of confidence whether the given handwritten samples are written by the same or different writer.,computer-vision65683009cf3-4b1c-4109-a057-daeb15addf6c,automatic-post-editing,Automatic Post-Editing,Automatic post-editing (APE) is used to correct errors in the translation made by the machine translation systems.,computer-vision6575fe519ad-6051-4b87-a57d-d7fd414d4855,sensor-fusion,Sensor Fusion,Sensor fusion is the process of combining sensor data or data derived from disparate sources such that the resulting information has less uncertainty than would be possible when these sources were used individually. [Wikipedia],computer-vision6583e009e95-8a9d-4680-86d2-7519e45c5514,road-segementation,Road Segmentation,Road Segmentation is a pixel wise binary classification in order to extract underlying road network. Various Heuristic and data driven models are proposed. Continuity and robustness still remains one of the major challenges in the area.,computer-vision659309e1a47-4797-4d06-8893-d762175b1238,steganographics,Steganographics,,computer-vision6604715a4b1-8c8c-43ee-bb48-84603307b074,fish-detection,Fish Detection,,computer-vision6610706d656-aa1f-4739-a5f6-2143114a8be6,keypoint-detection,Keypoint Detection,"**Keypoint Detection** involves simultaneously detecting people and localizing their keypoints. Keypoints are the same thing as interest points. They are spatial locations, or points in the image that define what is interesting or what stand out in the image. They are invariant to image rotation, shrinkage, translation, distortion, and so on.
662
663<span style=""color:grey; opacity: 0.6"">( Image credit: [PifPaf: Composite Fields for Human Pose Estimation](https://github.com/vita-epfl/openpifpaf); ""Learning to surf"" by fotologic, license: CC-BY-2.0 )</span>",computer-vision6644d002420-4be0-4d7f-9191-9db02db58379,jpeg-decompression,JPEG Decompression,Image credit: [Palette: Image-to-Image Diffusion Models](https://paperswithcode.com/paper/palette-image-to-image-diffusion-models),computer-vision6654a70ee08-7708-4c6d-a5e8-7caf732565ce,image-to-video-person-re-identification,Image-To-Video Person Re-Identification,,computer-vision666838020a6-86ff-4e4b-b6a7-0c00549798cf,gaze-redirection,gaze redirection,,computer-vision66728e63089-6873-49e1-8464-16c158e2dbeb,set-matching,set matching,,computer-vision6684c95b2a2-c9f3-4f67-b947-fc08dd2f0135,boundary-grounding,Boundary Grounding,"Provided with a description of a boundary inside a video, the machine is required to locate that boundary in the video.",computer-vision66992763350-b118-4f26-8db8-ef200c5c9144,grounded-situation-recognition,Grounded Situation Recognition,"Grounded Situation Recognition aims to produce the structured image summary which describes the primary activity (verb), its relevant entities (nouns), and their bounding-box groundings.",computer-vision67017c37e61-5c91-4bd8-b890-aed3be677b5f,gan-image-forensics,GAN image forensics,,computer-vision6714b685b4b-7212-4543-985d-02dcb0f137e3,video-synchronization,Video Synchronization,,computer-vision672b5754d58-4798-400e-a1af-e377c037aeef,rgb-t-tracking,Rgb-T Tracking,,computer-vision673c1409f98-bbdc-4b80-9a8c-109ca0f5584f,temporal-action-proposal-generation,Temporal Action Proposal Generation,,computer-vision6740edc4b91-96f5-40ac-aea9-8d292b294f59,disjoint-15-5,Disjoint 15-5,,computer-vision675498cb08e-aa4e-4940-9fdd-8234e6c68f8e,image-reconstruction,Image Reconstruction,,computer-vision6765cc9ef69-fe0d-4b8c-8a40-27e42f8ac1ce,single-view-3d-reconstruction,Single-View 3D Reconstruction,,computer-vision6777654f7bb-7aef-4076-8f48-41ad01fe22fb,image-declipping,Image Declipping,,computer-vision6788e64cada-c4f4-4324-b504-10355a3a4c6f,sign-language-recognition,Sign Language Recognition,"**Sign Language Recognition** is a computer vision and natural language processing task that involves automatically recognizing and translating sign language gestures into written or spoken language. The goal of sign language recognition is to develop algorithms that can understand and interpret sign language, enabling people who use sign language as their primary mode of communication to communicate more easily with non-signers.
679
680<span style=""color:grey; opacity: 0.6"">( Image credit: [Word-level Deep Sign Language Recognition from Video:
681A New Large-scale Dataset and Methods Comparison](https://arxiv.org/pdf/1910.11006v1.pdf) )</span>",computer-vision68295b75309-61c5-4d47-a44c-4ffa2f48f8ff,embodied-question-answering,Embodied Question Answering,,computer-vision6834251c6fe-d7cf-483e-be13-58c9ef259be3,concept-alignment,Concept Alignment,**Concept Alignment** aims to align the learned representations or concepts within a model with the intended or target concepts. It involves adjusting the model's parameters or training process to ensure that the learned concepts accurately reflect the underlying patterns in the data.,computer-vision68408f7375e-890e-482e-bf62-3b2c6b3450f8,spatial-relation-recognition,Spatial Relation Recognition,,computer-vision685aa29b25c-e975-4466-92f6-63a3d759e6ff,story-continuation,Story Continuation,"The task involves providing an initial scene that can be obtained in real world use cases. By including this scene, a model can then copy and adapt elements from it as it generates subsequent images.
686
687Source: [StoryDALL-E: Adapting Pretrained Text-to-Image Transformers for Story Continuation](https://paperswithcode.com/paper/storydall-e-adapting-pretrained-text-to-image)",computer-vision68813af0442-7fa8-467a-8d1b-e4e211ebaa41,attentive-segmentation-networks,Attentive segmentation networks,,computer-vision6891c47b7cf-1b4c-4eaf-8bd6-85c3cd78ebdf,3d-multi-person-pose-estimation-absolute,3D Multi-Person Pose Estimation (absolute),"This task aims to solve absolute 3D multi-person pose Estimation (camera-centric coordinates). No ground truth human bounding box and human root joint coordinates are used during testing stage.
690
691<span style=""color:grey; opacity: 0.6"">( Image credit: [RootNet](https://github.com/mks0601/3DMPPE_ROOTNET_RELEASE) )</span>",computer-vision692427c6069-186a-4e81-9887-0cd3eac1e1fa,hand-keypoint-localization,Hand Keypoint Localization,,computer-vision693f5ea788b-1455-4f81-8271-a4cff0e6e01f,3d-facial-expression-recognition,3D Facial Expression Recognition,"3D facial expression recognition is the task of modelling facial expressions in 3D from an image or video.
694
695<span style=""color:grey; opacity: 0.6"">( Image credit: [Expression-Net](https://github.com/fengju514/Expression-Net) )</span>",computer-vision696d8f0667d-8f3b-498b-a266-5836fc6048b9,self-supervised-learning,Self-Supervised Learning,"**Self-Supervised Learning** is proposed for utilizing unlabeled data with the success of supervised learning. Producing a dataset with good labels is expensive, while unlabeled data is being generated all the time. The motivation of Self-Supervised Learning is to make use of the large amount of unlabeled data. The main idea of Self-Supervised Learning is to generate the labels from unlabeled data, according to the structure or characteristics of the data itself, and then train on this unsupervised data in a supervised manner. Self-Supervised Learning is wildly used in representation learning to make a model learn the latent features of the data. This technique is often employed in computer vision, video processing and robot control.
697
698
699<span class=""description-source"">Source: [Self-supervised Point Set Local Descriptors for Point Cloud Registration ](https://arxiv.org/abs/2003.05199)</span>
700
701Image source: [LeCun](https://www.youtube.com/watch?v=7I0Qt7GALVk)",computer-vision70273ac2a1c-1536-47c2-9dc9-5bcbcb0f3cb6,anomaly-detection-at-30-anomaly,Anomaly Detection at 30% anomaly,Performance of unsupervised anomaly detection at specific anomaly percentage.,computer-vision703ff3089e8-fdda-4f90-9925-ea3900ea12d1,activity-recognition,Activity Recognition,"Human **Activity Recognition** is the problem of identifying events performed by humans given a video input. It is formulated as a binary (or multiclass) classification problem of outputting activity class labels. Activity Recognition is an important problem with many societal applications including smart surveillance, video search/retrieval, intelligent robots, and other monitoring systems.704 705 706<span class=""description-source"">Source: [Learning Latent Sub-events in Activity Videos Using Temporal Attention Filters ](https://arxiv.org/abs/1605.08140)</span>",computer-vision707d482e1c9-b03f-48c4-a654-042a4a7ac84b,action-triplet-recognition,Action Triplet Recognition,"Recognising action as a triplet of subject verb and object. Example HOI = Human Object Interaction, Surgical IVT = Instrument Verb Target, etc.",computer-vision708c1e7f1ee-ea14-4e5b-9449-5e11697737ae,image-based-localization,Image-Based Localization,Determining the location of an image without GPS based on cross-view matching. In most of the cases a database of satellite images is used to match the ground images to them.,computer-vision70955c4d34d-d56c-4f22-9ca0-eae35af324ce,robust-object-detection,Robust Object Detection,"A Benchmark for the:
710Robustness of Object Detection Models to Image Corruptions and Distortions
711
712To allow fair comparison of robustness enhancing methods all models have to use a standard ResNet50 backbone because performance strongly scales with backbone capacity. If requested an unrestricted category can be added later.
713
714Benchmark Homepage: https://github.com/bethgelab/robust-detection-benchmark
715
716
717Metrics:
718
719mPC [AP]: Mean Performance under Corruption [measured in AP]
720
721rPC [%]: Relative Performance under Corruption [measured in %]
722
723Test sets:
724Coco: val 2017; Pascal VOC: test 2007; Cityscapes: val;
725
726<span style=""color:grey; opacity: 0.6"">( Image credit: [Benchmarking Robustness in Object Detection](https://arxiv.org/pdf/1907.07484v1.pdf) )</span>",computer-vision727e5252aeb-4531-4b8f-b5cf-0193b79b3109,video-quality-assessment,Video Quality Assessment,"Video Quality Assessment is a computer vision task aiming to mimic video-based human subjective perception. The goal is to produce a mos score, where higher score indicates better perceptual quality. Some well-known benchmarks for this task are KoNViD-1k, LIVE-VQC, YouTube-UGC and LSVQ. SROCC/PLCC/RMSE are usually used to evaluate the performance of different models.",computer-vision728cccbb91e-3d3b-441b-9289-b1dec185dac5,semi-supervised-image-classification,Semi-Supervised Image Classification,"Semi-supervised image classification leverages unlabelled data as well as labelled data to increase classification performance.
729
730You may want to read some blog posts to get an overview before reading the papers and checking the leaderboards:
731
732- [An overview of proxy-label approaches for semi-supervised learning](https://ruder.io/semi-supervised/) - Sebastian Ruder
733- [Semi-Supervised Learning in Computer Vision](https://amitness.com/2020/07/semi-supervised-learning/) - Amit Chaudhary
734
735<span style=""color:grey; opacity: 0.6"">( Image credit: [Self-Supervised Semi-Supervised Learning](https://arxiv.org/pdf/1905.03670v2.pdf) )</span>",computer-vision73651bb8851-5260-47a3-a891-7a294415b460,action-analysis,Action Analysis,,computer-vision737c48e0134-f37f-4181-a4c3-2020414544e2,holdout-set,Holdout Set,,computer-vision73854851971-3e60-4878-bcca-dc21925b5852,gait-recognition-in-the-wild,Gait Recognition in the Wild,"Gait Recognition in the Wild refers to methods under real-world senses, i.e., unconstrained environment.",computer-vision7391580af53-1ee2-4e37-9994-2556c5dd3ae1,3d-object-reconstruction,3D Object Reconstruction,Image: [Choy et al](https://arxiv.org/pdf/1604.00449v1.pdf),computer-vision7402ec2a183-f152-46db-8aca-66179a09448e,materials-imaging,Materials Imaging,,computer-vision741d66875e2-eb68-4ecd-a28a-bd55b16e4797,video-matting,Video Matting,Image credit: [https://arxiv.org/pdf/2012.07810v1.pdf](https://arxiv.org/pdf/2012.07810v1.pdf),computer-vision742e1e18e08-7308-4cbc-9055-c6e0206ac1dc,gait-recognition,Gait Recognition,"<span style=""color:grey; opacity: 0.6"">( Image credit: [GaitSet: Regarding Gait as a Set for Cross-View Gait Recognition](https://github.com/AbnerHqC/GaitSet) )</span>",computer-vision7433a215e28-b496-4b48-83f0-09831d0fe7f3,monocular-3d-object-localization,Monocular 3D Object Localization,,computer-vision7443498f8f1-6b7f-46b3-90d7-d6d3a0b561d4,birds-eye-view-object-detection,Birds Eye View Object Detection,KITTI birds eye view detection task,computer-vision745ac65c824-3010-4034-a2ce-965195961c57,moving-object-detection,Moving Object Detection,,computer-vision74620319665-81a3-45de-a43d-8247a3666405,crosslingual-text-to-image-generation,Crosslingual Text-to-Image Generation,,computer-vision747619ebeb6-340d-49f7-8763-8f8dd5b52557,document-to-image-conversion,Document To Image Conversion,,computer-vision7489c8f015a-01c2-48bc-884e-59bbeb9fce3a,person-identification,Person Identification,,computer-vision7491d12bbed-23e3-47c7-b69a-39dfae3ab9da,medical-image-enhancement,Medical Image Enhancement,Aims to improve the perceptual quality of low-quality medical images,computer-vision750735601e2-1096-406a-b544-030254d1208c,semi-supervised-fashion-compatibility,Semi-Supervised Fashion Compatibility,,computer-vision7512f502ba6-e817-4cea-a277-fced1c22a2c0,visual-tracking,Visual Tracking,"**Visual Tracking** is an essential and actively researched problem in the field of computer vision with various real-world applications such as robotic services, smart surveillance systems, autonomous driving, and human-computer interaction. It refers to the automatic estimation of the trajectory of an arbitrary target object, usually specified by a bounding box in the first frame, as it moves around in subsequent video frames.752 753 754<span class=""description-source"">Source: [Learning Reinforced Attentional Representation for End-to-End Visual Tracking ](https://arxiv.org/abs/1908.10009)</span>",computer-vision755a9c84d55-97bb-4319-8487-d1cb432216dc,overlapped-25-25,Overlapped 25-25,,computer-vision7565ab2491c-c346-41fd-a340-912cb175479b,3d-human-pose-tracking,3D Human Pose Tracking,,computer-vision75746786ae8-834c-4543-87cf-9f44c826b2ec,transparency-separation,Transparency Separation,,computer-vision758a1615da7-9ef2-47ab-af78-da9503db4d62,3d-lane-detection,3D Lane Detection,"The goal of **3D Lane Detection** is to perceive lanes that provide guidance for autonomous vehicles. A lane can be represented as a visible laneline or a conceptual centerline. Furthermore, a lane obtains extra attributes from the understanding of the surrounding environment.
759
760<span style=""color:grey; opacity: 0.6"">( Image credit: [OpenLane-V2](https://github.com/OpenDriveLab/OpenLane-V2 ) )</span>",computer-vision761ae774ea2-0f29-425d-bb5b-1604ba7cabbf,indoor-monocular-depth-estimation,Indoor Monocular Depth Estimation,,computer-vision76290ef2fc9-c460-4104-a6b3-396098161bc9,landmine,Landmine,,computer-vision763ef271038-c8c9-4eac-be62-fbe3870cbc2e,multiple-object-tracking-with-transformer,Multiple Object Tracking with Transformer,,computer-vision7643a14b5bd-0d6d-4f41-8796-27c6fcf2874a,large-scale-person-re-identification,Large-Scale Person Re-Identification,,computer-vision765decc1855-49d2-4045-935e-da8ccaf165de,face-detection,Face Detection,"**Face Detection** is a computer vision task that involves automatically identifying and locating human faces within digital images or videos. It is a fundamental technology that underpins many applications such as face recognition, face tracking, and facial analysis.
766
767<span style=""color:grey; opacity: 0.6"">( Image credit: [insightface](https://github.com/deepinsight/insightface) )</span>",computer-vision76804a6680c-ba5b-478a-ae27-be40703db6ad,infrared-and-visible-image-fusion,Infrared And Visible Image Fusion,Image fusion with paired infrared and visible images,computer-vision7692a633a94-f7f8-4c9a-8254-db2c587ec1e0,3d-point-cloud-linear-classification,3D Point Cloud Linear Classification,Training a linear classifier(e.g. SVM) on the embeddings/representations of 3D point clouds. The embeddings/representations are usually trained in an unsupervised manner.,computer-vision77035d5b3e3-91bc-426b-ab28-bea1cfa5b075,3d-object-detection,3D Object Detection,"**3D Object Detection** is a task in computer vision where the goal is to identify and locate objects in a 3D environment based on their shape, location, and orientation. It involves detecting the presence of objects and determining their location in the 3D space in real-time. This task is crucial for applications such as autonomous vehicles, robotics, and augmented reality.
771
772<span style=""color:grey; opacity: 0.6"">( Image credit: [AVOD](https://github.com/kujason/avod) )</span>",computer-vision773af226e4f-76ec-4c9c-9e32-31c73a059c09,semantic-segmentation,Semantic Segmentation,"**Semantic Segmentation** is a computer vision task in which the goal is to categorize each pixel in an image into a class or object. The goal is to produce a dense pixel-wise segmentation map of an image, where each pixel is assigned to a specific class or object. Some example benchmarks for this task are Cityscapes, PASCAL VOC and ADE20K. Models are usually evaluated with the Mean Intersection-Over-Union (Mean IoU) and Pixel Accuracy metrics.
774
775<span style=""color:grey; opacity: 0.6"">( Image credit: [CSAILVision](https://github.com/CSAILVision/semantic-segmentation-pytorch) )</span>",computer-vision776ff2cdb06-03a2-4da6-bbb8-1396c9cefd05,composite-action-recognition,Composite action recognition,,computer-vision7770a20c938-15cf-48b1-a046-09d6135f1dfc,histopathological-segmentation,Histopathological Segmentation,,computer-vision7783ff46dfa-bf79-4fc9-81a3-e83f5abbe96e,plan2scene,Plan2Scene,Converting floorplans + RGB photos to textured 3D mesh models of houses.,computer-vision779cc1292e0-da78-46f0-95dd-29b7e969bb3d,offline-handwritten-chinese-character,Offline Handwritten Chinese Character Recognition,Handwritten Chinese characters recognition is the task of detecting and interpreting the components of Chinese characters (i.e. radicals and two-dimensional structures).,computer-vision7809cc240b8-8dea-4301-9ca5-3294a1e00a59,unet-quantization,UNET Quantization,,computer-vision7817bc76389-29b5-4d79-979f-fc42fda837ba,de-aliasing,De-aliasing,De-aliasing is the problem of recovering the original high-frequency information that has been aliased during the acquisition of an image.,computer-vision782f3a5e8be-0cb8-489d-ac15-81ae4f6a5544,video-recognition,Video Recognition,"**Video Recognition** is a process of obtaining, processing, and analysing data that it receives from a visual source, specifically video.",computer-vision78398ffaf5b-425a-48dd-b8f6-c7bb548e874a,3d-plane-detection,3D Plane Detection,Image: [Liu et al](https://arxiv.org/pdf/1812.04072v2.pdf),computer-vision7846b972076-1cee-44f5-b3fe-c7e6da180bae,cube-engraving-classification,Cube Engraving Classification,,computer-vision785cfb4e9a9-5e03-4244-847c-535b34c22aeb,supervised-video-summarization,Supervised Video Summarization,"**Supervised video summarization** rely on datasets with human-labeled ground-truth annotations (either in the form of video summaries, as in the case of the [SumMe](https://paperswithcode.com/dataset/summe) dataset, or in the form of frame-level importance scores, as in the case of the [TVSum](https://paperswithcode.com/dataset/tvsum-1) dataset), based on which they try to discover the underlying criterion for video frame/fragment selection and video summarization.
786
787Source: [Video Summarization Using Deep Neural Networks: A Survey](https://arxiv.org/abs/2101.06072)",computer-vision7880b5f5810-4659-4f9f-a9b2-34fe069a42b1,multi-view-subspace-clustering,Multi-view Subspace Clustering,,computer-vision78952fe80d4-26ba-4788-8ae2-1a0e7a0db3ae,self-supervised-anomaly-detection,Self-Supervised Anomaly Detection,Self-Supervision towards anomaly detection,computer-vision790e377656c-298a-492c-ba51-f03e5ff86ece,rgb-d-salient-object-detection,RGB-D Salient Object Detection,"RGB-D Salient object detection (SOD) aims at distinguishing the most visually distinctive objects or regions in a scene from the given RGB and Depth data. It has a wide range of applications, including video/image segmentation, object recognition, visual tracking, foreground maps evaluation, image retrieval, content-aware image editing, information discovery, photosynthesis, and weakly
791supervised semantic segmentation. Here, depth information plays an important complementary role in finding salient objects. Online benchmark: http://dpfan.net/d3netbenchmark.
792
793
794<span style=""color:grey; opacity: 0.6"">( Image credit: [Rethinking RGB-D Salient Object Detection: Models, Data Sets, and Large-Scale Benchmarks, TNNLS20](https://ieeexplore.ieee.org/abstract/document/9107477) )</span>",computer-vision795bef16087-6cbc-4bc5-8c17-406ce0fe977d,synthetic-image-detection,Synthetic Image Detection,Identify if the image is real or generated/manipulated by any generative models (GAN or Diffusion).,computer-vision796cd808f35-4f0b-46f3-9d72-909114b57159,sketch-to-image-translation,Sketch-to-Image Translation,,computer-vision7976489e7e7-59ef-4c78-92e8-d448778765f7,vocabulary-free-image-classification,Vocabulary-free Image Classification,"Recent advances in large vision-language models have revolutionized the image classification paradigm. Despite showing impressive zero-shot capabilities, a pre-defined set of categories, a.k.a. the vocabulary, is assumed at test time for composing the textual prompts. However, such assumption can be impractical when the semantic context is unknown and evolving. Vocabulary-free Image Classification (VIC) aims to assign to an input image a class that resides in an unconstrained language-induced semantic space, without the prerequisite of a known vocabulary.",computer-vision79801c37915-7537-4b25-88eb-fa4aea23fa3e,video-generation,Video Generation,"<span style=""color:grey; opacity: 0.6"">( Various Video Generation Tasks.
799Gif credit: [MaGViT](https://paperswithcode.com/paper/magvit-masked-generative-video-transformer) )</span>",computer-vision800d60fc433-fc1a-468e-8094-2fddd663ce4b,one-shot-instance-segmentation,One-Shot Instance Segmentation,"<span style=""color:grey; opacity: 0.6"">( Image credit: [Siamese Mask R-CNN
801](https://github.com/bethgelab/siamese-mask-rcnn) )</span>",computer-vision80234dd8178-de32-4170-b68d-59e986f88328,explainable-artificial-intelligence,Explainable artificial intelligence,"XAI refers to methods and techniques in the application of artificial intelligence (AI) such that the results of the solution can be understood by humans. It contrasts with the concept of the ""black box"" in machine learning where even its designers cannot explain why an AI arrived at a specific decision. XAI may be an implementation of the social right to explanation. XAI is relevant even if there is no legal right or regulatory requirement—for example, XAI can improve the user experience of a product or service by helping end users trust that the AI is making good decisions. This way the aim of XAI is to explain what has been done, what is done right now, what will be done next and unveil the information the actions are based on. These characteristics make it possible (i) to confirm existing knowledge (ii) to challenge existing knowledge and (iii) to generate new assumptions.",computer-vision8030240288e-72a4-4279-98eb-6940ee1a5d9c,one-shot-3d-action-recognition,One-Shot 3D Action Recognition,,computer-vision8049c493ccc-86ed-4749-9687-bf0878c0a1f2,jpeg-forgery-localization,Jpeg Forgery Localization,,computer-vision8051b51655e-f4ad-4ddb-ad09-37ec41150dc6,image-deblocking,Image Deblocking,,computer-vision8068cb3b2ae-ac31-4970-b7e7-cd09fda1edc3,key-frame-based-video-super-resolution-k-15,Key-Frame-based Video Super-Resolution (K = 15),"Key-Frame-based Video Super-Resolution is a sub-task of [Video Super-Resolution](https://paperswithcode.com/task/video-super-resolution), where, in addition to the low-resolution frames, high-resolution ground-truth frames for every Kth input frame are also provided as inputs to the model. For example, if `[LR-frame-1, LR-frame-2, LR-frame-3, ..., LR-frame-100]` is the sequence of low-resolution frames to be upscaled, the Key-Frame-based Video Super-Resolution (K = 15) model is also provided with the high-resolution frames `[HR-frame-1, HR-frame-16, ..., HR-frame-91]` . Key-frames are excluded when measuring the evaluation metrics.",computer-vision80782c3f07d-a9a2-4b85-8bc0-168b61a772f3,amodal-instance-segmentation,Amodal Instance Segmentation,"Different from traditional segmentation which only focuses on visible regions, amodal instance segmentation also predicts the occluded parts of object instances.
808
809Description Credit: [Deep Occlusion-Aware Instance Segmentation with Overlapping BiLayers, CVPR'21](https://openaccess.thecvf.com/content/CVPR2021/papers/Ke_Deep_Occlusion-Aware_Instance_Segmentation_With_Overlapping_BiLayers_CVPR_2021_paper.pdf)",computer-vision8103d4c14bb-dfec-4e55-bc60-9480d48e08aa,deepfake-detection,DeepFake Detection,"**DeepFake Detection** is the task of detecting fake videos or images that have been generated using deep learning techniques. Deepfakes are created by using machine learning algorithms to manipulate or replace parts of an original video or image, such as the face of a person. The goal of deepfake detection is to identify such manipulations and distinguish them from real videos or images.
811
812Description source: [DeepFakes: a New Threat to Face Recognition? Assessment and Detection](https://arxiv.org/pdf/1812.08685.pdf)
813
814Image source: [DeepFakes: a New Threat to Face Recognition? Assessment and Detection](https://paperswithcode.com/paper/deepfakes-a-new-threat-to-face-recognition)",computer-vision815a5d4923b-b2f3-4d1c-8b85-657cfd641b3a,action-detection,Action Detection,"Action Detection aims to find both where and when an action occurs within a video clip and classify what the action is taking place. Typically results are given in the form of action tublets, which are action bounding boxes linked across time in the video. This is related to temporal localization, which seeks to identify the start and end frame of an action, and action recognition, which seeks only to classify which action is taking place and typically assumes a trimmed video.",computer-vision816f41febfd-8254-4f95-bbf8-999642aa4f75,human-object-interaction-detection,Human-Object Interaction Detection,"Human-Object Interaction (HOI) detection is a task of identifying ""a set of interactions"" in an image, which involves the i) localization of the subject (i.e., humans) and target (i.e., objects) of interaction, and ii) the classification of the interaction labels.",computer-vision8175b01076c-0964-4e4b-8a84-8d9dc4e55f86,lighting-estimation,Lighting Estimation,Lighting Estimation analyzes given images to provide detailed information about the lighting in a scene.,computer-vision81883dfde13-600d-447a-bce5-f6037bd57006,document-image-skew-estimation,Document Image Skew Estimation,,computer-vision8199c05da45-ed29-4a38-9c36-5729a5acf347,road-damage-detection,Road Damage Detection,"Road damage detection is the task of detecting damage in roads.
820
821<span style=""color:grey; opacity: 0.6"">( Image credit: [Road Damage Detection And Classification In Smartphone Captured Images Using Mask R-CNN](https://arxiv.org/pdf/1811.04535v1.pdf) )</span>",computer-vision822c5b220f6-d28f-4092-b82b-ba72f0a0520f,multiple-object-tracking,Multiple Object Tracking,"**Multiple Object Tracking** is the problem of automatically identifying multiple objects in a video and representing them as a set of trajectories with high accuracy.823 824 825<span class=""description-source"">Source: [SOT for MOT ](https://arxiv.org/abs/1712.01059)</span>",computer-vision826be79ed0d-5a2c-4159-812f-20c2e80877fe,depth-and-camera-motion,Depth And Camera Motion,,computer-vision827b8e3f6bb-55a8-40bd-997e-7d3d592b4aac,film-simulation,Film Simulation,Simulate the appearance of film camera.,computer-vision8286cc080ce-5886-4e6a-b9e2-1850951a2539,subject-driven-video-generation,Subject-driven Video Generation,,computer-vision829221cccb4-7481-400c-8735-5f91d5901bc9,occluded-face-detection,Occluded Face Detection,,computer-vision83083d406fc-72da-435d-9725-5762f6b7b2a1,food-recognition,Food Recognition,,computer-vision8315b17ff0e-7078-4419-a321-9f2716ead47e,medical-image-denoising,Medical Image Denoising,Image credit: [Learning Medical Image Denoising with Deep Dynamic Residual Attention Network](https://paperswithcode.com/paper/learning-medical-image-denoising-with-deep),computer-vision8324e97419c-1154-4bd4-b893-16f67a8143a6,3d-semantic-scene-completion-from-a-single,3D Semantic Scene Completion from a single RGB image,This task relies on a single RGB image to infer the dense 3D voxelized semantic scene.,computer-vision8331be75958-7e31-4cdf-a8d7-97e91930ede5,unrolling,Rolling Shutter Correction,Rolling Shutter Correction,computer-vision8343f74793a-206f-45ec-bd6f-b703b1783925,learning-with-coarse-labels,Learning with coarse labels,"Learning fine-grained representation with coarsely-labelled dataset, which can significantly reduce the labelling cost. As a simple example, for the task of differentiation between different pets, we need a knowledgeable cat lover to distinguish between ‘British short’ and ‘Siamese’, but even a child annotator may help to discriminate between ‘cat’ and ‘non-cat’.",computer-vision835932463ce-d625-4775-965f-46dbd51cced8,person-retrieval,Person Retrieval,,computer-vision83685daacc9-e963-4e54-b04d-41ca628cbda3,open-vocabulary-object-detection,Open Vocabulary Object Detection,"Open-vocabulary detection (OVD) aims to generalize beyond the limited number of base classes labeled during the training phase. The goal is to detect novel classes defined by an unbounded
837(open) vocabulary at inference.",computer-vision838e8d08f18-1bda-4c49-86ca-b6e65070d10f,video-emotion-detection,Video Emotion Detection,,computer-vision839099a2a00-b6e0-4d5a-9e75-3adfbd7fb267,surface-normals-estimation-from-point-clouds,Surface Normals Estimation from Point Clouds,Parent task: 3d Point Clouds Analysis,computer-vision84030c76c9c-7cd0-4510-9299-6faf06830340,video-captioning,Video Captioning,"**Video Captioning** is a task of automatic captioning a video by understanding the action and event in the video which can help in the retrieval of the video efficiently through text.
841
842
843<span class=""description-source"">Source: [NITS-VC System for VATEX Video Captioning Challenge 2020 ](https://arxiv.org/abs/2006.04058)</span>",computer-vision844aab59969-aad0-4ab1-a187-6636822e32b5,action-recognition-in-still-images,Action Recognition In Still Images,,computer-vision845f2129983-da52-4135-85fd-b5e71c1f254a,video-compression,Video Compression,"**Video Compression** is a process of reducing the size of an image or video file by exploiting spatial and temporal redundancies within an image or video frame and across multiple video frames. The ultimate goal of a successful Video Compression system is to reduce data volume while retaining the perceptual quality of the decompressed data.
846
847
848<span class=""description-source"">Source: [Adversarial Video Compression Guided by Soft Edge Detection ](https://arxiv.org/abs/1811.10673)</span>",computer-vision849373f644b-22d8-4608-87e3-16c1808b5753,real-time-instance-segmentation,Real-time Instance Segmentation,"Similar to its parent task, instance segmentation, but with the goal of achieving real-time capabilities under a defined setting.
850
851Image Credit: [SipMask: Spatial Information Preservation for Fast Image and Video Instance Segmentation](https://arxiv.org/pdf/2007.14772v1.pdf)",computer-vision8527109440e-1735-4b5c-ba73-ca6d975181ad,few-shot-video-object-detection,Few-Shot Video Object Detection,"Few-Shot Video Object Detection
853(FSVOD): given only a few support images of the target
854object in an unseen class, detect all the objects belonging to
855the same class in a given query video.",computer-vision8564ce7c08b-1d0f-4fe5-8fd5-80f8eb6b6a6e,material-recognition,Material Recognition,,computer-vision857e8f4e05e-c051-41f3-9129-0b8bc2abd155,small-data,Small Data Image Classification,Supervised image classification with tens to hundreds of labeled training examples.,computer-vision85866fcaa84-8a19-4613-994e-075fa3608712,3d-shape-recognition,3D Shape Recognition,Image: [Wei et al](https://arxiv.org/pdf/1908.10098v1.pdf),computer-vision85914bf5bac-340d-4a53-8645-89723be5209e,video-narrative-grounding,Video Narrative Grounding,"**Video Narrative Grounding** is the task of linking video narratives to specific video segments. The input is a video with a text description (the narrative) and the positions of certain nouns marked. For each marked
860noun, the method must output a segmentation mask for the
861object it refers to, in each video frame.
862
863Source: [Connecting Vision and Language with Video Localized Narratives](/paper/connecting-vision-and-language-with-video)",computer-vision864b8c6620f-9ed8-42c9-bcf5-be3f8fdb22c1,hand-joint-reconstruction,Hand Joint Reconstruction,,computer-vision86598c4b4b7-5434-4e0d-9818-5fef4cb6cc19,action-anticipation,Action Anticipation,"Next action anticipation is defined as observing 1, ... , T frames and predicting the action that happens after a gap of T_a seconds. It is important to note that a new action starts after T_a seconds that is not seen in the observed frames. Here T_a=1 second.",computer-vision8663c7e8f5d-a08b-461a-bae8-8253cf1799c0,text-to-image,Text-To-Image,,computer-vision8675d6de600-0ff2-40d9-9e6a-dc4c576cb474,depth-estimation,Depth Estimation,"**Depth Estimation** is the task of measuring the distance of each pixel relative to the camera. Depth is extracted from either monocular (single) or stereo (multiple views of a scene) images. Traditional methods use multi-view geometry to find the relationship between the images. Newer methods can directly estimate depth by minimizing the regression loss, or by learning to generate a novel view from a sequence. The most popular benchmarks are KITTI and NYUv2. Models are typically evaluated according to a RMS metric.
868
869<span class=""description-source"">Source: [DIODE: A Dense Indoor and Outdoor DEpth Dataset ](https://arxiv.org/abs/1908.00463)</span>",computer-vision8701896daef-cdeb-46a5-a9ff-04c744d436db,motion-synthesis,Motion Synthesis,"Image source: [Multi-View Motion Synthesis via Applying Rotated Dual-Pixel Blur Kernels
871](https://paperswithcode.com/paper/multi-view-motion-synthesis-via-applying)",computer-vision8728edc8f5a-a3dd-4573-8768-193b2e272f51,visual-question-answering-1,Visual Question Answering,,computer-vision8733a79202e-62ed-4e1d-8876-cba586729e50,intrinsic-image-decomposition,Intrinsic Image Decomposition,"**Intrinsic Image Decomposition** is the process of separating an image into its formation components such as reflectance (albedo) and shading (illumination). Reflectance is the color of the object, invariant to camera viewpoint and illumination conditions, whereas shading, dependent on camera viewpoint and object geometry, consists of different illumination effects, such as shadows, shading and inter-reflections. Using intrinsic images, instead of the original images, can be beneficial for many computer vision algorithms. For instance, for shape-from-shading algorithms, the shading images contain important visual cues to recover geometry, while for segmentation and detection algorithms, reflectance images can be beneficial as they are independent of confounding illumination effects. Furthermore, intrinsic images are used in a wide range of computational photography applications, such as material recoloring, relighting, retexturing and stylization.874 875 876<span class=""description-source"">Source: [CNN based Learning using Reflection and Retinex Models for Intrinsic Image Decomposition ](https://arxiv.org/abs/1712.01056)</span>",computer-vision87770a001d8-3016-4f2f-a8e6-2ebde8882c30,neural-rendering,Neural Rendering,"Given a representation of a 3D scene of some kind (point cloud, mesh, voxels, etc.), the task is to create an algorithm that can produce photorealistic renderings of this scene from an arbitrary viewpoint. Sometimes, the task is accompanied by image/scene appearance manipulation.",computer-vision8786c6c2c46-6d6b-4a96-8ca3-733759bac32c,segmenting-flooded-buildings,Flooded Building Segmentation,,computer-vision879ab2f7783-5cf9-492f-805d-0cd94ff3dcf7,text-to-face-generation,Text-to-Face Generation,,computer-vision8807f4a57f5-ef98-436b-8482-f328c94a0cb5,one-shot-visual-object-segmentation,One-shot visual object segmentation,,computer-vision881efaa1caf-a09d-400b-970e-e2e09d50d199,weakly-supervised-segmentation,Weakly supervised segmentation,,computer-vision8823c25212c-bc58-498c-a2a4-0f963960bee6,real-time-visual-tracking,Real-Time Visual Tracking,,computer-vision88318f7908f-e35a-4000-b87a-c960ec746287,plant-phenotyping,Plant Phenotyping,,computer-vision8849297b2f2-92f5-47a2-89ac-9c815dfffe32,3d-human-dynamics,3D Human Dynamics,Image: [Zhang et al](https://openaccess.thecvf.com/content_ICCV_2019/papers/Zhang_Predicting_3D_Human_Dynamics_From_Video_ICCV_2019_paper.pdf),computer-vision885722ffd4e-0e86-4a31-b55f-7ff081a65011,unsupervised-anomaly-detection-with-specified-5,Unsupervised Anomaly Detection with Specified Settings -- 1% anomaly,,computer-vision8867e4d8c53-28c3-4518-9d0f-539ed0fc7be4,3d-object-recognition,3D Object Recognition,"3D object recognition is the task of recognising objects from 3D data.
887
888Note that there are related tasks you can look at, such as [3D Object Detection](https://paperswithcode.com/task/3d-object-detection) which have more leaderboards.
889
890<span style=""color:grey; opacity: 0.6"">(Image credit: [Look Further to Recognize Better](https://arxiv.org/pdf/1907.12924v1.pdf))</span>",computer-vision89127920077-c03b-4842-afa3-219313f41f54,object-proposal-generation,Object Proposal Generation,"Object proposal generation is a preprocessing technique that has been widely used in current object detection pipelines to guide the search of objects and avoid exhaustive sliding window search across images.
892
893<span style=""color:grey; opacity: 0.6"">( Image credit: [Multiscale Combinatorial Grouping
894for Image Segmentation and Object Proposal Generation](https://arxiv.org/pdf/1503.00848v4.pdf) )</span>",computer-vision89589a8a2a1-d9c7-4235-a2f4-b64e47a3999f,3d-shape-generation,3D Shape Generation,Image: [Mo et al](https://arxiv.org/pdf/1908.00575v1.pdf),computer-vision896e36051d6-c65a-43e6-ae75-6da02ab89688,reflection-removal,Reflection Removal,,computer-vision8976b005646-1a0d-4f48-ac77-b83253edfe14,multi-animal-tracking-with-identification,Multi-Animal Tracking with identification,Tracking all animals in a video maintaining their identities after touches or occlusions.,computer-vision8984a6244ba-ce3b-4a37-bc19-34f966c1d0d6,multi-person-pose-forecasting,Multi-Person Pose forecasting,,computer-vision89948462ea5-ebf9-43dd-84f1-d9b6d0c09862,2d-cyclist-detection,2D Cyclist Detection,,computer-vision90012466ee9-d9ce-4039-9ec8-a4e17541ab31,real-time-semantic-segmentation,Real-Time Semantic Segmentation,"Semantic Segmentation is a computer vision task that involves assigning a semantic label to each pixel in an image. In **Real-Time Semantic Segmentation**, the goal is to perform this labeling quickly and accurately in real-time, allowing for the segmentation results to be used for tasks such as object recognition, scene understanding, and autonomous navigation.
901
902<span style=""color:grey; opacity: 0.6"">( Image credit: [TorchSeg](https://github.com/ycszen/TorchSeg) )</span>",computer-vision903fc2c0b6b-1212-40ce-9c22-7b69feb06e85,generalized-referring-expression-segmentation,Generalized Referring Expression Segmentation,"Generalized Referring Expression Segmentation (GRES), introduced by [Liu et al in CVPR 2023](https://henghuiding.github.io/GRES/), allows expressions indicating any number of target objects. GRES takes an image and a referring expression as input, and requires mask prediction of the target object(s).",computer-vision9045eb83ba1-a912-4490-8883-3e184428da17,disparity-estimation,Disparity Estimation,The Disparity Estimation is the task of finding the pixels in the multiscopic views that correspond to the same 3D point in the scene.,computer-vision905433e341b-4e12-4c20-8a2a-0e7640733073,disjoint-10-1,Disjoint 10-1,,computer-vision90619007a7d-13fb-4fc8-94fa-8ee3270b9b75,mental-workload-estimation,Mental Workload Estimation,,computer-vision9075448e0eb-8b0a-43c6-9440-4d4e19f9434a,repetitive-action-counting,Repetitive Action Counting,Repetitive action counting aims to count the number of repetitive actions in a video.,computer-vision9083badd92c-66de-402f-bb73-8199bf66f81d,person-recognition,Person Recognition,,computer-vision909522bb7fa-1145-4c9c-bac7-847a7bd60092,video-emotion-recognition,Video Emotion Recognition,,computer-vision910366d4503-edae-4b00-b6eb-5639fdf34161,hybrid-positioning,Hybrid Positioning,Hybrid Positioning using CV and dead reckoning,computer-vision91196866937-eec4-4d38-9d92-4129b838aa95,indoor-localization,Indoor Localization,Indoor localization is a fundamental problem in indoor location-based applications.,computer-vision912238702ee-479b-4ab8-b410-0854afe06c7e,face-sketch-synthesis,Face Sketch Synthesis,"Face sketch synthesis is the task of generating a sketch from an input face photo.
913
914<span style=""color:grey; opacity: 0.6"">( Image credit: [High-Quality Facial Photo-Sketch Synthesis Using Multi-Adversarial Networks](https://arxiv.org/pdf/1710.10182v2.pdf) )</span>",computer-vision9159d1b8893-13c5-48d9-8922-4bb8488497c3,3d-scene-reconstruction,3D Scene Reconstruction,Creating 3D scene either using conventional SFM pipelines or latest deep learning approaches.,computer-vision9163d1b8779-92fe-4028-b197-b5df1044cd1e,facial-emotion-recognition,Facial Emotion Recognition,Emotion Recognition from facial images,computer-vision917b33fc841-b20c-41bc-ac7b-d8ef19113e78,lipreading,Lipreading,"Lipreading is a process of extracting speech by watching lip movements of a speaker in the absence of sound. Humans lipread all the time without even noticing. It is a big part in communication albeit not as dominant as audio. It is a very helpful skill to learn especially for those who are hard of hearing.
918
919Deep Lipreading is the process of extracting speech from a video of a silent talking face using deep neural networks. It is also known by few other names: Visual Speech Recognition (VSR), Machine Lipreading, Automatic Lipreading etc.
920
921The primary methodology involves two stages: i) Extracting visual and temporal features from a sequence of image frames from a silent talking video ii) Processing the sequence of features into units of speech e.g. characters, words, phrases etc. We can find several implementations of this methodology either done in two separate stages or trained end-to-end in one go.",computer-vision922abe03496-1ebb-4cf2-ba6b-ca2e096c360f,multi-exposure-image-fusion,Multi-Exposure Image Fusion,,computer-vision9238fc6c93d-6649-44f6-b293-5fe359fb7d93,3d-point-cloud-matching,3D Point Cloud Matching,Image: [Gojic et al](https://openaccess.thecvf.com/content_CVPR_2019/papers/Gojcic_The_Perfect_Match_3D_Point_Cloud_Matching_With_Smoothed_Densities_CVPR_2019_paper.pdf),computer-vision924d1487389-eec1-4b99-8f7f-978fcf49f585,camera-localization,Camera Localization,,computer-vision9257a20033c-5375-4a7e-926b-85a5c72502de,hd-semantic-map-learning,HD semantic map learning,"The goal of task is to generate map elements in a vectorized form using data from onboard sensors, e.g., RGB cameras and/or LiDARs. These map elements include but are not limited to : Road boundaries, boundaries of roads that split roads and sidewalks.",computer-vision926e74623ba-f141-46c2-b992-f0c68433c620,3d-volumetric-reconstruction,3D Volumetric Reconstruction,Image: [Grinvald et al](https://arxiv.org/pdf/1903.00268.pdf),computer-vision9273bc173b6-e5ad-436d-b2b0-16ae7abb3b7e,camera-auto-calibration,Camera Auto-Calibration,,computer-vision92874bcf158-5beb-41b1-8d2c-3a357e3eddec,face-swapping,Face Swapping,"Face swapping refers to the task of swapping faces between images or in an video, while maintaining the rest of the body and environment context.
929
930<span style=""color:grey; opacity: 0.6"">( Image credit: [Swapped Face Detection using Deep Learning and Subjective Assessment](https://arxiv.org/pdf/1909.04217v1.pdf) )</span>",computer-vision931c6ba86ff-e52b-4dae-8a41-4ed07e113fc9,change-detection,Change Detection,"**Change Detection** is a computer vision task that involves detecting changes in an image or video sequence over time. The goal is to identify areas in the image or video that have undergone changes, such as appearance changes, object disappearance or appearance, or even changes in the scene's background.
932
933Image credit: [""A TRANSFORMER-BASED SIAMESE NETWORK FOR CHANGE DETECTION""](https://arxiv.org/pdf/2201.01293v1.pdf)",computer-vision934577c91e2-f8ca-4e78-abe0-1a4cc5588107,gesture-recognition,Gesture Recognition,"**Gesture Recognition** is an active field of research with applications such as automatic recognition of sign language, interaction of humans and robots or for new ways of controlling video games.935 936 937<span class=""description-source"">Source: [Gesture Recognition in RGB Videos Using Human Body Keypoints and Dynamic Time Warping ](https://arxiv.org/abs/1906.12171)</span>",computer-vision938c92e916a-5310-4520-b9a8-ce4a512a210b,aesthetics-quality-assessment,Aesthetics Quality Assessment,Automatic assessment of aesthetic-related subjective ratings.,computer-vision9396d326e9a-913b-476f-b476-8688f6f5f00a,anomaly-detection-in-surveillance-videos,Anomaly Detection In Surveillance Videos,,computer-vision9406b5ddf9e-5b3c-42d8-b266-49e57523697d,metric-learning,Metric Learning,"The goal of **Metric Learning** is to learn a representation function that maps objects into an embedded space. The distance in the embedded space should preserve the objects’ similarity — similar objects get close and dissimilar objects get far away. Various loss functions have been developed for Metric Learning. For example, the **contrastive loss** guides the objects from the same class to be mapped to the same point and those from different classes to be mapped to different points whose distances are larger than a margin. **Triplet loss** is also popular, which requires the distance between the anchor sample and the positive sample to be smaller than the distance between the anchor sample and the negative sample.
941
942
943<span class=""description-source"">Source: [Road Network Metric Learning for Estimated Time of Arrival ](https://arxiv.org/abs/2006.13477)</span>",computer-vision944a27c88bf-5f84-44ed-ab8c-166f9ad77470,reverse-style-transfer,Reverse Style Transfer,,computer-vision945680e3f1c-b337-4e25-a1e6-46ead492aadb,handwriting-recognition,Handwriting Recognition,Image source: [Handwriting Recognition of Historical Documents with few labeled data](https://arxiv.org/pdf/1811.07768v1.pdf),computer-vision9461e45e47c-5ef1-48bf-9978-92ba3e93de7a,dial-meter-reading,Dial Meter Reading,,computer-vision947ff3f3ed5-9599-4b53-98ca-89a1b141bfd7,cross-domain-few-shot,Cross-Domain Few-Shot,,computer-vision94841d011cc-0899-448d-b0ad-d2fe997c8ae7,hurricane-forecasting,Hurricane Forecasting,"Tropical Cyclone Forecasting using Computer Vision, Deep Learning, and Time-Series methods",computer-vision94912bc82d6-1cea-4831-9176-450e113f2a88,supervised-dimensionality-reduction,Supervised dimensionality reduction,,computer-vision9503c80b37d-9882-4d8a-b276-7699a011b266,saliency-detection,Saliency Detection,"**Saliency Detection** is a preprocessing step in computer vision which aims at finding salient objects in an image.951 952 953<span class=""description-source"">Source: [An Unsupervised Game-Theoretic Approach to Saliency Detection ](https://arxiv.org/abs/1708.02476)</span>",computer-vision9547b4642e9-b9e5-4ebd-91e9-1c052f646a64,semantic-image-matting,Semantic Image Matting,,computer-vision95573e3e913-48f1-40f1-a51b-0cc9b8038e2d,license-plate-recognition,License Plate Recognition,,computer-vision95650a086e1-cd47-43c1-accc-910bcfa86c53,disjoint-15-1,Disjoint 15-1,,computer-vision9573640034f-3d62-483d-91ff-4ed897d201dd,dense-video-captioning,Dense Video Captioning,"Most natural videos contain numerous events. For example, in a video of a “man playing a piano”, the video might also contain “another man dancing” or “a crowd clapping”. The task of dense video captioning involves both detecting and describing events in a video.",computer-vision958312e03b9-6d87-406f-aacc-76b1a9491fd0,cross-domain-few-shot-learning,cross-domain few-shot learning,Its essence is transfer learning. The model needs to be trained in the source domain and then migrated to the target domain. Compliant with (1) the category in the target domain has never appeared in the source domain (2) the data distribution of the target domain is inconsistent with the source domain (3) each class in the target domain has very few labels,computer-vision9591a68b54f-e0eb-4024-be24-562b4d9ea2ab,age-and-gender-classification,Age And Gender Classification,"Age and gender classification is a dual-task of identifying the age and gender of a person from an image or video.
960
961<span style=""color:grey; opacity: 0.6"">( Image credit: [Multi-Expert Gender Classification on Age Group by Integrating Deep Neural Networks](https://arxiv.org/pdf/1809.01990v2.pdf) )</span>",computer-vision9622898c35c-b2ee-495c-bd92-6b7c6c85c63e,removing-text-from-natural-images,Image Text Removal,,computer-vision96385eedb32-017a-47bc-a451-423a0d4c7204,drone-navigation,Drone navigation,"(Satellite -> Drone) Given one satellite-view image, the drone intends to find the most relevant place (drone-view images) that it has passed by. According to its flight history, the drone could be navigated back to the target place.",computer-vision964d092e653-447e-4a28-920d-6ae72f61cd3b,handwritten-document-recognition,Handwritten Document Recognition,,computer-vision965eac0d5a3-3369-445a-a53e-e5ffa9a93dfe,disguised-face-verification,Disguised Face Verification,,computer-vision9665d3ff9f4-ec25-49c3-967f-b6d368d12ae4,layout-to-image-generation,Layout-to-Image Generation,"Layout-to-image generation its the task to generate a scene based on the given layout. The layout describes the location of the objects to be included in the output image.
967In this section, you can find state-of-the-art leaderboards for Layout-to-image generation.",computer-vision9685c868094-7a4b-4361-85a3-faaa15f55730,transform-a-video-into-a-comics,Transform A Video Into A Comics,,computer-vision96947ffac89-be1e-4ac9-9f45-5c12c41d0b97,image-stitching,Image Stitching,"**Image Stitching** is a process of composing multiple images with narrow but overlapping fields of view to create a larger image with a wider field of view.
970
971
972<span class=""description-source"">Source: [Single-Perspective Warps in Natural Image Stitching ](https://arxiv.org/abs/1802.04645)</span>
973
974( Image credit: [Kornia](https://github.com/kornia/kornia) )",computer-vision975f6942401-b4d6-4b14-94ed-9f7d89235d36,markerless-motion-capture,Markerless Motion Capture,,computer-vision97695c02a6c-9678-4c84-98ba-2e8b8274abaa,visual-crowd-analysis,Visual Crowd Analysis,,computer-vision9770584f78c-9e01-40ae-b3d7-eed262d507ec,highlight-detection,Highlight Detection,,computer-vision97801293107-0aab-4fa7-92f7-3a6b301acaf4,saliency-ranking,Saliency Ranking,,computer-vision979e0eca96f-5ee9-48fd-9ca7-39f4ec192793,point-set-upsampling,Point Set Upsampling,,computer-vision980a9a325c9-4678-44ba-96c0-287addecd3fb,scene-generation,Scene Generation,,computer-vision981e14eee3d-3ed6-4cc4-8db1-b2d595cda7cb,optical-flow-estimation,Optical Flow Estimation,"**Optical Flow Estimation** is a computer vision task that involves computing the motion of objects in an image or a video sequence. The goal of optical flow estimation is to determine the movement of pixels or features in the image, which can be used for various applications such as object tracking, motion analysis, and video compression.
982
983Approaches for optical flow estimation include correlation-based, block-matching, feature tracking, energy-based, and more recently gradient-based.
984
985Further readings:
986
987- [Optical Flow Estimation](https://www.cs.toronto.edu/~fleet/research/Papers/flowChapter05.pdf)
988- [Performance of Optical Flow Techniques](https://www.cs.toronto.edu/~fleet/research/Papers/ijcv-94.pdf)
989
990Definition source: [Devon: Deformable Volume Network for Learning Optical Flow ](https://arxiv.org/abs/1802.07351)
991
992Image credit: [Optical Flow Estimation](https://www.cs.toronto.edu/~fleet/research/Papers/flowChapter05.pdf)",computer-vision993fdb1da93-4426-4ce0-bd5e-903dad058e08,image-clustering,Image Clustering,"Models that partition the dataset into semantically meaningful clusters without having access to the ground truth labels.
994
995<span style=""color:grey; opacity: 0.6""> Image credit: ImageNet clustering results of [SCAN: Learning to Classify Images without Labels (ECCV 2020)](https://arxiv.org/abs/2005.12320) </span>",computer-vision996235dac22-1f0c-491b-a66e-2c6c93191c62,image-classification,Image Classification,"**Image Classification** is a fundamental task that attempts to comprehend an entire image as a whole. The goal is to classify the image by assigning it to a specific label. Typically, Image Classification refers to images in which only one object appears and is analyzed. In contrast, object detection involves both classification and localization tasks, and is used to analyze more realistic cases in which multiple objects may exist in an image.
997
998
999<span class=""description-source"">Source: [Metamorphic Testing for Object Detection Systems ](https://arxiv.org/abs/1912.12162)</span>",computer-vision1000e7c378b1-002f-4c49-a28a-a6a888442c07,video-to-video-synthesis,Video-to-Video Synthesis,,computer-vision1001ffafcd7e-2770-47a9-8855-f6bcc68d447c,referring-image-matting,Referring Image Matting,"Extracting the meticulous alpha matte of the specific object from the image that can best match the given natural language description, e.g., a keyword or a expression.",computer-vision1002e3598820-eadc-4448-ac36-af1e3d4b0361,moment-retrieval,Moment Retrieval,"Moment retrieval can de defined as the task of ""localizing moments in a video given a user query"".
1003
1004Description from: [QVHIGHLIGHTS: Detecting Moments and Highlights in Videos via Natural Language Queries](https://arxiv.org/pdf/2107.09609v1.pdf)
1005
1006Image credit: [QVHIGHLIGHTS: Detecting Moments and Highlights in Videos via Natural Language Queries](https://arxiv.org/pdf/2107.09609v1.pdf)",computer-vision1007b6562155-6600-48ae-8107-853385151224,3d-depth-estimation,3D Depth Estimation,Image: [monodepth2](https://github.com/nianticlabs/monodepth2),computer-vision1008efcd6359-9071-4c02-9c7b-fb23c16645b1,reference-based-video-super-resolution,Reference-based Video Super-Resolution,"Reference-based video super-resolution (RefVSR) is an expansion of reference-based super-resolution (RefSR) to the video super-resolution (VSR). RefVSR inherits the objectives of both RefSR and VSR tasks and utilizes a Ref video for reconstructing an HR video from an LR video
1009video from an LR video.",computer-vision10108b062738-1c44-410a-89e7-a3cebdeed796,multi-class-one-shot-image-synthesis,Multi class one-shot image synthesis,The goal of Multi-class one-shot image synthesis is to learn a generative model that can generate samples with visual attributes from as few as one or more images of at least 2 related classes.,computer-vision10113a9a82ad-57ca-4033-99d5-ab381f9abad7,image-enhancement,Image Enhancement,"**Image Enhancement** is basically improving the interpretability or perception of information in images for human viewers and providing ‘better’ input for other automated image processing techniques. The principal objective of Image Enhancement is to modify attributes of an image to make it more suitable for a given task and a specific observer.1012 1013 1014<span class=""description-source"">Source: [A Comprehensive Review of Image Enhancement Techniques ](https://arxiv.org/abs/1003.4053)</span>",computer-vision1015eb0cb877-bcc6-49c8-ab43-fe4280c08811,video-text-retrieval,Video-Text Retrieval,Video-Text retrieval requires understanding of both video and language together. Therefore it's different to video retrieval task.,computer-vision1016b00cc83c-a383-4f55-ae12-3975bc69dd35,video-description,Video Description,"The goal of automatic **Video Description** is to tell a story about events happening in a video. While early Video Description methods produced captions for short clips that were manually segmented to contain a single event of interest, more recently dense video captioning has been proposed to both segment distinct events in time and describe them in a series of coherent sentences. This problem is a generalization of dense image region captioning and has many practical applications, such as generating textual summaries for the visually impaired, or detecting and describing important events in surveillance footage.
1017
1018
1019<span class=""description-source"">Source: [Joint Event Detection and Description in Continuous Video Streams ](https://arxiv.org/abs/1802.10250)</span>",computer-vision1020794365a3-4ef4-45f7-8566-099100e6c8ff,lake-ice-detection,Lake Ice Monitoring,,computer-vision1021d1f7e2a2-d406-4885-be52-6023ab525297,partial-video-copy-detection,Partial Video Copy Detection,The PVCD goal is identifying and locating if one or more segments of a long testing video have been copied (transformed) from the reference videos dataset.,computer-vision10220d6a7f57-92b3-42b2-ac58-d84727981a2e,overlapped-15-5,Overlapped 15-5,,computer-vision10233a4c8a5b-bb9c-42a1-91c3-2ac271ea659c,lidar-absolute-pose-regression,lidar absolute pose regression,,computer-vision1024b256b553-fe42-43a6-ae47-64e5fd58ed6c,class-agnostic-object-detection,Class-agnostic Object Detection,Class-agnostic object detection aims to localize objects in images without specifying their categories.,computer-vision102580d8f236-3fd8-4040-9af5-72f334278b64,unbalanced-segmentation,Unbalanced Segmentation,,computer-vision1026518fdaa7-4672-480e-ad4f-59cf3d7171dd,scene-classification,Scene Classification,"**Scene Classification** is a task in which scenes from photographs are categorically classified. Unlike object classification, which focuses on classifying prominent objects in the foreground, Scene Classification uses the layout of objects within the scene, in addition to the ambient context, for classification.1027 1028 1029<span class=""description-source"">Source: [Scene classification with Convolutional Neural Networks ](http://cs231n.stanford.edu/reports/2017/pdfs/102.pdf)</span>",computer-vision10302a46d2cf-c12f-42ff-b787-c93197d7c531,one-shot-segmentation,One-Shot Segmentation,"<span style=""color:grey; opacity: 0.6"">( Image credit: [One-Shot Learning for Semantic
1031Segmentation](https://arxiv.org/pdf/1709.03410v1.pdf) )</span>",computer-vision1032cbc8b264-5fbe-4bb4-805e-e76ad9d5235a,human-instance-segmentation,Human Instance Segmentation,"Instance segmentation is the task of detecting and delineating each distinct object of interest appearing in an image.
1033
1034Image Credit: [Deep Occlusion-Aware Instance Segmentation with Overlapping BiLayers](https://arxiv.org/abs/2103.12340)",computer-vision1035a4732122-7612-4f6c-bf42-6248fcf87cfc,wavelet-structure-similarity-loss,wavelet structure similarity loss,,computer-vision1036ec257ac0-622b-4565-9859-fa32ef0dfb17,generalized-zero-shot-learning-unseen,Generalized Zero-Shot Learning - Unseen,"The average of the normalized top-1 prediction scores of unseen classes in the generalized zero-shot learning setting, where the label of a test sample is predicted among all (seen + unseen) classes.",computer-vision1037b9074f29-65d5-44f2-85fa-9720794e2e0b,head-pose-estimation,Head Pose Estimation,"Estimating the head pose of a person is a crucial problem that has a large amount of applications such as aiding in gaze estimation, modeling attention, fitting 3D models to video and performing face alignment.
1038
1039<span style=""color:grey; opacity: 0.6"">( Image credit: [FSA-Net: Learning Fine-Grained Structure Aggregation for Head Pose
1040Estimation from a Single Image](http://openaccess.thecvf.com/content_CVPR_2019/papers/Yang_FSA-Net_Learning_Fine-Grained_Structure_Aggregation_for_Head_Pose_Estimation_From_CVPR_2019_paper.pdf) )</span>",computer-vision1041166d00c1-3212-41ef-ac18-e61f30534035,scene-aware-dialogue,Scene-Aware Dialogue,,computer-vision1042db41e40f-0c62-4b37-ab6b-b54ad3cedd0d,3d-object-super-resolution,3D Object Super-Resolution,"3D object super-resolution is the task of up-sampling 3D objects.
1043
1044<span style=""color:grey; opacity: 0.6"">( Image credit: [Multi-View Silhouette and Depth Decomposition for High Resolution 3D Object Representation](https://github.com/EdwardSmith1884/Multi-View-Silhouette-and-Depth-Decomposition-for-High-Resolution-3D-Object-Representation) )</span>",computer-vision1045415397d5-740e-4d7c-8e2e-7e5549ce968f,few-shot-point-cloud-classification,Few-Shot Point Cloud Classification,Few-Shot Learning on point cloud classification task,computer-vision104624c1c91e-5e09-45a0-a710-700de20d7c4b,referring-expression-segmentation,Referring Expression Segmentation,"The task aims at labeling the pixels of an image or video that represent an object instance referred by a linguistic expression. In particular, the referring expression (RE) must allow the identification of an individual object in a discourse or scene (the referent). REs unambiguously identify the target instance.",computer-vision10471f2886ab-6ac2-4308-b43d-c98ae7f43ac0,landslide-segmentation,Landslide segmentation,,computer-vision104819a0fcfe-da41-427b-8970-53973135eea6,pose-transfer,Pose Transfer,,computer-vision10493df60d58-8f47-4222-8ce3-430bf4c61ac2,color-image-compression-artifact-reduction,Color Image Compression Artifact Reduction,,computer-vision105031b0e2f9-d673-465a-8c69-50a58488bc66,dynamic-region-segmentation,Dynamic Region Segmentation,,computer-vision1051e80e492b-4bc6-44db-802d-c2b3fe9e80b3,multilingual-text-to-image-generation,Multilingual Text-to-Image Generation,,computer-vision10526b629472-1dd2-4009-9781-de6cb2fc974f,multi-object-discovery,Multi-object discovery,,computer-vision1053d51ec404-208a-4ff2-a848-1c0fc327bba1,symmetry-detection,Symmetry Detection,,computer-vision1054c8c879f9-eada-413f-b9b7-72cf405e5f6d,user-constrained-thumbnail-generation,User Constrained Thumbnail Generation,"Thumbnail generation is the task of generating image thumbnails from an input image.
1055
1056<span style=""color:grey; opacity: 0.6"">( Image credit: [User Constrained Thumbnail Generation using Adaptive Convolutions](https://arxiv.org/pdf/1810.13054v3.pdf) )</span>",computer-vision1057f11b4ca3-e1ef-4a09-9b68-887f96fb7e6c,metamerism,Metamerism,,computer-vision10585cdd4f97-6266-470f-bc11-9ded1c782966,action-assessment,Action Assessment,,computer-vision1059bf571ef3-5c54-4e3b-8ea2-b8695a483958,unsupervised-image-to-image-translation,Unsupervised Image-To-Image Translation,"Unsupervised image-to-image translation is the task of doing image-to-image translation without ground truth image-to-image pairings.
1060
1061<span style=""color:grey; opacity: 0.6"">( Image credit: [Unpaired Image-to-Image Translation
1062using Cycle-Consistent Adversarial Networks](https://arxiv.org/pdf/1703.10593v6.pdf) )</span>",computer-vision1063ca3bc6e8-c805-4c66-859d-42b4a6f2faa2,video-object-detection,Video Object Detection,"Video object detection is the task of detecting objects from a video as opposed to images.
1064
1065<span style=""color:grey; opacity: 0.6"">( Image credit: [Learning Motion Priors for Efficient Video Object Detection](https://arxiv.org/pdf/1911.05253v1.pdf) )</span>",computer-vision1066693610f9-8a38-4f22-a2ba-2bab83bce50f,gaze-estimation,Gaze Estimation,"**Gaze Estimation** is a task to predict where a person is looking at given the person’s full face. The task contains two directions: 3-D gaze vector and 2-D gaze position estimation. 3-D gaze vector estimation is to predict the gaze vector, which is usually used in the automotive safety. 2-D gaze position estimation is to predict the horizontal and vertical coordinates on a 2-D screen, which allows utilizing gaze point to control a cursor for human-machine interaction.1067 1068 1069<span class=""description-source"">Source: [A Generalized and Robust Method Towards Practical Gaze Estimation on Smart Phone ](https://arxiv.org/abs/1910.07331)</span>",computer-vision107088d8eda0-2192-4312-a338-93d1658fe4ef,action-classification,Action Classification,Image source: [The Kinetics Human Action Video Dataset](https://arxiv.org/pdf/1705.06950.pdf),computer-vision10718bcc4d8e-5cdb-4fa0-85ba-9715e6d2f56e,flare-removal,Flare Removal,"When a camera is pointed at a strong light source, the resulting photograph may contain lens flare artifacts. Flares appear in a wide variety of patterns (halos, streaks, color bleeding, haze, etc.) and this diversity in appearance makes flare removal challenging.",computer-vision10727127a14b-f31d-4199-8799-36cf364f0143,motion-magnification,Motion Magnification,"Motion magnification is a technique that acts like a microscope for visual motion. It can amplify subtle motions in a video sequence, allowing for visualization of deformations that would otherwise be invisible. To achieve motion magnification, we need to accurately measure visual motions, and group the pixels to be modified.
1073
1074There are different approaches to motion magnification, such as Lagrangian and Eulerian methods. Lagrangian methods track the trajectories of moving objects and exaggerate them, while Eulerian methods manipulate the motions at fixed positions. Eulerian methods can be further divided into linear and phase-based methods. Linear methods apply a temporal bandpass filter to boost the linear term of a Taylor series expansion of the displacement function, while phase-based methods use complex wavelet transforms to manipulate the phase of the signal.
1075
1076Motion magnification has various applications, such as measuring the human pulse, visualizing the heat plume of candles, revealing the oscillations of a wine glass, and detecting structural defects.",computer-vision1077bf565d52-18f1-4e13-8fe3-2c94aee3b126,color-constancy,Color Constancy,"**Color Constancy** is the ability of the human vision system to perceive the colors of the objects in the scene largely invariant to the color of the light source. The task of computational Color Constancy is to estimate the scene illumination and then perform the chromatic adaptation in order to remove the influence of the illumination color on the colors of the objects in the scene.1078 1079 1080<span class=""description-source"">Source: [CroP: Color Constancy Benchmark Dataset Generator ](https://arxiv.org/abs/1903.12581)</span>",computer-vision10811ac214af-8ce9-43ec-92bc-90a3a3d86eb3,visual-relationship-detection,Visual Relationship Detection,"Visual relationship detection (VRD) is one newly developed computer vision task aiming to recognize relations or interactions between objects in an image. It is a further learning task after object recognition and is essential for fully understanding images, even the visual world.",computer-vision10821a1e1b8d-6ae5-49ee-b355-a619f58a4169,unsupervised-facial-landmark-detection,Unsupervised Facial Landmark Detection,"Facial landmark detection in the unsupervised setting popularized by [1]. The evaluation occurs in two stages:
1083(1) Embeddings are first learned in an unsupervised manner (i.e. without labels);
1084(2) A simple regressor is trained to regress landmarks from the unsupervised embedding.
1085
1086[1] Thewlis, James, Hakan Bilen, and Andrea Vedaldi. ""Unsupervised learning of object landmarks by factorized spatial embeddings."" Proceedings of the IEEE International Conference on Computer Vision. 2017.
1087
1088<span style=""color:grey; opacity: 0.6"">( Image credit: [Unsupervised learning of object landmarks by factorized spatial embeddings](https://www.robots.ox.ac.uk/~vedaldi/assets/pubs/thewlis17unsupervised.pdf) )</span>",computer-vision108915225091-c6d9-495c-ba5b-b538f13a6fef,generative-visual-question-answering,Generative Visual Question Answering,Generating answers in free form to questions posed about images.,computer-vision1090503d1a04-1f41-430f-89e2-2edecd83a8b1,unsupervised-face-recognition,Unsupervised face recognition,,computer-vision1091c066f268-a5fc-44f1-8d42-058e882e8af8,text-line-extraction,Text-Line Extraction,,computer-vision10928b3ea835-3e05-4e9a-88df-1ed4097ae5b9,unsupervised-anomaly-detection-with-specified-4,Unsupervised Anomaly Detection with Specified Settings -- 20% anomaly,,computer-vision1093af8c4e03-4e28-463b-81df-55a377f7f604,supervised-anomaly-detection,Supervised Anomaly Detection,"In the training set, the amount of abnormal samples is limited and significant fewer than normal samples, producing data distributions that lead to a naturally imbalanced learning problem.",computer-vision10946f890032-8afd-4727-aef0-7438e8bf1d3b,personalized-segmentation,Personalized Segmentation,"Given a one-shot image with a reference mask, the models are required to segment the indicated target object in any other images.",computer-vision1095fccfea97-232f-457e-abf0-654570f8d3b5,3d-human-pose-estimation,3D Human Pose Estimation,"**3D Human Pose Estimation** is a computer vision task that involves estimating the 3D positions and orientations of body joints and bones from 2D images or videos. The goal is to reconstruct the 3D pose of a person in real-time, which can be used in a variety of applications, such as virtual reality, human-computer interaction, and motion analysis.",computer-vision1096f6386616-661c-4524-8806-f2fd84802cba,video-background-subtraction,Video Background Subtraction,,computer-vision109715cc180b-1350-4a07-a7f1-9a49748d9b67,direct-transfer-person-re-identification,Direct Transfer Person Re-identification,,computer-vision1098aa706edd-b3eb-45f3-8344-8d43fe8eb329,chat-based-image-retrieval,Chat-based Image Retrieval,"ChatIR: a Chat-based Image Retrieval system that engages in a conversation with the user to elicit information, in addition to an initial query, in order to clarify the user’s search intent.",computer-vision1099b01e9cb4-3884-4d18-bcb0-284a9596cf6a,3d-human-pose-and-shape-estimation,3D human pose and shape estimation,Estimate 3D human pose and shape (e.g. SMPL) from images,computer-vision1100d0a3ebcf-c552-4ce6-bb96-6cf9008a81d8,conditional-text-to-image-synthesis,Conditional Text-to-Image Synthesis,"Introducing extra conditions based on the text-to-image generation process, similar to the paradigm of ControlNet.",computer-vision110181ec785a-6cf3-4324-9c5a-8f24eb5c1200,human-action-generation,Human action generation,"Yan et al. (2019) CSGN:
1102
1103""When the dancer is stepping, jumping and spinning on the
1104stage, attentions of all audiences are attracted by the streamof the fluent and graceful movements. Building a model that is capable of dancing is as fascinating a task as appreciating the performance itself. In this paper, we aim to generate long-duration human actions represented as skeleton sequences, e.g. those that cover the entirety of a dance, with hundreds of moves and countless possible combinations.""
1105
1106
1107<span style=""color:grey; opacity: 0.6"">( Image credit: [Convolutional Sequence Generation for Skeleton-Based Action Synthesis](http://www.dahualin.org/publications/dhl19_csgn.pdf) )</span>",computer-vision1108e336a2b5-b263-457e-8d48-1a7fd4ecdd05,unity,Unity,,computer-vision11098568940e-ebc1-4a06-88d3-c7bab486a666,lip-sync,Unconstrained Lip-synchronization,"Given a video of an arbitrary person, and an arbitrary driving speech, the task is to generate a lip-synced video that matches the given speech.
1110
1111This task requires the approach to not be constrained by identity, voice, or language.",computer-vision11123418a5fa-da0e-4cb3-ae00-262ff2a22ac3,face-alignment,Face Alignment,"Face alignment is the task of identifying the geometric structure of faces in digital images, and attempting to obtain a canonical alignment of the face based on translation, scale, and rotation.
1113
1114<span style=""color:grey; opacity: 0.6"">( Image credit: [3DDFA_V2](https://github.com/cleardusk/3DDFA_V2) )</span>",computer-vision11154eebf9dd-0859-4c92-808b-e76780cbf31f,image-quality-assessment,Image Quality Assessment,,computer-vision111631c55176-061b-405b-b54d-0e2c519f1086,semi-supervised-object-detection,Semi-Supervised Object Detection,Semi-supervised object detection uses both labeled data and unlabeled data for training. It not only reduces the annotation burden for training high-performance object detectors but also further improves the object detector by using a large number of unlabeled data.,computer-vision1117579fc4db-7376-45ed-8bed-f5c06b98d534,medical-image-detection,medical image detection,,computer-vision1118757a4bf8-ec58-4146-a26d-8e1cf47fa709,image-retargeting,Image Retargeting,,computer-vision1119c50fd80e-1393-45eb-b3b4-f6d47720ee13,visual-sentiment-prediction,Visual Sentiment Prediction,,computer-vision1120fc287133-6a0b-4004-9186-15f50410f2b4,object-recognition,Object Recognition,"Object recognition is a computer vision technique for detecting + classifying objects in images or videos. Since this is a combined task of object detection plus image classification, the state-of-the-art tables are recorded for each component task [here](https://www.paperswithcode.com/task/object-detection) and [here](https://www.paperswithcode.com/task/image-classification2).
1121
1122<span style=""color:grey; opacity: 0.6"">( Image credit: [Tensorflow Object Detection API
1123](https://github.com/tensorflow/models/tree/master/research/object_detection) )</span>",computer-vision112454580acb-dc43-433f-8fd8-7b06e78f95be,rotated-mnist,Rotated MNIST,,computer-vision112551e7b2c3-7d41-4fea-8b6d-c569def29e41,partially-relevant-video-retrieval,Partially Relevant Video Retrieval,"In the Partially Relevant Video Retrieval (PRVR) task, an untrimmed video is considered to be partially relevant w.r.t. a given textual query if it contains a moment relevant to the query. PRVR aims to retrieve such partially relevant videos from a large collection of untrimmed videos.",computer-vision1126e838ace4-dfcc-40c2-b530-dd05883dafb6,single-image-blind-deblurring,Single-Image Blind Deblurring,,computer-vision112705f98391-210d-472d-a2bf-553d151ce538,face-reenactment,Face Reenactment,"**Face Reenactment** is an emerging conditional face synthesis task that aims at fulfilling two goals simultaneously: 1) transfer a source face shape to a target face; while 2) preserve the appearance and the identity of the target face.
1128
1129
1130<span class=""description-source"">Source: [One-shot Face Reenactment ](https://arxiv.org/abs/1908.03251)</span>",computer-vision1131421fefdf-c89d-413b-8749-09f0add697ef,point-cloud-pre-training,Point Cloud Pre-training,,computer-vision11328898695f-b78a-448a-b1c0-811c8b94b829,thermal-infrared-object-tracking,Thermal Infrared Object Tracking,,computer-vision11335af36ecc-a2e1-4750-b120-ab8c72e2b581,neural-stylization,Neural Stylization,,computer-vision11349436f70a-e50a-4168-bf24-f0819a2ffed3,synthetic-image-attribution,Synthetic Image Attribution,"Determine the source or origin of a generated image, such as identifying the model or tool used to create it. This information can be useful for detecting copyright infringement or for investigating digital crimes.",computer-vision11359f47e689-68b1-4466-92e0-2bbac4ba8795,active-observation-completion,Active Observation Completion,,computer-vision113655758a3b-0461-47ff-b30a-d6a4b139e91a,electron-microscopy-image-segmentation,Electron Microscopy Image Segmentation,,computer-vision1137e1c80782-5a00-4660-9139-228e8b02b7d3,scene-graph-generation,Scene Graph Generation,"A scene graph is a structured representation of an image, where nodes in a scene graph correspond to object bounding boxes with their object categories, and edges correspond to their pairwise relationships between objects. The task of **Scene Graph Generation** is to generate a visually-grounded scene graph that most accurately correlates with an image.1138 1139 1140<span class=""description-source"">Source: [Scene Graph Generation by Iterative Message Passing ](https://arxiv.org/abs/1701.02426)</span>",computer-vision1141701e39e5-dc78-4134-abe2-59aeb81225f4,depth-aleatoric-uncertainty-estimation,Depth Aleatoric Uncertainty Estimation,,computer-vision1142eb53566c-8c80-4147-b7a0-9fcd979a1cc2,multiple-affordance-detection,Multiple Affordance Detection,"Affordance detection is the task of detecting objects that are usable (or graspable) by a human.
1143
1144<span style=""color:grey; opacity: 0.6"">( Image credit: [What can I do here? Leveraging Deep 3D saliency and geometry for fast and scalable multiple affordance detection](https://github.com/eduard626/deep-interaction-tensor) )</span>",computer-vision1145906029bf-475c-403c-812d-ac9770fcbe52,handwriting-generation,Handwriting generation,The inverse of handwriting recognition. From text generate and image of handwriting (offline) of trajectory of handwriting (online).,computer-vision1146f3c129cc-436c-4d39-b620-e50d9b42ab33,emotion-recognition,Emotion Recognition,"**Emotion Recognition** is an important area of research to enable effective human-computer interaction. Human emotions can be detected using speech signal, facial expressions, body language, and electroencephalography (EEG). <span class=""description-source"">Source: [Using Deep Autoencoders for Facial Expression Recognition ](https://arxiv.org/abs/1801.08329)</span>",computer-vision1147a1f6f742-6399-4df3-8529-b0452a4a8c80,image-generation-from-scene-graphs,Image Generation from Scene Graphs,,computer-vision114879439ef5-952a-439a-9277-3e7d5a2ad792,image-denoising,Image Denoising,"**Image Denoising** is a computer vision task that involves removing noise from an image. Noise can be introduced into an image during acquisition or processing, and can reduce image quality and make it difficult to interpret. Image denoising techniques aim to restore an image to its original quality by reducing or removing the noise, while preserving the important features of the image.
1149
1150<span style=""color:grey; opacity: 0.6"">( Image credit: [Wide Inference Network for Image Denoising via
1151Learning Pixel-distribution Prior](https://arxiv.org/pdf/1707.05414v5.pdf) )</span>",computer-vision115282cc0c24-2df3-42d6-be23-5ef452b01488,stereoscopic-image-quality-assessment,Stereoscopic image quality assessment,,computer-vision1153bebbda6d-4a83-493c-bb85-334d25a5bc9d,drone-based-object-tracking,drone-based object tracking,drone-based object tracking,computer-vision11546936d65f-2f81-487a-ab51-ce46a0fa37ed,video-question-answering,Video Question Answering,"Video Question Answering (VideoQA) aims to answer natural language questions according to the
1155given videos. Given a video and a question in natural language, the model produces accurate answers according
1156to the content of the video.",computer-vision1157d406fc7b-defc-4519-8c03-653d2936d796,image-based-automatic-meter-reading,Image-based Automatic Meter Reading,,computer-vision1158d1e479a3-118b-42a6-9cd8-c0cd9fbd75e8,facial-inpainting,Facial Inpainting,"Facial inpainting (or face completion) is the task of generating plausible facial structures for missing pixels in a face image.
1159
1160<span style=""color:grey; opacity: 0.6"">( Image credit: [SymmFCNet](https://github.com/csxmli2016/SymmFCNet) )</span>",computer-vision1161fcf1d936-c961-466a-ba51-7aa7078639ff,vehicle-speed-estimation,Vehicle Speed Estimation,Vehicle speed estimation is the task of detecting and tracking vehicles whose real-world speeds are then estimated. The task is usually evaluated with recall and precision of the detected vehicle tracks as well as the mean or median errors of the estimated vehicle speeds.,computer-vision1162e3cffc81-ae1a-4c1a-852c-a31447e63bf8,cross-view-person-re-identification,Cross-Modal Person Re-Identification,,computer-vision11632ec7f17e-6e8d-4160-b4cd-e5ca4fa021f1,laminar-turbulent-flow-localisation,Laminar-Turbulent Flow Localisation,It is a segmentation task on thermographic measurement images in order to separate laminar and turbulent flow regions on flight body parts.,computer-vision1164d18a71d3-f0cb-40b8-9bb7-9d18f2a0345b,geometrical-view,Geometrical View,,computer-vision1165fc878640-d7da-42c4-a91d-3574333a8ed9,covid-19-image-segmentation,COVID-19 Image Segmentation,,computer-vision11660b6d3163-1834-4a54-a7ef-c014cc0cd908,motion-retargeting,motion retargeting,,computer-vision11671c498b6b-d034-4765-94ea-5752f41ff6d1,neural-radiance-caching,Neural Radiance Caching,"Involves the task of predicting photorealistic pixel colors from feature buffers.
1168
1169Image source: [Instant Neural Graphics Primitives with a Multiresolution Hash Encoding](https://arxiv.org/pdf/2201.05989v1.pdf)",computer-vision1170c6ac3350-5f6d-4343-8de1-17b0dff002b6,facial-expression-recognition,Facial Expression Recognition (FER),"**Facial Expression Recognition (FER)** is a computer vision task aimed at identifying and categorizing emotional expressions depicted on a human face. The goal is to automate the process of determining emotions in real-time, by analyzing the various features of a face such as eyebrows, eyes, mouth, and other features, and mapping them to a set of emotions such as anger, fear, surprise, sadness and happiness.
1171
1172<span style=""color:grey; opacity: 0.6"">( Image credit: [DeXpression](https://arxiv.org/pdf/1509.05371v2.pdf) )</span>",computer-vision11736aaf6a93-4c50-4a29-a829-3782733b345a,camera-shot-segmentation,Camera shot segmentation,"Camera shot temporal segmentation consists in classifying each video frame according to the type of camera used to record said frame. This task is introduced with the SoccerNet-v2 dataset, where 13 camera classes are considered (main camera, behind the goal, corner camera, etc.).",computer-vision1174b049dfec-ab6c-47ac-b35b-8f9c96767e34,local-color-enhancement,Local Color Enhancement,"Enhancement techniques for improving the contrast between lesion and background skin on dermatological macro-images are limited in the literature. To fill this gap, a modified sigmoid transform is applied in the HSV color space. The crossover point in the modified sigmoid transform that divides the macro-image into lesion and background is predicted using a modified EfficientNet regressor to exclude manual intervention and subjectivity.",computer-vision11751bd2c36d-0c85-419a-aeca-d7aa9172e613,video-reconstruction,Video Reconstruction,"<span class=""description-source"">Source: [Deep-SloMo](https://github.com/avinashpaliwal/Deep-SloMo)</span>",computer-vision117647294fca-e6cd-4364-940f-688b2734af75,vgsi,VGSI,"Given a textual goal and multiple images representing candidate events, a model must choose one image which constitutes a reason- able step towards the given goal.
1177A model should correctly recognize not only the specific action illustrated in an image (e.g., “turning on the oven”), but also the intent of the action (“baking fish”).",computer-vision11782be217ea-83fa-41ce-a8ed-a457699cb48b,monocular-cross-view-road-scene-parsing,Monocular Cross-View Road Scene Parsing(Vehicle),,computer-vision1179495975f8-389a-49b2-9069-ff0318cdb585,video-instance-segmentation,Video Instance Segmentation,"The goal of video instance segmentation is simultaneous detection, segmentation and tracking of instances in videos. In words, it is the first time that the image instance segmentation problem is extended to the video domain.
1180
1181To facilitate research on this new task, a large-scale benchmark called YouTube-VIS, which consists of 2,883 high-resolution YouTube videos, a 40-category label set and 131k high-quality instance masks is built.",computer-vision1182859bf4db-247b-4808-9181-246fe8fe0745,camera-relocalization,Camera Relocalization,"""Camera relocalization, or image-based localization is a fundamental problem in robotics and computer vision. It refers to the process of determining camera pose from the visual scene representation and it is essential for many applications such as navigation of autonomous vehicles, structure from motion (SfM), augmented reality (AR) and simultaneous localization and mapping (SLAM)."" ([Source](https://paperswithcode.com/paper/camera-relocalization-by-computing-pairwise))",computer-vision1183c52d70b9-433b-4e8e-85c6-c5bc8c04fcd2,denoising,Denoising,"**Denoising** is a task in image processing and computer vision that aims to remove or reduce noise from an image. Noise can be introduced into an image due to various reasons, such as camera sensor limitations, lighting conditions, and compression artifacts. The goal of denoising is to recover the original image, which is considered to be noise-free, from a noisy observation.
1184
1185<span style=""color:grey; opacity: 0.6"">( Image credit: [Beyond a Gaussian Denoiser](https://arxiv.org/pdf/1608.03981v1.pdf) )</span>",computer-vision11868e818f1e-2844-4f21-9d72-3f4c898b49bf,cross-domain-iris-presentation-attack,Cross-Domain Iris Presentation Attack Detection,,computer-vision1187c692bea2-e921-4e4e-ad17-55302cab8ca9,few-shot-image-segmentation,Few-Shot Semantic Segmentation,Few-shot semantic segmentation (FSS) learns to segment target objects in query image given few pixel-wise annotated support image.,computer-vision118808c4cf37-6ad9-4513-8565-adabdd43613e,lake-detection,Lake Detection,,computer-vision1189c0c28ec2-db8c-46a8-b4c2-2ccae9bb5c88,video-interlacing,Video Interlacing,,computer-vision1190d7cc855a-0f7b-4d4e-9086-c696057d7d58,6d-pose-estimation-using-rgbd,6D Pose Estimation using RGBD,Image: [Zeng et al](https://arxiv.org/pdf/1609.09475v3.pdf),computer-vision11917c4179e9-1417-4e43-bebe-1f727125b8ea,face-model,Face Model,,computer-vision11928887db77-d0d5-468d-8206-a74c470c83ef,image-similarity-search,Image Similarity Search,Image credit: [The 2021 Image Similarity Dataset and Challenge](https://paperswithcode.com/paper/the-2021-image-similarity-dataset-and),computer-vision11936e8c8429-a01d-400a-b2c8-81df31bbe511,self-supervised-person-re-identification,Self-Supervised Person Re-Identification,"Currently, self-supervised representation learning is mainly tested on image classification tasks, which is not insufficient to verify its effectiveness. It should also be tested in the visual matching task, and pedestrian re-recognition is just such an appropriate task.",computer-vision119406be337f-7fd9-402b-8ee4-267c8c232f39,online-clustering,Online Clustering,"Models that learn to label each image (i.e. cluster the dataset into its ground truth classes) without seeing the ground truth labels. Under the online scenario, data is in the form of streams, i.e., the whole dataset could not be accessed at the same time and the model should be able to make cluster assignments for new data without accessing the former data.
1195
1196Image Credit: [Online Clustering by Penalized Weighted GMM](https://arxiv.org/pdf/1902.02544v1.pdf)",computer-vision1197782812ee-db82-4308-b1f9-df2a5e333ca7,human-object-interaction-concept-discovery,Human-Object Interaction Concept Discovery,"Discovering the reasonable HOI concepts/categories from known categories and their instances. Actually, it is also a matrix (verb-object matrix) complementation problem.",computer-vision11985c50dd36-66a5-4553-96ea-4add826eecd1,video-boundary-captioning,Video Boundary Captioning,"Provided with the timestamp of a boundary inside a video, the machine is required to generate sentences describing the status change at the boundary.",computer-vision1199e9b5d209-a9f5-45b1-9ff8-effe11208038,spectral-super-resolution,Spectral Super-Resolution,,computer-vision12006408f1e9-b631-41b6-bf7e-f9839ed6a4cc,superpixel-image-classification,Superpixel Image Classification,A **Superpixel Image classification** can be classified the group of pixels that share common characteristics (like pixel intensity ) or segementize the common pixel value in to one group.,computer-vision