Self-supervised pretrained models exhibit competitive performance in automatic speech recognition (ASR) on finetuning, even with limited in-domain supervised data. However, popular pretrained models are not suitable for streaming ASR because they are train ...
Institute of Electrical and Electronics Engineers2025
Automatic speech recognition (ASR) systems are well known to perform poorly on dysarthric speech. Previous works have addressed this by speaking rate modification to reduce the mismatch with typical speech. Unfortunately, these approaches rely on transcrib ...
Institute of Electrical and Electronics Engineers Inc.2025
Faithful human performance capture and free-view rendering from sparse RGB observations is a long-standing problem in Vision and Graphics. The main challenges are the lack of observations and the inherent ambiguities of the setting, e.g. occlusions and dep ...
Springer Science and Business Media Deutschland GmbH2025
A robust anomaly detection mechanism should possess the capability to effectively remediate anomalies, restoring them to a healthy state, while preserving essential healthy information. Despite the efficacy of existing generative models in learning the und ...
Given a ground-level query image and a geo-referenced aerial image that covers the query’s local surroundings, fine-grained cross-view localization aims to estimate the location of the ground camera inside the aerial image. Recent works have focused on dev ...
Springer Science and Business Media Deutschland GmbH2025
The digitization of 3D deformable objects remains a significant challenge in computer graphics and vision, particularly in the accurate modeling of garments. Garments exhibit complex shape variability, non-rigid deformations, and frequent self-occlusion, m ...
Photorealistic synthetic data and novel rendering techniques significantly advanced computer vision research. However, datasets focused on computer vision applications cannot be easily applied to robotics because they typically lack physics-related informa ...
Jointly estimating hand and object shape facilitates the grasping task in human-to-robot handovers. Relying on handcrafted prior knowledge about the geometric structure of the object fails when generalising to unseen objects, and depth sensors fail to dete ...
Object detection, a fundamental task in computer vision, is crucial for various intelligent edge computing applications. However, object detection algorithms are usually heavy in computation, hindering their deployments on resource-constrained edge devices ...
The field of text-to-video retrieval has advanced significantly with the evolution of language models and large-scale pre-training on generated caption-video pairs. Current methods predominantly focus on visual and event-based details, making retrieval lar ...
Institute of Electrical and Electronics Engineers Inc.2024
Autonomous driving is a revolutionary technology that has seen considerable advancements through the adoption of deep learning solutions. One of the major challenges in this field is the interaction with other road users. This interaction necessitates a "t ...
State-of-the-art results in large language models (LLMs) often rely on scale, which becomes computationally expensive. This has sparked a research agenda to reduce these models' parameter counts and computational costs without significantly impacting their ...
This dissertation on data-driven music theory is centered around curatorial practices concerning the creation, publication, and evaluation of large, expert-annotated symbolic datasets. With its primary interest in the harmony of European tonal music from i ...
In this paper, we present EdgeFace - a lightweight and efficient face recognition network inspired by the hybrid architecture of EdgeNeXt. By effectively combining the strengths of both CNN and Transformer models, and a low rank linear layer, EdgeFace achi ...
Background and Objective: Cough audio signal classification is a potentially useful tool in screening for respiratory disorders, such as COVID-19. Since it is dangerous to collect data from patients with contagious diseases, many research teams have turned ...
While sensory representations in the brain depend on context, it remains unclear how such modulations are implemented at the biophysical level, and how processing layers further in the hierarchy can extract useful features for each possible contex-tual sta ...
Federated learning allows for training deep learning models from various sources (e.g., hospitals) without sharing patient information, but only the model weights. Two central problems arise when sending the updated weights to the central node in a federat ...
Deep neural networks may easily memorize noisy labels present in real-world data, which degrades their ability to generalize. It is therefore important to track and evaluate the robustness of models against noisy label memorization. We propose a metric, ca ...
Robustness of medical image classification models is limited by its exposure to the candidate disease classes. Generalized zero shot learning (GZSL) aims at correctly predicting seen and unseen classes and most current GZSL approaches have focused on the s ...
Laser Powder Bed Fusion (LPBF) is an Additive Manufacturing (AM) process consolidating parts layer by layer, from a metallic powder bed. It allows no limitation in terms of geometry and is therefore of particular interest to various industries. Metallic LP ...