Developing agents that can reliably act on our behalf is central to artificial intelligence (AI). These agents must seamlessly interact with tools, like search engines and databases, and collaborate. In this thesis, we study the abstractions, methods, and ...
Given a ground-level query image and a geo-referenced aerial image that covers the query’s local surroundings, fine-grained cross-view localization aims to estimate the location of the ground camera inside the aerial image. Recent works have focused on dev ...
Springer Science and Business Media Deutschland GmbH2025
Self-supervised pretrained models exhibit competitive performance in automatic speech recognition (ASR) on finetuning, even with limited in-domain supervised data. However, popular pretrained models are not suitable for streaming ASR because they are train ...
Institute of Electrical and Electronics Engineers2025
Continuous monitoring of various physiological conditions enables early detection of potential health issues. The monitoring offers real-time insights into patient well-being and rapid medical interventions. As a case study, in epilepsy, a prevalent neurol ...
Recently developed van der Waals magnets offer a promising platform for advancing spintronics. The weak interlayer antiferromagnetic exchange coupling in van der Waals antiferromagnets allows for unique spin dynamics and control over magnons. In this study ...
The first search for the Z boson decay to tau tau mu mu at the CERN LHC is presented, based on data collected by the CMS experiment at the LHC in proton-proton collisions at a center-of-mass energy of 13 TeV and corresponding to an integrated luminosity of ...
State-of-the-art results in large language models (LLMs) often rely on scale, which becomes computationally expensive. This has sparked a research agenda to reduce these models' parameter counts and computational costs without significantly impacting their ...
The field of text-to-video retrieval has advanced significantly with the evolution of language models and large-scale pre-training on generated caption-video pairs. Current methods predominantly focus on visual and event-based details, making retrieval lar ...
Institute of Electrical and Electronics Engineers Inc.2024
Determining the relative pose of an object between two images is pivotal to the success of generalizable object pose estimation. Existing approaches typically approximate the continuous pose representation with a large number of discrete pose hypotheses, w ...