Digitala Vetenskapliga Arkivet

Endre søk
RefereraExporteraLink to record
Permanent link

Direct link
Referera
Referensformat
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Annet format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annet språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf
Dynamic Visual Learning
Linköpings universitet, Institutionen för systemteknik, Datorseende. Linköpings universitet, Tekniska fakulteten. Zenseact AB, Gothenburg. (Computer Vision Laboratory)ORCID-id: 0000-0003-2553-3367
2022 (engelsk)Doktoravhandling, med artikler (Annet vitenskapelig)
Abstract [en]

Autonomous robots act in a \emph{dynamic} world where both the robots and other objects may move. The surround sensing systems of said robots therefore work with dynamic input data and need to estimate both the current state of the environment as well as its dynamics. One of the key elements to obtain a high-level understanding of the environment is to track dynamic objects. This enables the system to understand what the objects are doing; predict where they will be in the future; and in the future better estimate where they are. In this thesis, I focus on input from visual cameras, images. Images have, with the advent of neural networks, become a cornerstone in sensing systems. Image-processing neural networks are optimized to perform a specific computer vision task -- such as recognizing cats and dogs -- on vast datasets of annotated examples. This is usually referred to as \emph{offline training} and given a well-designed neural network, enough high-quality data, and a suitable offline training formulation, the neural network is expected to become adept at the specific task.

This thesis starts with a study of object tracking. The tracking is based on the visual appearance of the object, achieved via discriminative convolution filters (DCFs). The first contribution of this thesis is to decompose the filter into multiple subfilters. This serves to increase the robustness during object deformations or rotations. Moreover, it provides a more fine-grained representation of the object state as the subfilters are expected to roughly track object parts. In the second contribution, a neural network is trained directly for object tracking. In order to obtain a fine-grained representation of the object state, it is represented as a segmentation. The main challenge lies in the design of a neural network able to tackle this task. While the common neural networks excel at recognizing patterns seen during offline training, they struggle to store novel patterns in order to later recognize them. To overcome this limitation, a novel appearance learning mechanism is proposed. The mechanism extends the state-of-the-art and is shown to generalize remarkably well to novel data. In the third contribution, the method is used together with a novel fusion strategy and failure detection criterion to semi-automatically annotate visual and thermal videos.

Sensing systems need not only track objects, but also detect them. The fourth contribution of this thesis strives to tackle joint detection, tracking, and segmentation of all objects from a predefined set of object classes. The challenge here lies not only in the neural network design, but also in the design of the offline training formulation. The final approach, a recurrent graph neural network, outperforms prior works that have a runtime of the same order of magnitude.

Last, this thesis studies \emph{dynamic} learning of novel visual concepts. It is observed that the learning mechanisms used for object tracking essentially learns the appearance of the tracked object. It is natural to ask whether this appearance learning could be extended beyond individual objects to entire semantic classes, enabling the system to learn new concepts based on just a few training examples. Such an ability is desirable in autonomous systems as it removes the need of manually annotating thousands of examples of each class that needs recognition. Instead, the system is trained to efficiently learn to recognize new classes. In the fifth contribution, we propose a novel learning mechanism based on Gaussian process regression. With this mechanism, our neural network outperforms the state-of-the-art and the performance gap is especially large when multiple training examples are given.

To summarize, this thesis studies and makes several contributions to learning systems that parse dynamic visuals and that dynamically learn visual appearances or concepts.

sted, utgiver, år, opplag, sider
Linköping: Linköping University Electronic Press, 2022. , s. 59
Serie
Linköping Studies in Science and Technology. Dissertations, ISSN 0345-7524 ; 2196
HSV kategori
Identifikatorer
URN: urn:nbn:se:liu:diva-181604DOI: 10.3384/9789179291488ISBN: 9789179291471 (tryckt)ISBN: 9789179291488 (digital)OAI: oai:DiVA.org:liu-181604DiVA, id: diva2:1616651
Disputas
2022-01-19, Ada Lovelace, B Building, Campus Valla, Linköping, 09:00 (engelsk)
Opponent
Veileder
Prosjekter
WASP Industrial PhD student
Forskningsfinansiär
Wallenberg AI, Autonomous Systems and Software Program (WASP)Tilgjengelig fra: 2021-12-08 Laget: 2021-12-03 Sist oppdatert: 2026-06-24bibliografisk kontrollert
Delarbeid
1. DCCO: Towards Deformable Continuous Convolution Operators for Visual Tracking
Åpne denne publikasjonen i ny fane eller vindu >>DCCO: Towards Deformable Continuous Convolution Operators for Visual Tracking
2017 (engelsk)Inngår i: Computer Analysis of Images and Patterns: 17th International Conference, CAIP 2017, Ystad, Sweden, August 22-24, 2017, Proceedings, Part I / [ed] Michael Felsberg, Anders Heyden and Norbert Krüger, Springer, 2017, Vol. 10424, s. 55-67Konferansepaper, Publicerat paper (Fagfellevurdert)
Abstract [en]

Discriminative Correlation Filter (DCF) based methods have shown competitive performance on tracking benchmarks in recent years. Generally, DCF based trackers learn a rigid appearance model of the target. However, this reliance on a single rigid appearance model is insufficient in situations where the target undergoes non-rigid transformations. In this paper, we propose a unified formulation for learning a deformable convolution filter. In our framework, the deformable filter is represented as a linear combination of sub-filters. Both the sub-filter coefficients and their relative locations are inferred jointly in our formulation. Experiments are performed on three challenging tracking benchmarks: OTB-2015, TempleColor and VOT2016. Our approach improves the baseline method, leading to performance comparable to state-of-the-art.

sted, utgiver, år, opplag, sider
Springer, 2017
Serie
Lecture Notes in Computer Science, ISSN 0302-9743, E-ISSN 1611-3349 ; 10424
HSV kategori
Identifikatorer
urn:nbn:se:liu:diva-145373 (URN)10.1007/978-3-319-64689-3_5 (DOI)000432085900005 ()9783319646886 (ISBN)9783319646893 (ISBN)
Konferanse
17th International Conference, CAIP 2017, Ystad, Sweden, August 22-24, 2017, Proceedings, Part I
Merknad

Funding agencies: SSF (SymbiCloud); VR (EMC2) [2016-05543]; SNIC; WASP; Nvidia

Tilgjengelig fra: 2018-02-26 Laget: 2018-02-26 Sist oppdatert: 2025-02-01bibliografisk kontrollert
2. A generative appearance model for end-to-end video object segmentation
Åpne denne publikasjonen i ny fane eller vindu >>A generative appearance model for end-to-end video object segmentation
Vise andre…
2019 (engelsk)Inngår i: 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Institute of Electrical and Electronics Engineers (IEEE), 2019, s. 8945-8954Konferansepaper, Publicerat paper (Fagfellevurdert)
Abstract [en]

One of the fundamental challenges in video object segmentation is to find an effective representation of the target and background appearance. The best performing approaches resort to extensive fine-tuning of a convolutional neural network for this purpose. Besides being prohibitively expensive, this strategy cannot be truly trained end-to-end since the online fine-tuning procedure is not integrated into the offline training of the network. To address these issues, we propose a network architecture that learns a powerful representation of the target and background appearance in a single forward pass. The introduced appearance module learns a probabilistic generative model of target and background feature distributions. Given a new image, it predicts the posterior class probabilities, providing a highly discriminative cue, which is processed in later network modules. Both the learning and prediction stages of our appearance module are fully differentiable, enabling true end-to-end training of the entire segmentation pipeline. Comprehensive experiments demonstrate the effectiveness of the proposed approach on three video object segmentation benchmarks. We close the gap to approaches based on online fine-tuning on DAVIS17, while operating at 15 FPS on a single GPU. Furthermore, our method outperforms all published approaches on the large-scale YouTube-VOS dataset.

sted, utgiver, år, opplag, sider
Institute of Electrical and Electronics Engineers (IEEE), 2019
Serie
Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition, CVPR, IEEE Conference on Computer Vision and Pattern Recognition, ISSN 1063-6919, E-ISSN 2575-7075
Emneord
Segmentation; Grouping and Shape; Motion and Tracking
HSV kategori
Identifikatorer
urn:nbn:se:liu:diva-161037 (URN)10.1109/CVPR.2019.00916 (DOI)000542649302058 ()9781728132938 (ISBN)9781728132945 (ISBN)
Konferanse
IEEE Conference on Computer Vision and Pattern Recognition. 2019, Long Beach, CA, USA, USA, 15-20 June 2019
Forskningsfinansiär
Wallenberg AI, Autonomous Systems and Software Program (WASP)Swedish Foundation for Strategic ResearchSwedish Research Council
Tilgjengelig fra: 2019-10-17 Laget: 2019-10-17 Sist oppdatert: 2026-01-23bibliografisk kontrollert
3. Semi-automatic Annotation of Objects in Visual-Thermal Video
Åpne denne publikasjonen i ny fane eller vindu >>Semi-automatic Annotation of Objects in Visual-Thermal Video
Vise andre…
2019 (engelsk)Inngår i: 2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW), Institute of Electrical and Electronics Engineers (IEEE), 2019Konferansepaper, Publicerat paper (Fagfellevurdert)
Abstract [en]

Deep learning requires large amounts of annotated data. Manual annotation of objects in video is, regardless of annotation type, a tedious and time-consuming process. In particular, for scarcely used image modalities human annotationis hard to justify. In such cases, semi-automatic annotation provides an acceptable option.

In this work, a recursive, semi-automatic annotation method for video is presented. The proposed method utilizesa state-of-the-art video object segmentation method to propose initial annotations for all frames in a video based on only a few manual object segmentations. In the case of a multi-modal dataset, the multi-modality is exploited to refine the proposed annotations even further. The final tentative annotations are presented to the user for manual correction.

The method is evaluated on a subset of the RGBT-234 visual-thermal dataset reducing the workload for a human annotator with approximately 78% compared to full manual annotation. Utilizing the proposed pipeline, sequences are annotated for the VOT-RGBT 2019 challenge.

sted, utgiver, år, opplag, sider
Institute of Electrical and Electronics Engineers (IEEE), 2019
Serie
IEEE/CVF International Conference on Computer Vision Workshop (ICCVW), ISSN 2473-9936, E-ISSN 2473-9944
HSV kategori
Identifikatorer
urn:nbn:se:liu:diva-161076 (URN)10.1109/ICCVW.2019.00277 (DOI)000554591602039 ()2-s2.0-85082471785 (Scopus ID)978-1-7281-5023-9 (ISBN)978-1-7281-5024-6 (ISBN)
Konferanse
IEEE International Conference on Computer Vision Workshop (ICCVW)
Forskningsfinansiär
Swedish Research Council, 2013-5703Swedish Foundation for Strategic ResearchWallenberg AI, Autonomous Systems and Software Program (WASP)Vinnova, VS1810-Q
Merknad

Funding agencies: Swedish Research CouncilSwedish Research Council [2013-5703]; project ELLIIT (the Strategic Area for ICT research - Swedish Government); Wallenberg AI, Autonomous Systems and Software Program (WASP); Visual Sweden project ndimensional Modelling [VS1810-Q]

Tilgjengelig fra: 2019-10-21 Laget: 2019-10-21 Sist oppdatert: 2026-02-12
4. Video Instance Segmentation with Recurrent Graph Neural Networks
Åpne denne publikasjonen i ny fane eller vindu >>Video Instance Segmentation with Recurrent Graph Neural Networks
2021 (engelsk)Inngår i: Pattern Recognition: 43rd DAGM German Conference, DAGM GCPR 2021, Bonn, Germany, September 28 – October 1, 2021, Proceedings. / [ed] Bauckhage C., Gall J., Schwing A., Springer, 2021, s. 206-221Konferansepaper, Publicerat paper (Fagfellevurdert)
Abstract [en]

Video instance segmentation is one of the core problems in computer vision. Formulating a purely learning-based method, which models the generic track management required to solve the video instance segmentation task, is a highly challenging problem. In this work, we propose a novel learning framework where the entire video instance segmentation problem is modeled jointly. To this end, we design a graph neural network that in each frame jointly processes all detections and a memory of previously seen tracks. Past information is considered and processed via a recurrent connection. We demonstrate the effectiveness of the proposed approach in comprehensive experiments. Our approach, operating at over 25 FPS, outperforms previous video real-time methods. We further conduct detailed ablative experiments that validate the different aspects of our approach.

sted, utgiver, år, opplag, sider
Springer, 2021
Serie
Lecture Notes in Computer Science, ISSN 0302-9743, E-ISSN 1611-3349 ; 13024
HSV kategori
Identifikatorer
urn:nbn:se:liu:diva-183945 (URN)10.1007/978-3-030-92659-5_13 (DOI)001500565200013 ()2-s2.0-85124252424 (Scopus ID)978-3-030-92658-8 (ISBN)978-3-030-92659-5 (ISBN)
Konferanse
43rd DAGM German Conference, DAGM GCPR 2021, Bonn, Germany, September 28 – October 1, 2021
Tilgjengelig fra: 2022-03-28 Laget: 2022-03-28 Sist oppdatert: 2026-06-24bibliografisk kontrollert

Open Access i DiVA

fulltext(4093 kB)960 nedlastinger
Filinformasjon
Fil FULLTEXT01.pdfFilstørrelse 4093 kBChecksum SHA-512
7a9ab8c8b66f0a16bc3ad37710a3de43bf6a238b94d8e4cd8bb685ffb6fb4bb71f59eccc3537cd6162b6b8ac00499d46c81cf5a9759877f573cf747ef76c2a41
Type fulltextMimetype application/pdf
Bestill online >>

Andre lenker

Forlagets fulltekst

Søk i DiVA

Av forfatter/redaktør
Johnander, Joakim
Av organisasjonen

Søk utenfor DiVA

GoogleGoogle Scholar
Totalt: 966 nedlastinger
Antall nedlastinger er summen av alle nedlastinger av alle fulltekster. Det kan for eksempel være tidligere versjoner som er ikke lenger tilgjengelige

doi
isbn
urn-nbn

Altmetric

doi
isbn
urn-nbn
Totalt: 2645 treff
RefereraExporteraLink to record
Permanent link

Direct link
Referera
Referensformat
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Annet format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annet språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf