Delete search term
To content

Main navigation

School of Engineering

Service navigation

Machine Perception and Cognition

“Equipped with frontier AI expertise, we are committed to practical problem solving: developing, for example, common-sensical world models; deploying document foundation models; innovating medical & industrial AI under scarce resources; researching human-AI interaction patterns; and envisioning positive societal futures, advocating for pro-human AI."

Professor Dr Thilo Stadelmann

Expertise

  • Sensory pattern recognition with deep learning 
  • Document recognition and multimedia analysis (e.g., for industry and medicine) 
  • Biology-inspired neural system development, world models 
  • Societal effects of AI, pro-human AI design 

Real-world AI tasks often start with the detection of patterns in sensory data (like similarities and anomalies in image, video, time series, or audio data) and extend to making sense of these patterns (leading to actions like a classification, segmentation, prediction or decision). 

The MPC group spans this arc from perception to cognition, being rooted in pattern recognition research and focusing on deep neural network methodology. 

Tasks we study have diffferent learning target (e.g., detection, classification, clustering, segmentation, novelty detection, control) and corresponding practical use case (e.g., predictive maintenance for industrial assets, speaker recognition for multimedia indexing, document analysis for construction, computer vision for industrial quality control, automated machine learning in general data sciences, deep reinforcement learning for building control, video analysis for automatic media production, face recognition for biometric access control, human-AI co-learning for meaningful human-AI collaboration in high-consequence scenarios), which in turn shed light on different aspects of the learning process. 

We use this experience to create increasingly general real-world AI systems built on neural architectures. Fundamental challenges thereby lie in the systems’ robustness as well as their sample and label efficiency, and are often approached with transfer learning and domain adaptation. 

Beyond this, we take inspiration from biological learning to work on next-level AI methodology with world models, and are actively engaged at the interface of technology and society to contribute to a future worth living in with pro-human AI. 

Services

Additional Information

The group has very diverse backgrounds, which lets us complement each other’s skills and work on diverse problems with focused methodology. Get in touch to explore engaging with us as a practice partner in funded collaborative research & innovation or undertake master’s studies under our supervision. We post job offers (including for PhD students) on the institutional website if and when they become available.

Team

Projects

Publications

Other Releases

When Type Content
2023 Extended Abstract Thilo Stadelmann. KI als Chance für die angewandten Wissenschaften im Wettbewerb der Hochschulen. Workshop (“Atelier”) at the Bürgenstock-Konferenz der Schweizer Fachhochschulen und Pädagogischen Hochschulen 2023, Luzern, Schweiz, 20. Januar 2023
2022 Extended Abstract Christoph von der Malsburg, Benjamin F. Grewe, and Thilo Stadelmann. Making Sense of the Natural Environment. Proceedings of the KogWis 2022 - Understanding Minds Biannual Conference of the German Cognitive Science Society, Freiburg, Germany, September 5-7, 2022.
2022 Open Reserach Data Felix M. Schmitt-Koopmann, Elaine M. Huang, Hans-Peter Hutter, Thilo Stadelmann, and Alireza Darvishy. FormulaNet: A Benchmark Dataset for Mathematical Formula Detection. One unsolved sub-task of document analysis is mathematical formula detection (MFD). Research by ourselves and others has shown that existing MFD datasets with inline and display formula labels are small and have insufficient labeling quality. There is therefore an urgent need for datasets with better quality labeling for future research in the MFD field, as they have a high impact on the performance of the models trained on them. We present an advanced labeling pipeline and a new dataset called FormulaNet. At over 45k pages, we believe that FormulaNet is the largest MFD dataset with inline formula labels. Our dataset is intended to help address the MFD task and may enable the development of new applications, such as making mathematical formulae accessible in PDFs for visually impaired screen reader users.
2020 Open Research Data Lukas Tuggener, Yvan Putra Satyawan, Alexander Pacha, Jürgen Schmidhuber, and Thilo Stadelmann, DeepScoresV2. The DeepScoresV2 Dataset for Music Object Detection contains digitally rendered images of written sheet music, together with the corresponding ground truth to fit various types of machine learning models. A total of 151 Million different instances of music symbols, belonging to 135 different classes are annotated. The total Dataset contains 255,385 Images. For most researches, the dense version, containing 1714 of the most diverse and interesting images, is a good starting point.