Human–AI Teaming

The Human-AI Teaming research program explores how humans and artificial intelligence can collaborate to enhance creativity, decision-making, and support everyday tasks. Our research focuses on understanding similarities and differences between humans and machines. It also involves developing multimodal and multilingual interfaces integrating speech, non-verbal communication behaviors, and haptics, to enable intuitive human-AI interaction. Finally, we work on improving AI reliability through human feedback and enhanced interactions.

Grounded in Idiap’s expertise in multimodal interaction, cognitive systems, and human-robot collaboration, our work aims to create AI that works not just for people, but with people, to amplify human potential and societal impact. More specifically, we aim to design AI systems that improve access to knowledge, provide adaptive, personalized support across roles ranging from personal assistants to industrial and medical companions, and assist people in work and daily life through robotic technologies.

Ongoing projects

AI2PUB

Artificial Intelligence (AI) has become a powerful and pervasive technology in recent years, influencing numerous aspects of our daily lives. We encounter AI through recommendation algorithms in online stores, voice-activated smartphone assistants, or the widespread use of technologies like ChatGPT. However, its rapid growth and integration into society raise complex questions and concerns among the public. Public opinions on AI vary widely; while some people are enthusiastic about its potential to revolutionize industries and enable breakthroughs in fields such as medicine, others fear that AI could lead to undesirable outcomes, such as a loss of human control or privacy. These views are often shaped by media narratives and the competing interests of different stakeholders, which play a significant role in influencing both public opinion and policy decisions.

Our team of scientists and communication experts aims to enhance the understanding of AI technologies among the Swiss people, with a particular focus on teenagers and female students, to create a positive societal impact. Building on the foundations of our previous project, NewsOnAI, we will expand beyond traditional media such as newspapers and employ diverse methods, including artistic performances and interactive exhibitions. We will design these activities to be highly interactive, encouraging active participation and dialogue. Activities will include themed theater plays that explore AI’s impact on everyday life, exhibitions where participants can interact with AI tools, and workshops specifically designed for teenagers and female students to discuss AI’s future role in society. Feedback collected will include real-time audience reactions, structured questionnaires, and focus group discussions, which will be analyzed to continuously refine and adapt our engagement strategies.

Our primary audience includes Swiss citizens interested in cultural activities, particularly teenagers who are keen to follow new trends. Additionally, we are committed to addressing gender aspects by designing content and activities that specifically appeal to female students. We aim to inspire and empower young women to take on more prominent roles in shaping the digital world, acknowledging that they have historically been underrepresented in these fields. As societal attention shifts toward greater inclusion, our project will contribute to fostering a more balanced and equitable digital future.

While many individuals in our target groups may lack in-depth technical knowledge of AI, they often encounter new AI products, companies, and social issues through various media channels, including newspapers and science fiction movies. As a result, they may be aware of recent developments but also susceptible to misunderstandings and controversies related to technologies such as ChatGPT, Elon Musk's brain-chip startup, and other emerging AI applications. It is crucial to recognize that media portrayal significantly influences public opinion on AI, both positively and negatively. Media creators, even if they are not experts in AI, often produce content that captures public attention, which high-profile figures, including entrepreneurs, CEOs, and politicians, may leverage to advance their agendas. This can sometimes lead to skewed public perceptions, whether intentionally or unintentionally. Given this landscape, it is essential for AI scientists to collaborate with media creators, providing evidence-based insights to ensure accurate and balanced information is shared with the public. Our project fosters such collaboration, ensuring that both the potential and limitations of AI are clearly communicated. By sharing our findings through diverse media outlets, we aim to reach a broad audience, extending beyond Switzerland. Furthermore, our proactive engagement efforts will foster dynamic, two-way communication between scientists and the public, using interactive methods in exhibitions and theater plays to engage teenagers and female students specifically. Analyzing the feedback from these initiatives will provide invaluable insights into public perspectives on emerging technologies. This understanding will guide scientists in pursuing research directions that effectively address societal concerns, demonstrating the tangible benefits of our project for both scientific advancement and societal well-being. We anticipate that our efforts will have a multiplying social impact over time, promoting informed public discourse and a deeper understanding of AI technologies.

ALIGNAI

Large Language Models (LLMs) are trained on broad data, using self-supervision at scale, to complete a wide range of tasks. Wider use of LLMs has risen in recent months due to applications such as ChatGPT. Although LLMs bring many opportunities to improve our everyday lives, the impacts on humans and society have not yet been prioritized or fully understood. Given the rapid development of these tools, the risk of negative implications is significant if LLMs are not developed and deployed in a way that is aligned with human values and responds to individual needs and preferences. To mitigate any negative consequences, academia, in close collaboration with industry, needs to train the next generation of researchers to understand the complexities of the socio-technical implications surrounding the use of LLMs.
The alignAI Doctoral Network will train 17 doctoral candidates (DCs) to work in the international and highly interdisciplinary field of LLM research and development. The core of the project focuses on the alignment of LLMs with human values, identifying relevant values and methods for alignment implementation. Two principles provide a foundation for the approach. First, explainability is a key enabler for all aspects of trustworthiness, accelerating development, promoting usability, and facilitating human oversight and auditing of LLMs. Second, fairness is a key aspect of trustworthiness, facilitating access to AI applications and ensuring equal impact of AI-driven decision-making. The practical relevance of the project is ensured by three use cases in education, positive mental health, and news consumption. This approach allows us to develop specific guidelines and test prototypes and tools to promote value alignment. We follow a unique methodological approach, with DCs from social sciences and humanities “twinned” with DCs from technical disciplines for each use case (9 DCs in total), while the other 8 DCs carry out horizontal research across the use cases.

BALM

We address the controllability of large language models (LLMs) by giving them interpretable beliefs and programmable knowledge through leveraging the PI's work on understanding and improving transformer embeddings. Transformers' empirical success comes from the attention function's ability to induce graphs of relations from text. Our recent work has extended this ability to knowledge graphs, and to inducing the nodes of the graph as well, known as entity induction, with the first variational-Bayesian generalisation of the attention mechanism. This project will further develop this information-theoretic understanding of transformer embeddings and its sparsity-inducing regulariser, for learning graphs of higher-level abstract entities. The resulting Bayesian beliefs over generalised transformer embeddings of texts and graphs will give us the more interpretable, more programmable and more learnable abstract representations which are the core of this proposed project.

To leverage and extend these fundamental advances in representation learning, we will develop LLM architectures with a memory. Motivated by the success of Retrieval Augmented LLMs, our Belief Augmented Language Models (BALMs) will move knowledge extracted from training data out of large uninterpretable weight matrices into our interpretable Bayesian beliefs over large transformer embeddings. These beliefs will then be: augmented with human-editable knowledge graphs and selected new texts, refined with control objectives and multi-hop reasoning, and combined with inference of concensus beliefs and opinion summarisation. BALMs will be developed both to evaluate these beliefs and as a chat interface for specifying, accessing and editing the beliefs themselves, including the collaborative specification of shared beliefs. These fundamental advances in deep learning theory and architectures will allow us to control what an LLM says by controlling what it believes, thereby unlocking the power of AI for society.

BOVINE

In recent years, attention-based models like Transformers have radically improved the performance of natural language understanding (NLU), demonstrating the appropriateness of attention-based representation for language. In (Henderson, 2020) we show that these representation share many characteristics with those found in traditional computational linguistics (e.g. graph structure), except that they do not automatically learn multiple levels of representation nor their entities (morphemes, phrases, discourse entities, etc). Motivated by this challenge of entity induction, our recent work has discovered a very non-traditional perspective, which characterises attention-based models like Transformers as doing nonparametric Bayesian inference. Given an input text, our Nonparametric Variational Information Bottleneck (NVIB) Transformer infers distributions over nonparametric mixture distributions (Henderson and Fehr, 2023). We have even shown that pretrained Transformers can be converted into equivalent NVIB Transformers, and regularised post-training (Fehr and Henderson, 2023).

This reinterpretation of Transformers, combined with their unprecedented empirical success, leads us to postulate the hypothesis that natural language understanding is nonparametric variational Bayesian inference over mixture distributions. This claim of the adequacy of NVIB leads to two fundamental challenges which are not currently being addressed, each with an associated technological aim:

1. How can NVIB support inducing graph-structured representations at multiple levels of representation?
   Making deep learning representations interpretable.

2. How can NVIB enable controlling the information in representations?
   Making deep learning representations controllable.

For the first challenge, we will extend our previous structure processing methods (Mohammadshahi and Henderson, 2020, 2021, 2023; Miculicich and Henderson, 2022), developed for set-of-vector representations, to mixture-of-component distributions. And we will focus on unsupervised learning methods, rather than our previous supervised learning methods. To extend these models to multiple levels, we will take the approach of embedding all levels in one big mixture of non-homogeneous components, which are computed with iterative refinement. This extends our previous work on iterative graph refinement (Mohammadshahi and Henderson, 2021; Miculicich and Henderson, 2022), adding the induction of the nodes of the graph and the induction of multiple levels of representation. Learning representations which are interpretable as linguistic structures will be a testbed for the general aim of deep learning of interpretable representations.

For the second challenge, we will leverage the information theory behind NVIB to model both inferring implicit information and removing private information. We will investigate the use of KL divergence as a measure of entailment in semantic inference. We will apply the framework of Rényi differential privacy (Mironov, 2017) to provide privacy guarantees by adding noise which removes targeted information from Transformer embeddings. This method extends differential privacy to anything that can be embedded with a Transformer (especially text), with many important applications. These methods address the general aim of controlling the information in deep learning representations.

Addressing these challenges will lead to fundamental advances in machine learning, including novel deep learning architectures and fundamental insights into Transformers and their pretraining. We will do both intrinsic and extrinsic evaluations of our induced representations, expecting to show improvements on core NLP tasks, including privacy-preserving sharing of textual data. Given the current level of interest in the AI research and development community for Transformers and variational Bayesian methods, we expect the proposed research to have a profound impact on the field.

Past projects

AIML-VISIT

This proposal aims to support a visit by Dr. Damien Teney, head of the Machine Learning group at the Idiap Research Institute, to the Australian Institute for Machine Learning (AIML) in Adelaide. Dr. Teney has an extensive history of successful collaborations with several scientists from the AIML. This visit will enable rapid progress on two key projects requiring intense collaboration due to the combination of multiple skillsets and domains of expertise. More specifically, these projects aim to improve our scientific understanding of the capabilities, limitations, and reliability characteristcs of large machine learning models. These topics are increasingly relevant on scientific, societal, and economical levels due to the growing importance and adoption of machine learning and AI at large. This proposal is strongly supported by the AIML's director since it addresses topics of mutual interest. The visit will benefit the two parties through the accelerated production of high-impact scientific knowledge. It will also contribute to international visibility of Swiss research capacity. The candidate additionally plans to prepare joint grant proposals with AIML scientists, as well as to promote future opportunities for Australian scientists to visit Swiss institutions. These activities will help sustain the partnership and ensure that its benefits extend beyond the duration of the visit.

AI-SENSOR

This project aims to exploit the data generated by depth sensors in the realm of 3D computer vision. The goal is to develop and enhance state-of-the-art deep learning methods that can utilize this data to enable various applications such as dense depth map generation from structured-light sensors, novel view synthesis, and dense visual RGB-D SLAM.


Structured light sensors are one of the most commonly used depth sensors in computer vision applications. However, accurate depth map generation using structured light sensors remains a challenging task. This project proposes a solution that combines data from multi-view images to improve the accuracy of dense depth map generation using structured light sensors.


Another application of depth sensor data is in the field of novel view synthesis. Neural Radiance Fields (NeRF) have shown great potential in this domain. This project aims to explore and develop new algorithms based on NeRF that can generate novel views of an object from a given set of views. This can have significant applications in virtual reality and 3D content creation.


Dense Visual SLAM is another field that this project investigates. Visual SLAM is a popular technique in robotics and autonomous vehicles to create maps of the environment using visual data. However, traditional methods for visual SLAM often struggle to generate accurate dense depth maps in real-time. This project aims to develop a dense visual SLAM pipeline that leverages recent advances in neural rendering. By utilizing depth sensor data and an efficient neural rendering implementation, the proposed visual SLAM pipeline aims to generate more accurate and comprehensive maps of the environment while achieving higher processing speed compared to existing methods.


In conclusion, the AI-Sensor project aims to contribute to the advancement of deep learning methods in the field of 3D computer vision by exploiting the data generated by depth sensors. The project proposes novel approaches for dense depth map generation using structured light sensors, novel view synthesis based on Neural Radiance Fields, and a dense visual SLAM pipeline based on recent advances in neural rendering. These contributions have the potential to significantly impact various fields such as robotics, autonomous vehicles, virtual reality, and 3D content creation.

C-LING

Background Language is predominantly viewed as a means for information exchange in the field of AI and natural language processing. However, this overlooks the important role language plays as a vehicle for creative thought. This aspect of cognition is unique to mankind, crucial for innovation, and currently under-explored in AI. As an example, inventions such as the mobile phone, or the glass canoe, have come to life in people’s minds by means of combining known concepts in unseen ways. The field of NLP has mature methods to capture the meaning of words and the way they relate to each other. At the same time, a number of researchers have started questioning the learning capacity of these large language models and concerns regarding short-cut learning, bias and inability to generalise are growing. Creative tasks, such as the ones specified in this project, require a high level of generalisation that allows the system to cut across patterns found: across domains, time periods, as well as different languages. They are therefore a perfect testbed for measuring the generalisation power and actual level of intelligence of current systems. Objectives This project aims to investigate what aspects computational models need to perform creative cognitive tasks, from generating relatively simple novel concepts to more complex and structured ideas, across multiple domains and languages. More in particular, it aims to answer what types of structured and unstructured knowledge are needed and what models best integrate these types of knowledge. Methods Current neural language models are trained on large amounts of text data and used successfully in many NLP applications as models of language use. However, as shown in previous work, such statistical models might capture word associations well, but they are not sufficient for tasks that involve reasoning or generalisation across domains, which is the case for the creative cognitive task we study. Previous work from cognitive science used symbolic methods to model conceptual blending. However, these methods are often not scalable. We plan to use hybrid neural-symbolic methods for modelling creative thinking, which will allow us to add structure in our models and integrate insights from cognitive science. Furthermore, we will exploit differences in embeddings and knowledge across domains and languages/cultures to inform the models for novel concept creation. Expected results and impact We expect to build a variety of computational models for the range of creative tasks we defined from novel concept generation to the generation of more structured complex ideas, across domains and languages. This will lead to insights w.r.t. the limitations and capabilities of neural and neuro-symbolic approaches on such creative tasks. In addition, we will learn in how far we can leverage cognitive theory for building creative systems. Our research into creative processes will impact the field by pushing towards AI tools that are more flexible and resourceful than current technologies, which will lead to the level of innovation and diversity needed for human progress, while better exploring the wealth of data available.

CODIMAN

The Swiss economy is known for its productivity, precision and expertise. However the comparatively igh wages make it difficult for companies to stay worldwide competitive and avoid offshoring. The digitalization of work has the potential to help sustain and further improve the Swiss competitive advantages as well as strengthen their position in a global market. However, the implications of digitalization for work are not fully explored and give rise to contestation with respect to how labour in the future will look like and what competencies are needed in order to empower workers to take advantage of and participate in the digital transformation of work. This project will examine these questions by addressing factors that support the implementation and acceptance of new technologies with a particular focus on flexible automation and human-machine interactions. The project’s objectives are: a) to provide guidelines as to which digital skills are required by workers to make interactions with collaborative robotics empowering; and b) to provide insights into the necessary technical and educational tools for this empowerment to take place. Empirically, the project will employ a mixed-method approach consisting of interviews, non-participant observations and survey questionnaires as well as prototype testing. Besides guidelines for a successful implementation, the project will also deliver a platform for intuitive human-machine interaction that will provide a basis to acquire the necessary skills. To achieve these aims, the project is set up interdisciplinarily, combining expertise from computer sciences and engineering with social sciences. Building on research from those two strands, the project will make a significant and practical contribution to the advancement of a future of work that is not only more digital but also more humane.

Don't miss a Step - Join us
Whether you want to join our team, become part of our community, support us through a donation, or explore a partnership, you’ll find all the ways to connect with us right here.