Speech and Audio Processing

Including speech recognition, speech synthesis, paralinguistics, environmental audio analysis, probabilistic inference from acoustic data, and signal processing for communication.

Ongoing projects

AI2PUB

Artificial Intelligence (AI) has become a powerful and pervasive technology in recent years, influencing numerous aspects of our daily lives. We encounter AI through recommendation algorithms in online stores, voice-activated smartphone assistants, or the widespread use of technologies like ChatGPT. However, its rapid growth and integration into society raise complex questions and concerns among the public. Public opinions on AI vary widely; while some people are enthusiastic about its potential to revolutionize industries and enable breakthroughs in fields such as medicine, others fear that AI could lead to undesirable outcomes, such as a loss of human control or privacy. These views are often shaped by media narratives and the competing interests of different stakeholders, which play a significant role in influencing both public opinion and policy decisions.

Our team of scientists and communication experts aims to enhance the understanding of AI technologies among the Swiss people, with a particular focus on teenagers and female students, to create a positive societal impact. Building on the foundations of our previous project, NewsOnAI, we will expand beyond traditional media such as newspapers and employ diverse methods, including artistic performances and interactive exhibitions. We will design these activities to be highly interactive, encouraging active participation and dialogue. Activities will include themed theater plays that explore AI’s impact on everyday life, exhibitions where participants can interact with AI tools, and workshops specifically designed for teenagers and female students to discuss AI’s future role in society. Feedback collected will include real-time audience reactions, structured questionnaires, and focus group discussions, which will be analyzed to continuously refine and adapt our engagement strategies.

Our primary audience includes Swiss citizens interested in cultural activities, particularly teenagers who are keen to follow new trends. Additionally, we are committed to addressing gender aspects by designing content and activities that specifically appeal to female students. We aim to inspire and empower young women to take on more prominent roles in shaping the digital world, acknowledging that they have historically been underrepresented in these fields. As societal attention shifts toward greater inclusion, our project will contribute to fostering a more balanced and equitable digital future.

While many individuals in our target groups may lack in-depth technical knowledge of AI, they often encounter new AI products, companies, and social issues through various media channels, including newspapers and science fiction movies. As a result, they may be aware of recent developments but also susceptible to misunderstandings and controversies related to technologies such as ChatGPT, Elon Musk's brain-chip startup, and other emerging AI applications. It is crucial to recognize that media portrayal significantly influences public opinion on AI, both positively and negatively. Media creators, even if they are not experts in AI, often produce content that captures public attention, which high-profile figures, including entrepreneurs, CEOs, and politicians, may leverage to advance their agendas. This can sometimes lead to skewed public perceptions, whether intentionally or unintentionally. Given this landscape, it is essential for AI scientists to collaborate with media creators, providing evidence-based insights to ensure accurate and balanced information is shared with the public. Our project fosters such collaboration, ensuring that both the potential and limitations of AI are clearly communicated. By sharing our findings through diverse media outlets, we aim to reach a broad audience, extending beyond Switzerland. Furthermore, our proactive engagement efforts will foster dynamic, two-way communication between scientists and the public, using interactive methods in exhibitions and theater plays to engage teenagers and female students specifically. Analyzing the feedback from these initiatives will provide invaluable insights into public perspectives on emerging technologies. This understanding will guide scientists in pursuing research directions that effectively address societal concerns, demonstrating the tangible benefits of our project for both scientific advancement and societal well-being. We anticipate that our efforts will have a multiplying social impact over time, promoting informed public discourse and a deeper understanding of AI technologies.

CARMEN

Challenges related to “on the move” biometrics are 1/ lower quality live biometric data, and 2/ no time to read the ePassport. Also, fully automatic biometric border control solutions, even with a stop allowed, are currently deployed only for pedestrians in controlled environment. Carmen offers biometric solutions for non-stop border control, suitable for pedestrians and vehicles, in uncontrolled environmental conditions. Travellers’ authentication is achieved in two steps: 1) the biometric data (face, iris, periocular) of travellers is securely stored in their smartphones, thanks to a DTC (Digital Traveler Credential).2) the biometric data is securely transferred from the DTC to the police infrastructure, and compared to the live biometric data collected as the traveller crosses the border control point.
To address a variety of environmental conditions, Carmen uses both NIR and RGB live biometric images and compare them to the reference images that have been acquired with one type of lighting only. To make biometrics more robust, Carmen has a multimodal approach on face iris and periocular regions. In the fraud detection, Carmen detects presentation attacks on moving travellers and searches strange travellers’ behaviour. To address small and large border control points, Carmen enables the use of fix and body-worn cameras. To address travellers in cars or lorries, Carmen detects face images in slowly moving vehicles, through the windows. Travellers in coaches are controlled by border guards walking through the coach, using dedicated portable equipment as they walk. Carmen addresses the robustness of DTC via data injection attacks. Of course, Carmen complies with the existing legal and ethical standards, and privacy is a central concern. Carmen solutions will be demonstrated in operational conditions over the UK - France border, as French and UK border authorities are part of Carmen consortium, using the infrastructure proposed by the partner Brittany Ferries.

CERTAIN

Along the whole value chain in using data for economic purposes, guidelines and tools are required to make the business of the different stakeholders successful, and the end-users confident that none of their rights are endangered. CERTAIN addresses these needs and delivers solutions for data holders, dataspaces and AI systems providers, and AI systems deployers, which are the primary actors of the data and AI value chain. They must be compliant with applicable European regulations, must reach this compliance in a timely manner, and at reasonable cost.
CERTAIN delivers guidelines and technical tools to help with compliance, to assess data quality, to measure biases in datasets, and to protect privacy. CERTAIN sets the foundation of AI certification: it translates the regulations to business terms, builds a directory of certification entities per business, develops a platform to streamline the certification process, and tools for AI system providers and certification entities so that they could respectively prepare and run a certification process.
In case of security breach, not only privacy may get compromised, but also AI models may become useless and lead to extremely damageable decisions. To make sure that AI-based products are of high quality and reliability, CERTAIN develops security tools and methods, specifically suitable for dataspaces and AI systems.
CERTAIN addresses the environmental footprint of the AI value chain. Innovative techniques are elaborated to reduce energy consumption when building and running AI systems. This is beneficial not only for the green deal but to reduce cost for AI stakeholders.
As importantly, CERTAIN considers the end-users perspective, and provides templates and guidelines that may be used by AI systems deployers to reassure end-users on the use of their private data. The project tests its results on seven operational pilots in six different business areas, considering all the actors along the AI value chain.

CHASPEEPRO

Oral verbal communication represents the main communication channel among humans. In most communication contexts, speakers must speak clearly and accurately in order to be intelligible. Intelligible speech can be disrupted in a variety of conditions of motor speech disorders (MSD). MSD in adults refers to a broad set of altered speech dimensions (articulation, speech rate, voice, prosody) in the course of several neurological diseases, which can dramatically impact patients’ communication. MSDs are due to disruption in the processes transforming a linguistic message intoarticulated speech, i.e., (a) the retrieval/encoding, contextualization and coordination of speech goals into a speech plan, (b) the preparation of motor programs with detailed neuromuscular specifications, and (c) the execution of these programs. Impairments atthese different stages have been associated with different MSDs, with apraxia of speech(AoS) associated with impairments at the first stage, i.e., the planning stage, and dysarthria associated with impairments at the programming or at the execution stage. Nonetheless, defining planning and programming stages, as well as distinguishing impairments at these two levels in terms of speech features and clinical differential diagnosis, is far from being clear-cut. This proposal builds on the successful outcomes of the Sinergia MoSpeeDi project (2017-2021, https://www.unige.ch/fapse/mospeedi/) led by the same multidisciplinary consortium. Thanks to the complementary expertise in speech and language pathology, psycholinguistics, neurology, phonetics, and speech engineering, we have collected an impressive database of MSD speech, have developed procedures sensitive enough to assess and classify mild and moderate MSD, and have obtained converging experimental evidences for the characterization of processes occurring at the planning and motor programming stages. This knowledge gained from carefully designed experiments and laboratory settings should now be expanded to speech production elicited in a more natural clinical setting. The distinction between speech planning and programming processes should also be further tackled to overcome the difficulty in defining and operationalizing processes at these two stages. With the overarching goal of understanding and modelling speech planning and programming and their related disorders, we will pursue our synergic approach based on the integration of methods and on the convergence of evidence obtained with experimentally induced speech behaviours, electrophysiological brain signals, and acoustic analyses of typical and impaired speech. Based on the results and expertise developed in the ongoing project to pinpoint speech planning and programming and to classify speakers and speech samples, in this project we propose to (a) develop assessment and classification methods applicable to realistic clinical constraints and needs, (b) build on the convergence of phonetic knowledge-based approaches and knowledge-free approaches, (c) enrich our set of acoustic descriptors in order to capture alterations at different scales of speech organization, and (d) complement acoustic-based characterization of speech planning, programming, and MSD classification with EEG signals. The outcomes of the project will rely on substantial data of disordered speech collected from over 180 French speaking participants with different types of MSDs including AoS and subtypes of dysarthria following stroke or neurodegenerative diseases. Results will be used to challenge current models of speech production which need to integrate data from MSD and will contribute to the development of speech assessment systems adapted to atypical speech and to the needs of clinical practice.

Past projects

AAMASSE

In previous years, we have focused our automatic speech recognition (ASR) research with Samsung on accents and on multi-linguality. In this year, we propose to focus on “Natural” user interfaces. By natural, we mean that the interface should function in such a way that the user should not have to behave differently from when he or she interacts with a person. Of course, there are many facets to this; however, two are pertinent: conversational/spontaneous speech and recognition exploiting natural sensors. Speech user interfaces typically rely on being able to place a microphone close to the user’s mouth. This maximizes the volume and clarity of the speech signal, whilst minimizing the effect of other noise in the vicinity. Such an interface is natural for, say, a telephone. However, many applications do not lend themselves to this type of interface. Examples include most home electronics, where the user might typically be in the center of a room, but the device is near a wall. In the case of televisions, a useful intermediate device is the remote control. Nevertheless, it is still inconvenient to hold a remote control like a telephone in order to talk to it.

ADDG2SU

Current state-of-the-art automatic speech recognition (ASR) systems commonly use hidden Markov models (HMMs), where phonemes (phones) are assumed to be the intermediate subword units and each word to be recognized is explicitly modeled as a sequence of phonemes. Thus, despite availability of sophisticated statistical modeling or machine learning techniques, to develop an ASR system one requires prior knowledge, such as lexical resources (e.g., phoneme set, lexicon) and some minimum phonetic expertise. The lexicon in ASR system contains phonetic transcription of each word. One of the key aspect in lexicon development is learning the relation between graphemes/alphabets and phonemes. Often this is done by applying statistical methods such as, decision trees, conditional random fields which invariably rely on the availability of an initial lexicon that contains good quality pronunciations. Major languages such as, English, French, German, Spanish have well developed lexical resources. However, there are minority languages, such as Scottish Gaelic, Afrikaans that do not have such well developed lexical resources. Thus, development of ASR systems have mainly focussed towards major languages. Recently, at Idiap we have developed a novel approach which, with the aid of new statistical models, learns/captures probabilistic relation between graphemes and phonemes through acoustic data. This has opened up multiple opportunities for further development and research. For instance, this approach allows the possibility to exploit both lexical and acoustic resources from one or multiple languages to develop lexical resources for another language. In addition, it allows the possibility to develop an ASR system where instead of phonemes units automatically derived from acoustic data are used as subword units. Such systems are of utmost interest to all languages for rapid development and deployment of ASR systems. The goal of the present project is to exploit the novel approach to a) develop a framework for flexible development of lexical resources for both major and minority languages, and b) develop an ASR system that overcomes the need for linguistically motivated subword units (i.e., phonemes) or prior lexical resources, while yielding state-of-the-art performance.

ADDG2SU_EXT
ADEL

The goal of the here submitted proposal is to finance a first year of research as a concrete first step towards the creation of the Center for Leadership and New Technologies (Unil, Idiap/EPFL, IMD). The Center that we aim to create long term will include AI and virtual reality among other technologies in relation to leadership. The Center will develop tools for assessing and developing leadership, conduct research with respect to new technologies related to leadership, as well as showcase our developments and empirical results for the corporate world (e.g., writing white papers, organizing symposia and conferences). Ideally, firms would turn to the Center for advice, training, and thought leadership on the topic of new technologies and leadership. IMD will be crucial in creating the link with companies and will be able to use the new technologies for their teaching and training. A first concrete project for which we ask for seed funding from the Trans4 consortium concerns the development of a collection of software modules that will be able to automatically detect leadership skills from videotaped speeches using voice and body language information. The algorithms developed will be able to automatically detect perceived leadership based on voice and video samples. We will train an algorithm to infer leadership (e.g., trustworthiness, competence as a strategic leader, competence as a transformational leader etc.) automatically based on vocal cues and body language automatically detected by the machine. We will train the algorithms with ground truth data that we will collect from a panel of evaluators (e.g., MTurk workers) on either selfpresentation videos (e.g., video CVs on YouTube) or on public speaking videos (e.g., TED Talks). Given that the quality of the algorithm depends on the quality of the training data (i.e., ground truth), we will put extra care and effort in producing this training data. The so developed software modules can then be used for leadership skill assessment and for leadership skill training and development. It can be seen as a stand‐alone outcome but at the same time it can be incorporated to the Charismometer algorithm that John and Philip have already developed and it can be added to work Daniel and Marianne have been doing in the past (on automatic extraction of nonverbal behavior from video). Basing the new development on existing work ensures that we do not start from scratch and that we can achieve the goal within one year of funding. The seed money project is thus at the same time a continuation of existing work and an important extension of it.

Don't miss a Step - Join us
Whether you want to join our team, become part of our community, support us through a donation, or explore a partnership, you’ll find all the ways to connect with us right here.