Project Name
Trustworthy and Generative Large Language Models
Project Goal
This project deals with the challenges of trustworthiness and generation difficulties in large language models.
Project Description
Large language models (LLMs) using generative pretrained transformers (GPTs) have been established as the core technology in the era of generative artificial intelligence. A variety of applications in presence of speech, audio, text, image and video have been developed and implemented to build various LLMs in our daily life. Considering the crucial role of LLMs in current and future world, this project focuses on a number of challenging issues in construction and utilization of LLMs in different domains and systems under different modalities with a broad range of applications including speech recognition/synthesis, multilingual language processing, sentiment analysis/generation, empathetic dialogue generation, multi-modal communication, to name a few. The issues of model generalization, learning efficiency, domain adaptation, diverse dialogue, trustworthy generation in LLMs are investigated. Three types of research challenges are explored in learning representation for LLMs. Firstly, the richness and rigorousness of LLM training are enhanced by incorporating a powerful generation process based on diffusion model. In particular, this project develops the diffusion models to implement an easy-first generation of natural sentences as well as a disentangled generation of semantic images. A contrastive mixture diffusion model is proposed to carry out continuous-discrete diffusion steps for text generation where contrastive objective is enforced to encourage generation of common/easy words in the early steps and specific/difficult words in the late steps. Also, the disentangled attention is presented to generate high-quality images through text-to-image LLMs. This study further explores the capability of diffusion-based LLMs to implement a non-autoregressive generation in speech recognition based on connectionist temporal classification. Secondly, the hallucination in dialogue generation is tackled by mitigating the overconfident responses by means of increasing the reliability and enhancing the trustworthiness in LLMs. To preserve the safety of using flexible LLMs, this study formulates an optimization problem for retrieval-augmented generation (RAG) where the soft prompts are learned towards factual accuracy through variational dehallucination. The mutual information between the retrieved top-k and bottom-k examples is maximized. In addition, RAG is further improved by conducting the reinforcement learning (RL) to learn a knowledge encoder, a retriever embedder and an answering agent based on a so-called reinforced LLMs. A specialized reward function is presented. Alternatively, the white-box LLM is cooperative with the black-box LLM to exploit a confidence-aware prompt refinement where the overconfident response is handled by a knowledge-based confidence calibration using black-box LLM and then the soft prompt is refined to learn an adapter over white-box LLM through RL by using the confidence-aware reward. Thirdly, this project addresses the expressiveness and robustness in speech and text representations where the empathetic speech synthesis and text generation in dialogue responses using speech and text LLMs are consolidated. Overall, this proposal deals with various fundamental issues in LLMs and develops practical applications for speech, text and image generation to build a cross-domain multi-modal empathetic conversational system.
