LAPE Framework: Detecting Alzheimer's Disease by Fusing Speech Features with Large Language Models

LAPE anchors on LLMs to textualize paralinguistic cues like pauses, achieving SOTA Alzheimer's speech detection.
LAPE (LLM-Anchored Paralinguistic Enrichment) is a multimodal framework for early Alzheimer's speech detection. Its core innovation converts paralinguistic acoustic features — such as pauses and word prolongations — into text tokens that LLMs can process, then uses lexico-prosodic alignment and a NormGate dynamic fusion mechanism to let LLMs perceive acoustic cognitive decline signals alongside linguistic content. The framework achieves state-of-the-art performance across four evaluation settings on the ADReSS and ADReSSo benchmarks, offering a new paradigm for non-invasive, scalable cognitive screening.
Introduction: Speech as a New Window for Early Cognitive Screening
Early detection of Alzheimer's Disease (AD) has long been a major challenge in medicine. Traditional diagnostic approaches typically rely on invasive procedures or specialized cognitive assessments — costly, high-barrier, and difficult to scale. In recent years, automatic speech-based detection methods have emerged as a non-invasive, scalable alternative for early cognitive screening.
Why can speech reveal cognitive state? Research shows that Alzheimer's Disease affects not only patients' lexical-semantic organization (i.e., word choice and semantic expression), but also their speech production patterns — such as abnormal pauses and word prolongations. These subtle linguistic features turn out to be sensitive signals of cognitive decline.

However, existing methods have a notable shortcoming: they have not truly integrated these paralinguistic cues with linguistic content in a cohesive way. A paper published on arXiv proposes a framework called LAPE (LLM-Anchored Paralinguistic Enrichment), designed precisely to bridge this gap.
LAPE's Core Idea: Using LLMs as an Anchor for Paralinguistic Fusion
LAPE's central idea can be summarized as: using linguistic representations extracted by a large language model (LLM) as an anchor, then weaving paralinguistic cues into them. Rather than simply concatenating text and speech features, LAPE enables the LLM to simultaneously understand lexical content while "perceiving" acoustic-level changes such as pauses and prolongations.
To achieve this, the research team designed three synergistic innovations that form the technical backbone of the LAPE framework.
Prosodic Event Textualization: Teaching Language Models to Read Acoustic Signals
The first innovation is prosodic event textualization. Its elegance lies in encoding prosodic events — such as pauses and word prolongations — as explicit markers, using a duration-aware bounded repetition mechanism to represent the intensity of these events.
This approach transforms pause and prolongation information, which originally belongs to the acoustic domain, into text tokens that LLMs can directly process, enabling the model to jointly model prosody and lexical content. It's an inspiring cross-modal translation strategy — converting hard-to-quantify speech behaviors into the LLM's "native language."
Lexico-Prosodic Unitization and Chunking: Precisely Aligning Two Modalities
The second innovation is lexico-prosodic unitization and chunking. To preserve the identity and magnitude of events across both text and speech modalities, this method pools only contiguous word units.
This design prevents the loss of critical prosodic information during feature fusion. By aggregating only contiguous units, the model can precisely align lexical content with its corresponding prosodic features, ensuring that the exact location and degree of pauses or prolongations are not obscured.
Text-Anchored Paralinguistic Fusion and the NormGate Mechanism
The third innovation is the fusion hub of the entire framework: text-anchored paralinguistic fusion. It integrates both local-level and utterance-level speech features, and introduces a dynamic regulation mechanism called NormGate.
NormGate normalizes speech features and dynamically scales them relative to the text features. This means speech features never overshadow the text; instead, they are always adjusted relative to the text representations, resulting in a more robust and balanced multimodal fusion.
Experimental Validation: SOTA Performance Across Four Evaluation Settings
To validate LAPE's effectiveness, the research team conducted comprehensive evaluations on two widely recognized benchmark datasets — ADReSS and ADReSSo — both authoritative standards in the field of AD speech detection.

For evaluation methodology, the team employed rigorous participant-level cross-validation and leave-one-subject-out strategies. These approaches effectively prevent data leakage and ensure reliable model generalization — particularly critical in medical settings where the same subject's data must not appear in both training and test sets.
The results are impressive: LAPE achieves state-of-the-art performance across all four primary evaluation settings. This strongly validates the effectiveness and reliability of the approach: anchoring on LLMs and fusing paralinguistic cues.
Research Significance and Future Outlook
LAPE's value lies not only in setting new performance benchmarks, but in offering a new paradigm for fusing linguistic content with paralinguistic features. Traditional methods typically treat text and speech as two separate processing pipelines, whereas LAPE — through prosodic textualization and the NormGate mechanism — makes the LLM a true hub for cross-modal understanding.
From an application standpoint, speech-based Alzheimer's detection offers natural scalability advantages: a brief natural conversation recording may be sufficient for early cognitive screening. This has significant real-world implications given the enormous potential screening demand in aging societies.
Notably, the research team has indicated that the code will be publicly released upon paper acceptance, opening the door for reproduction and follow-up research. Of course, as a cutting-edge study, LAPE still has some distance to travel before clinical deployment — real-world speech data tends to be far noisier and more diverse, and model performance across cross-lingual and cross-cultural settings warrants further validation.
Regardless, LAPE demonstrates the tremendous potential of large language models for cross-modal fusion in healthcare, providing a highly valuable technical framework for speech-driven disease detection research.
Related articles

Insufficient Source Material to Generate a Valid Article
The provided source material is a single unrelated tweet with no AI or tech relevance — insufficient to support a complete, valid technical article.

Insufficient Source Material to Generate a Valid AI/Tech Article
This source material is a tweet about the ages of Underworld members — unrelated to AI or tech, and insufficient to support a full article.

Insufficient Material: Unable to Generate a Valid AI/Tech Article
The provided material is a condolence tweet about a San Diego mosque attack — unrelated to AI/tech and too limited to generate a valid technical article.