CaringAI has reported preliminary real-world evidence for CaringAI Listen, its artificial intelligence-powered telephone cognitive assessment platform, after approximately 80% of 68 system-generated reports were accepted without modification by a Certified Dementia Practitioner. Presented at the Alzheimer’s Association International Conference 2026 in London, the study evaluated telephone screenings involving adults aged 60 and older in a primary care population that was 88% Black or African American.
The findings suggest that automated telephone screening can reproduce clinician-accepted scoring and triage decisions in many cases, potentially reducing the practical burden of conducting routine cognitive assessments. However, the research remains an early feasibility and preliminary validation study, rather than evidence that CaringAI Listen can independently diagnose mild cognitive impairment, Alzheimer’s disease or another form of dementia.
That distinction matters because primary care practices face a genuine detection problem. Cognitive concerns frequently emerge during appointments that are already crowded with chronic disease management, medication reviews and preventive care. Even when a clinician recognises the need for assessment, limited appointment time, staff availability and follow-up capacity can prevent a structured cognitive screen from being completed.
CaringAI Listen attempts to move part of that workload outside the traditional consultation. The platform uses a conversational voice agent to administer cognitive assessments by telephone and generate a scored report for clinical review. Patients do not need to install an application, own a smartphone, use a patient portal or maintain a reliable internet connection, giving the model a potentially broader reach than app-dependent digital health tools.
Why does 80% clinician acceptance matter for the future of automated cognitive assessment?
The headline result is operationally relevant because automation only creates value if clinicians can trust its outputs without routinely rescoring or rewriting them. In the CaringAI study, the reviewer accepted the platform’s scoring and triage without changes in approximately four out of every five cases. Agreement across the cognitive instruments used in the assessments was described as ranging from good to excellent.
For a primary care practice, that level of acceptance could translate into fewer minutes spent administering tests, calculating scores and preparing documentation. A telephone assessment completed before or after an appointment could allow the clinician to begin with a structured report, examine areas of concern and decide whether additional evaluation is appropriate.
The remaining 20% of reports are equally important. The available findings do not explain how substantial the modifications were, whether the differences affected triage decisions, or whether errors were concentrated among certain types of patients or assessments. Minor documentation corrections would have a different safety and workflow meaning from changes that moved a patient between normal, borderline and impaired categories.
Clinician acceptance is also not identical to clinical accuracy. It measures whether an expert agreed with the system-generated scoring and triage, not whether the patient ultimately had mild cognitive impairment, Alzheimer’s disease or another condition. A platform and reviewer can agree with one another while still differing from a comprehensive diagnostic evaluation.
The study therefore supports a human-in-the-loop model more strongly than autonomous screening. CaringAI Listen may be able to perform the repetitive administration and scoring work, while clinicians retain responsibility for reviewing the report, interpreting the wider clinical picture and deciding what happens next.

How much clinical confidence can be drawn from a real-world study involving 68 reports?
The study’s real-world design is a strength because assessments were conducted through both in-clinic calls and scheduled outbound telephone calls. This is closer to the environment in which the platform would be used than a tightly controlled laboratory test involving highly selected participants and ideal technical conditions.
At the same time, 68 reviewed reports represent a limited evidence base. Small studies can establish feasibility, identify usability problems and provide early estimates of agreement, but they are rarely large enough to characterise uncommon failures or demonstrate consistent performance across multiple clinical settings.
The publicly disclosed results do not provide sensitivity, specificity, positive predictive value or negative predictive value against an independent diagnostic reference standard. They also do not disclose detailed confidence intervals, performance by impairment severity, call completion rates, test-retest reliability or the frequency with which technical problems affected an assessment.
The use of a Certified Dementia Practitioner as the independent reviewer provides clinical oversight, but a stronger validation programme would involve multiple blinded reviewers and measure agreement between those reviewers as well as agreement with the artificial intelligence system. That would help distinguish platform-related discrepancies from the ordinary variation that exists when humans interpret cognitive assessment results.
The evidence was presented as a conference poster under the title “Feasibility and Preliminary Validation of a Voice Agent-administered Cognitive Screening Protocol in Healthy and Impaired Older Adults.” A full peer-reviewed publication would allow clinicians and health systems to evaluate the methodology, participant selection, scoring rules and statistical analysis in greater detail.
Why is the predominantly Black study population clinically important but not sufficient by itself?
The demographic composition of the study is one of its most meaningful features. Black Americans have historically been underrepresented in Alzheimer’s disease research despite facing a higher documented burden of Alzheimer’s disease and related dementias, as well as greater risks of delayed or missed diagnosis.
Testing the platform in a population that was 88% Black or African American provides more relevant evidence than attempting to generalise from a largely White technology-validation cohort. It also addresses an important weakness in voice-based artificial intelligence, where performance can be affected by accents, dialects, speech patterns and the composition of the data used to train or calibrate the system.
The telephone format may reduce several participation barriers. It does not require broadband access, smartphone proficiency, transportation to a specialist centre or familiarity with a digital interface. For older adults who are comfortable answering a conventional telephone call but reluctant to navigate an application, that simplicity could improve participation.
Yet demographic representation alone does not prove equitable performance. Future studies will need to report whether scoring agreement remains consistent across age groups, education levels, dialects, socioeconomic backgrounds and levels of cognitive impairment. Performance should also be examined among people with hearing loss, speech disorders, limited English proficiency, strong regional accents and medical conditions that can influence verbal responses.
Cognitive testing is particularly sensitive to cultural, educational and linguistic context. A scalable platform must do more than include diverse patients in aggregate. It must demonstrate that its error rates do not systematically disadvantage particular subgroups or create avoidable disparities in referrals and follow-up care.
Can telephone-based screening fit into primary care without creating a new follow-up bottleneck?
CaringAI Listen targets an attractive point in the clinical workflow because the telephone remains one of the most widely accessible communication tools available to healthcare providers. A practice could theoretically schedule assessments before annual wellness visits, after a patient or family member reports memory concerns, or as part of a broader population-health programme.
The commercial value would depend on more than the accuracy of scoring. Health systems will want evidence that the platform increases completed assessments, reduces staff time, integrates with electronic health records and produces reports that clinicians can review quickly. Accountable care organisations may also examine whether earlier identification improves enrolment in care-management programmes or helps coordinate appropriate specialist referrals.
Implementation will introduce practical complications. Poor call quality, background noise, interruptions, fatigue, hearing difficulty and unfamiliarity with an automated caller could affect performance. Health systems would also need procedures for confirming patient identity, obtaining consent, handling incomplete calls and escalating concerning results.
A positive cognitive screen does not complete the diagnostic process. Patients may require medical history review, functional assessment, medication evaluation, laboratory testing, imaging or specialist consultation. If automated screening expands rapidly without sufficient follow-up capacity, it could shift the bottleneck from initial assessment to diagnostic evaluation.
False-positive results could generate anxiety and unnecessary referrals, while false negatives could provide misplaced reassurance and delay further investigation. The operational case for CaringAI Listen will therefore depend on whether it helps practices identify the right patients while keeping the number of avoidable escalations manageable.
What regulatory, privacy and governance questions will CaringAI need to answer?
The regulatory position of an artificial intelligence cognitive assessment platform depends heavily on its intended use, claims and outputs. Software that organises information for independent clinician review may be treated differently from software that produces a specific diagnostic conclusion or directs patient management.
CaringAI has described Listen as a cognitive assessment and triage platform, but the announcement did not disclose a United States Food and Drug Administration clearance or authorisation. As the company builds its evidence base, it will need to define whether the product is positioned as a workflow and decision-support tool, a cognitive assessment aid, or software intended to detect a medical condition.
That decision will shape the evidence, quality systems and post-deployment monitoring expected of the platform. If future commercial claims extend beyond administration and scoring into automated detection or diagnosis, regulatory scrutiny is likely to increase.
Voice data also carry distinctive privacy risks. Recordings and transcripts may reveal health information, identity characteristics and conversational details beyond the answers required for the assessment. Health systems will want clarity on whether calls are recorded, how long data are retained, whether information is used for model training and how access is controlled.
Artificial intelligence performance can change when models, prompts, scoring logic or speech-recognition components are updated. Buyers will therefore expect version controls, audit trails and procedures for monitoring performance after deployment. Clinician agreement measured with one version of the system cannot automatically be assumed to apply to every later version.
What evidence would move CaringAI Listen from promising feasibility to scalable clinical adoption?
CaringAI plans to extend the research through larger, multisite clinical evaluations. Those studies will be decisive because they can test whether the initial agreement holds across different practices, patient populations, telephone environments and levels of cognitive impairment.
A stronger evidence package should include prespecified thresholds, blinded comparison with accepted clinical reference standards, multiple independent reviewers and detailed reporting of false-positive and false-negative results. Sensitivity and specificity will be particularly important if the platform is used to determine which patients require further evaluation.
Researchers should also measure assessment completion, patient acceptance, time saved, clinician review burden and the proportion of generated reports requiring meaningful correction. These measures would reveal whether apparently strong technical agreement translates into a practical improvement for primary care.
Longer-term studies could examine whether telephone screening leads to earlier diagnostic evaluation, more appropriate referrals, improved care planning or better support for patients and caregivers. Without that downstream evidence, the platform may remain an efficient way to administer a test without proving that it improves the care pathway.
The CaringAI study nevertheless addresses a real and commercially significant healthcare problem. A cognitive assessment that is accessible through an ordinary phone call could reach patients who are missed by app-based tools and time-constrained clinic workflows. The next challenge is to demonstrate that convenience, inclusivity and automation can be delivered without weakening clinical accuracy, patient trust or human oversight.
