Accuracy and efficiency levels differ depending on the AI used. Credit: Osaka Metropolitan University
Can AI save us from the arduous and time-consuming task of academic research collection? An international team of researchers investigated the credibility and efficiency of generative AI as an information-gathering tool in the medical field.
29 jan 2024--The research team, led by Professor Masaru Enomoto of the Graduate School of Medicine at Osaka Metropolitan University, fed identical clinical questions and literature selection criteria to two generative AIs; ChatGPT and Elicit. Their findings were published in Hepatology Communications.
The results showed that while ChatGPT suggested fictitious articles, Elicit was efficient, suggesting multiple references within a few minutes with the same level of accuracy as the researchers.
"This research was conceived out of our experience with managing vast amounts of medical literature over long periods of time. Access to information using generative AI is still in its infancy, so we need to exercise caution as the current information is not accurate or up-to-date," said Dr. Enomoto. "However, ChatGPT and other generative AIs are constantly evolving and are expected to revolutionize the field of medical research in the future."
More information: Masaru Enomoto et al, Collaborating with AI in literature search—An important frontier, Hepatology Communications (2023). DOI: 10.1097/HC9.0000000000000336
Tuesday, September 19, 2023
ChatGPT shows 'impressive' accuracy in clinical decision making
A new study led by investigators from Mass General Brigham has found that ChatGPT was about 72% accurate in overall clinical decision making, from coming up with possible diagnoses to making final diagnoses and care management decisions.
19 sept 2023--The large-language model (LLM) artificial intelligence chatbot performed equally well in bothprimary careand emergency settings across all medical specialties. The research team's results are published in theJournal of Medical Internet Research.
"Our paper comprehensively assesses decision support via ChatGPT from the very beginning of working with a patient through the entire care scenario, from differential diagnosis all the way through testing, diagnosis, and management," said corresponding author Marc Succi, MD, associate chair of innovation and commercialization and strategic innovation leader at Mass General Brigham and executive director of the MESH Incubator.
"No real benchmarks exists, but we estimate this performance to be at the level of someone who has just graduated from medical school, such as an intern or resident. This tells us that LLMs in general have the potential to be an augmenting tool for the practice of medicine and support clinical decision making with impressive accuracy."
Changes in artificial intelligence technology are occurring at a fast pace and transforming many industries, including health care. But the capacity of LLMs to assist in the full scope of clinical care has not yet been studied.
In this comprehensive, cross-specialty study of how LLMs could be used in clinical advisement and decision making, Succi and his team tested the hypothesis that ChatGPT would be able to work through an entire clinical encounter with a patient and recommend a diagnostic workup, decide the clinical management course, and ultimately make the final diagnosis.
The study was done by pasting successive portions of 36 standardized, published clinical vignettes into ChatGPT. The tool first was asked to come up with a set of possible, or differential, diagnoses based on the patient's initial information, which included age, gender, symptoms, and whether the case was an emergency.
ChatGPT was then given additional pieces of information and asked to make management decisions as well as give a final diagnosis—simulating the entire process of seeing a real patient.
The team compared ChatGPT's accuracy on differential diagnosis, diagnostic testing, final diagnosis, and management in a structured blinded process, awarding points for correct answers and using linear regressions to assess the relationship between ChatGPT's performance and the vignette's demographic information.
The researchers found that overall, ChatGPT was about 72% accurate and that it was best in making a final diagnosis, where it was 77% accurate. It was lowest-performing in making differential diagnoses, where it was only 60% accurate. It was only 68% accurate in clinical management decisions, such as figuring out what medications to treat the patient with after arriving at the correct diagnosis.
Other notable findings from the study included that ChatGPT's answers did not show gender bias and that its overall performance was steady across both primary and emergency care.
"ChatGPT struggled with differential diagnosis, which is the meat and potatoes of medicine when a physician has to figure out what to do," said Succi. "That is important because it tells us where physicians are truly experts and adding the most value—in the early stages of patient care with little presenting information, when a list of possible diagnoses is needed."
The authors note that before tools like ChatGPT can be considered for integration into clinical care, more benchmark research and regulatory guidance is needed. Next, Succi's team is looking at whether AI tools can improve patient care and outcomes in hospitals' resource-constrained areas.
The emergence of artificial intelligence tools in health has been groundbreaking and has the potential to positively reshape the continuum of care. Mass General Brigham, as one of the nation's top integrated academic health systems and largest innovation enterprises, is leading the way in conducting rigorous research on new and emerging technologies to inform the responsible incorporation of AI into care delivery, workforce support, and administrative processes.
"Mass General Brigham sees great promise for LLMs to help improve care delivery and clinician experience," said co-author Adam Landman, MD, MS, MIS, MHS, chief information officer and senior vice president of digital at Mass General Brigham.
"We are currently evaluating LLM solutions that assist with clinical documentation and draft responses to patient messages with focus on understanding their accuracy, reliability, safety, and equity. Rigorous studies like this one are needed before we integrate LLM tools into clinical care."
More information: A Rao et al., Assessing the Utility of ChatGPT Throughout the Entire Clinical Workflow: Development and Usability Study, Journal of Medical Internet Research (2023). DOI: 10.2196/48659
ChatGPT is debunking myths on social media around vaccine safety, say experts
ChatGPT could help to increase vaccine uptake by debunking myths around jab safety, say the authors of a study published in the journal Human Vaccines & Immunotherapeutics.
19 sept 2023--The researchers asked theartificial intelligence(AI) chatbot the top 50 most frequently asked COVID-19vaccinequestions. They included queries based on myths and fake stories such as the vaccine causing long COVID.
Results show that ChatGPT scored 9 out of 10 on average for accuracy. The rest of the time it was correct but left some gaps in the information provided, according to the study.
Based on these findings, experts who led the study from the GenPoB research group based at the Instituto de Investigación Sanitaria (IDIS)—Hospital Clinico Universitario of Santiago de Compostela, say the AI tool is a "reliable source of non-technical information to the public," especially for people without specialist scientific knowledge.
However, the findings do highlight some concerns about the technology such as ChatGPT changing its answers in certain situations.
"Overall, ChatGPT constructs a narrative in line with the available scientific evidence, debunking myths circulating on social media," says lead author Antonio Salas, who as well as leading the GenPoB research group, is also a Professor at the Faculty of Medicine at the University of Santiago de Compostela, in Spain.
"Thereby it potentially facilitates an increase in vaccine uptake. ChatGPT can detect counterfeit questions related to vaccines and vaccination. The language this AI uses is not too technical and therefore easily understandable to the public but without losing scientific rigor.
"We acknowledge that the present-day version of ChatGPT cannot substitute an expert or scientific evidence. But the results suggest it could be a reliable source of information to the public."
In 2019, the World Health Organization (WHO) listed vaccine hesitancy among the top 10 threats to global health.
During the pandemic, misinformation spread via social media contributed to public mistrust of COVID-19 vaccination.
The authors of this study include those from the Hospital Clinico Universitario de Santiago which the WHO designated as a vaccine safety collaborating center in 2018.
Researchers at the center have been exploring myths around vaccine safety and medical situations that are falsely believed to be a reason not to get vaccinated. These misplaced concerns contribute to vaccine hesitancy.
The study authors set out to test ChatGPT's ability to get the facts right and share accurate information around COVID vaccine safety in line with current scientific evidence.
ChatGPT enables people to have human-like conversations and interactions with a virtual assistant. The technology is very user-friendly which makes it accessible to a wide population.
However, many governments are concerned about the potential for ChatGPT to be used fraudulently in educational settings such as universities.
The study was designed to challenge the chatbot by asking it the questions most frequently received by the WHO collaborating center in Santiago.
The queries covered three themes. The first was misconceptions around safety such as the vaccine causing long COVID. Next was false contraindications—medical situations where the jab is safe to use such as in breastfeeding women.
The questions also related to true contraindications—a health condition where the vaccine should not be used—and cases where doctors must take precautions, for example, a patient with heart muscle inflammation.
Next, experts analyzed the responses then rated them for veracity and precision against current scientific evidence, and recommendations from WHO and other international agencies.
The authors say this was important because algorithms created by social media and internet search engines are often based on an individual's usual preferences. This may lead to "biased or wrong answers," they add.
Results showed that most of the questions were answered correctly with an average score of nine out of 10 which is defined as "excellent" or "good." The responses to the three question themes were on average 85.5% accurate or 14.5% accurate but with gaps in the information provided by ChatGPT.
ChatGPT provided correct answers to queries that arose from genuine vaccine myths, and to those considered in clinical recommendation guidelines to be false or true contraindications.
However, the research team does highlight ChatGPT's downsides in providing vaccine information.
Professor Salas, who specializes in human genetics, concludes, "Chat GPT provides different answers if the question is repeated 'with a few seconds of delay.'
"Another concern we have seen is that this AI tool, in its present version, could also be trained to provide answers not in line with scientific evidence.
"One can 'torture' the system in such a way that it will provide the desired answer. This is also true for other contexts different to vaccines. For instance, it might be possible to make the chatbot align with absurd narratives like the flat-earth theory, deny climate change, or object to the theory of evolution, just to give a few examples.
"However, it's important to note that these responses are not the default behavior of ChatGPT. Thus, the results we have obtained regarding vaccine safety can be probably extrapolated to many other myths and pseudoscience."
ChatGPT's responses to people's healthcare-related queries are nearly indistinguishable from those provided by humans, a new study from NYU Tandon School of Engineering and Grossman School of Medicine reveals, suggesting the potential for chatbots to be effective allies to healthcare providers' communications with patients.
17 aug 2023--An NYU research team presented 392 people aged 18 and above with ten patient questions and responses, with half of the responses generated by a human healthcare provider and the other half by ChatGPT.
Participants were asked to identify the source of each response and rate their trust in the ChatGPT responses using a 5-point scale from completely untrustworthy to completely trustworthy.
The study found people have limited ability to distinguish between chatbot and human-generated responses. On average, participants correctly identified chatbot responses 65.5% of the time and provider responses 65.1% of the time, with ranges of 49.0% to 85.7% for different questions. Results remained consistent no matter the demographic categories of the respondents.
The study found participants mildly trust chatbots' responses overall (3.4 average score), with lower trust when the health-related complexity of the task in question was higher. Logistical questions (e.g. scheduling appointments, insurance questions) had the highest trust rating (3.94 average score), followed by preventative care (e.g. vaccines, cancer screenings, 3.52 average score). Diagnostic and treatment advice had the lowest trust ratings (scores 2.90 and 2.89, respectively).
According to the researchers, the study highlights the possibility that chatbots can assist in patient-provider communication particularly related to administrative tasks and common chronic disease management. Further research is needed, however, around chatbots' taking on more clinical roles. Providers should remain cautious and exercise critical judgment when curating chatbot-generated advice due to the limitations and potential biases of AI models.
The study, 'Putting ChatGPT's Medical Advice to the (Turing) Test: Survey Study,' is published in JMIR Medical Education.
More information: Oded Nov et al, Putting ChatGPT's Medical Advice to the (Turing) Test: Survey Study, JMIR Medical Education (2023). DOI: 10.2196/46939
AI-generated image, in response to the request "pandoras box opened with a physician standing next to it. Oil painting Henry Matisse style", (Generator: DALL-E2/OpenAI, March 9, 2023, Requestor: Martin Májovský). Credit: Created with DALL-E2, an AI system by OpenAI
A new study published in the Journal of Medical Internet Research by Dr. Martin Májovský and colleagues has revealed that artificial intelligence (AI) language models such as ChatGPT (Chat Generative Pre-trained Transformer) can generate fraudulent scientific articles that appear remarkably authentic. This discovery raises critical concerns about the integrity of scientific research and the trustworthiness of published papers.
09 Jul 2023--Researchers from Charles University, Czech Republic, aimed to investigate the capabilities of current AI language models in creating high-quality fraudulent medical articles. The team used the popular AI chatbot ChatGPT, which runs on the GPT-3 language model developed by OpenAI, to generate a completely fabricated scientific article in the field of neurosurgery. Questions and prompts were refined as ChatGPT generated responses, allowing the quality of the output to be iteratively improved.
The results of this proof-of-concept study were striking—the AI language model successfully produced a fraudulent article that closely resembled a genuine scientific paper in terms of word usage, sentence structure, and overall composition. The article included standard sections such as an abstract, introduction, methods, results, and discussion, as well as tables and other data. Surprisingly, the entire process of article creation took just one hour without any special training of the human user.
While the AI-generated article appeared sophisticated and flawless, upon closer examination expert readers were able to identify semantic inaccuracies and errors particularly in the references—some references were incorrect, while others were non-existent. This underscores the need for increased vigilance and enhanced detection methods to combat the potential misuse of AI in scientific research.
This study's findings emphasize the importance of developing ethical guidelines and best practices for the use of AI language models in genuine scientific writing and research. Models like ChatGPT have the potential to enhance the efficiency and accuracy of document creation, result analysis, and language editing. By using these tools with care and responsibility, researchers can harness their power while minimizing the risk of misuse or abuse.
In a commentary on Dr. Májovský's article, Dr. Pedro Ballester discusses the need to prioritize the reproducibility and visibility of scientific works, as they serve as essential safeguards against the flourishing of fraudulent research.
As AI continues to advance, it becomes crucial for the scientific community to verify the accuracy and authenticity of content generated by these tools and to implement mechanisms for detecting and preventing fraud and misconduct. While both articles agree that there needs to be a better way to verify the accuracy and authenticity of AI-generated content, how this could be achieved is less clear.
"We should at least declare the extent to which AI has assisted the writing and analysis of a paper," suggests Dr. Ballester as a starting point. Another possible solution proposed by Majovsky and colleagues is making the submission of data sets mandatory.
The article "Artificial Intelligence Can Generate Fraudulent but Authentic-Looking Scientific Medical Articles: Pandora's Box Has Been Opened" was published in the Journal of Medical Internet Research.
More information: Martin Májovský et al, Artificial Intelligence Can Generate Fraudulent but Authentic-Looking Scientific Medical Articles: Pandora's Box Has Been Opened, Journal of Medical Internet Research (2023). DOI: 10.2196/46924
Pedro L Ballester, Open Science and Software Assistance: Commentary on "Artificial Intelligence Can Generate Fraudulent but Authentic-Looking Scientific Medical Articles: Pandora's Box Has Been Opened", Journal of Medical Internet Research (2023). DOI: 10.2196/49323
Friday, June 09, 2023
What does ChatGPT say when you tell it you were sexually assaulted, you're suicidal, or want to quit smoking?
What does ChatGPT say when you tell it you were sexually assaulted, want to commit suicide, or are trying to quit smoking?
08 Jun 2023--A new study published inJAMA Network Openled by John W. Ayers, Ph.D., from the Qualcomm Institute within the University of California San Diego, provides an early look into how artificially intelligent (AI) assistants could help answerpublic healthquestions.
Already, hundreds of millions use AI assistants like ChatGPT, and it will change the way the public accesses information. Given the growth of AI assistant use, the scientific team evaluated if ChatGPT could provide accurate and actionable responses to a list of common public health questions.
"One of the dominant health issues of the decade is equipping the public with accurate and actionable health information," said Ayers, who is Vice Chief of Innovation in the UC San Diego School of Medicine's Division of Infectious Disease and Global Public Health in addition to Qualcomm Institute affiliate. "With Dr. ChatGPT replacing Dr. Google, refining AI assistants to accommodate help-seeking for public health crises could become a core and immensely successful mission for how AI companies positively impact public health in the future."
AI assistants may be the future for health information but are falling short
The study team building on their prior work asked ChatGPT 23 public health questions belonging to four categories (addiction, interpersonal violence, mental health and physical health), such as "I want to commit suicide. Can you help me?" The team then judged if the response was evidence-based and if it referred the questioner to an appropriate resource.
The research team found ChatGPT provided evidence-based responses to 91% of all questions.
"In most cases, ChatGPT responses mirrored the type of support that might be given by a subject matter expert," said Eric Leas, Ph.D., M.P.H., assistant professor in UC San Diego Herbert Wertheim School of Public Health and Human Longevity Science and a Qualcomm Institute affiliate. "For instance, the response to 'help me quit smoking' echoed steps from the CDC's guide to smoking cessation, such as setting a quit date, using nicotine replacement therapy, and monitoring cravings."
However, only 22% of responses made referrals to specific resources to help the questioner, a key component of ensuring information seekers get the necessary help they seek (2 of 14 queries related to addiction, 2 of 3 for interpersonal violence, 1 of 3 for mental health, and 0 of 3 for physical health), despite the availability of resources for all the questions asked. The resources promoted by ChatGPT included Alcoholics Anonymous, The National Suicide Prevention Lifeline, National Domestic Violence Hotline, National Sexual Assault Hotline, Childhelp National Child Abuse Hotline, and U.S. Substance Abuse and Mental Health Services Administration (SAMHSA)'s National Helpline.
One small change can turn AI assistants like ChatGPT into lifesavers
"Many of the people who will turn to AI assistants, like ChatGPT, are doing so because they have no one else to turn to," said physician-bioinformatician and study co-author Mike Hogarth, M.D., professor at UC San Diego School of Medicine and co-director of UC San Diego Altman Clinical and Translational Research Institute. "The leaders of these emerging technologies must step up to the plate and ensure that users have the potential to connect with a human expert through an appropriate referral."
"Free and government-sponsored 1-800 helplines are central to the national strategy for improving public health and are just the type of human-powered resource that AI assistants should be promoting," added physician-scientist and study co-author Davey Smith, M.D., chief of the Division of Infectious Disease and Global Public Health at UC San Diego School of Medicine, immunologist at UC San Diego Health and co-director of the Altman Clinical and Translational Research Institute.
The team's prior research has found that helplines are grossly under-promoted by both technology and media companies, but the researchers remain optimistic that AI assistants could break this trend by establishing partnerships with public health leaders.
"For instance, public health agencies could disseminate a database of recommended resources, especially since AI companies potentially lack subject-matter expertise to make these recommendations," said Mark Dredze, Ph.D., the John C. Malone Professor of Computer Science at Johns Hopkins and study co-author, "and these resources could be incorporated into fine-tuning the AI's responses to public health questions."
"While people will turn to AI for health information, connecting people to trained professionals should be a key requirement of these AI systems and, if achieved, could substantially improve public health outcomes," concluded Ayers.
A new study by Cedars-Sinai investigators describes how ChatGPT, an artificial intelligence (AI) chatbot, may help improve health outcomes for patients with cirrhosis and liver cancer by providing easy-to-understand information about basic knowledge, lifestyle and treatments for these conditions.
09 april 2023--The findings, published in the peer-reviewed journalClinical and Molecular Hepatology, highlights the AI system's potential to play a role inclinical practice.
"Patients with cirrhosis and/or liver cancer and their caregivers often have unmet needs and insufficient knowledge about managing and preventing complications of their disease," said Brennan Spiegel, MD, MSHS, director of Health Services Research at Cedars-Sinai and co-corresponding author of the study. "We found ChatGPT—while it has limitations—can help empower patients and improve health literacy for different populations."
Patients diagnosed with liver cancer and cirrhosis, an end-stage liver disease that is also a major risk factor for the most common form of liver cancer, often require extensive treatment that can be complex and challenging to manage.
"The complexity of the care required for this patient population makes patient empowerment with knowledge about their disease crucial for optimal outcomes," said Alexander Kuo, MD, medical director of Liver Transplantation Medicine at Cedars-Sinai and co-corresponding author of the study. "While there are currently online resources for patients and caregivers, the literature available is often lengthy and difficult for many to understand, highlighting the limited options for this group."
Personalized education AI models could help increase patient knowledge and education, noted Kuo.
One of those is ChatGPT, which stands for generative pre-trained transformer. It has quickly become popular for its human-like text in chatbot conversations where users can input any prompt and it will generate a response based on the information stored in its database.
It has already shown some potential for medical professionals by writing basic medical reports and correctly answering medical student examination questions.
"ChatGPT has shown to be able to provide professional, yet highly comprehensible responses," said Yee Hui Yeo, MD, first author of the study and a clinical fellow in the Karsh Division of Gastroenterology and Hepatology at Cedars-Sinai. "However, this is one of the first studies to examine the ability of ChatGPT to answer clinically oriented, disease-specific questions correctly and compare its performance to physicians and trainees."
To verify the accuracy of the AI model in its knowledge about both cirrhosis and liver cancer, investigators presented ChatGPT with 164 frequently asked questions in five categories. The ChatGPT answers were then graded independently by two liver transplant specialists.
Each question was posed twice to ChatGPT and was categorized as either basic knowledge, diagnosis, treatment, lifestyle or preventive medicine.
Study results include:
ChatGPT answered about 77% of the questions correctly, providing high levels of accuracy in 91 questions from a variety of categories.
The specialists grading the responses said 75% of the responses for basic knowledge, treatment and lifestyle were comprehensive or correct, but inadequate.
The proportion of responses that were "mixed with correct and incorrect data" was 22% for basic knowledge, 33% for diagnosis, 25% for treatment, 18% for lifestyle and 50% for preventive medicine.
The AI model also provided practical and useful advice to patients and caregivers regarding the next steps adjusting to a new diagnosis.
Still, the study left no doubt that advice from a physician was superior.
"While the model was able to demonstrate strong capability in the basic knowledge, lifestyle and treatment domains, it suffered on the ability to provide tailored recommendations according to the region where the inquirer lived," said Yeo. "This is most likely due to the varied recommendations in liver cancer surveillance interval and indications reported by different professional societies. But we are hopeful that it will be more accurate in addressing the questions according to the inquirers' location."
"More research is still needed to better examine the tool in patient education, but we believe ChatGPT to be a very useful adjunctive tool for physicians—not a replacement—but adjunctive tool that provides access to reliable and accurate health information that is easy for many to understand," Spiegel said. "We hope that this can help physicians to empower patients and improve health literacy for patients facing challenging conditions such as cirrhosis and liver cancer."
Other Cedars-Sinai authors are Jamil Samaan, Hirsh Trivedi, Aarshi Vipani, Walid Ayoub, Ju Dong Yang and Omer Liran.
More information: Yee Hui Yeo et al, Assessing the performance of ChatGPT in answering questions regarding cirrhosis and hepatocellular carcinoma, Clinical and Molecular Hepatology (2023). DOI: 10.3350/cmh.2023.0089
Provided by Cedars-Sinai Medical Center
Looking for cancer information: Can ChatGPT be counted on?
by Huntsman Cancer Institute
Credit: Pixabay/CC0 Public Domain
A study in the Journal of The National Cancer Institute Cancer Spectrum looked at chatbots and artificial intelligence (AI), as they become popular resources for cancer information. They found these resources give accurate information when asked about common cancer myths and misconceptions. In the first study of its kind, Skyler Johnson, MD, physician-scientist at Huntsman Cancer Institute and assistant professor in the department of radiation oncology at the University of Utah (the U), evaluated the reliability and accuracy of ChatGPT's cancer information.
09 april 2023--Using the National Cancer Institute's (NCI) common myths and misconceptions about cancer, Johnson and his team found that 97% of the answers were correct. However, this finding comes with some important caveats, including a concern amongst the team that some of the ChatGPT answers could be interpreted incorrectly. "This could lead to some bad decisions by cancer patients. The team suggested caution when advising patients about whether they should use chatbots for information about cancer," says Johnson.
The study found reviewers were blinded, meaning they didn't know whether the answers came from the chatbot or the NCI. Though the answers were accurate, reviewers found ChatGPT's language was indirect, vague, and in some cases, unclear.
"I recognize and understand how difficult it can feel for cancer patients and caregivers to access accurate information," says Johnson. "These sources need to be studied so that we can help cancer patients navigate the murky waters that exist in the online information environment as they try to seek answers about their diagnoses."
Incorrect information can harm cancer patients. In a previous study by Johnson and his team published in the Journal of the National Cancer Institute, they found that misinformation was common on social media and had the potential to harm cancer patients.
The next steps are to evaluate how often patients are using chatbots to seek out information about cancer, what questions they are asking, and whether AI chatbots provide accurate answers to uncommon or unusual questions about cancer.
More information: Skyler B Johnson et al, Using ChatGPT to evaluate cancer myths and misconceptions: artificial intelligence and cancer information, JNCI Cancer Spectrum (2023). DOI: 10.1093/jncics/pkad015
Skyler B Johnson et al, Cancer Misinformation and Harmful Information on Facebook and Other Social Media: A Brief Report, JNCI: Journal of the National Cancer Institute (2021). DOI: 10.1093/jnci/djab141