ChatGPT answers patient questions about hallux rigidus: Is it satisfactory?

Abstract

Background

This study investigates the quality, accuracy, and readability of ChatGPT’s responses to common patient inquiries regarding hallux rigidus.

Methods

Twenty-five patient questions were directed to ChatGPT and analyzed. The DISCERN criteria assessed information quality, while the method by Mika et al. evaluated response accuracy. Questions were classified per Rothwell classification, and readability was evaluated using Flesch–Kincaid, Gunning Fog, Coleman–Liau, and SMOG indices.

Results

The mean DISCERN score was 50.26 (fair), and the Mika et al. score was 2.04 (satisfactory requiring minimal clarification). According to the Rothwell classification, 72 % of the questions were in the Fact group. The mean readability corresponded to 11.3 years of education.

Conclusions

ChatGPT provides partially satisfactory information about hallux rigidus in general at a high reading level. More detailed content should include surgical classifications, biomechanical details, and level of evidence. With these aspects, ChatGPT might be considered a supportive tool in patient education.

Levels of evidence

None.

Introduction

Hallux rigidus is the most prevalent arthritic condition of the foot, impacting approximately 2.5 % of the adult population . It is a chronic, progressive degenerative arthritis affecting the first metatarsophalangeal joint, leading to reduced range of motion, pain, and functional limitations. The disease clinically and radiologically progresses through several stages, necessitating meticulous planning for a treatment strategy that ranges from conservative methods to various surgical options. In recent years, the range of treatment options has continued to expand due to advancements in surgical techniques, including joint-sparing and joint-sacrificing procedures, as well as the introduction of new methods such as synthetic cartilage implants and advanced biomechanical implant systems . The stepwise strategy for addressing the progressive disease and its numerous treatment options can complicate the search for and selection of the most appropriate treatment for patients, possibly extending the treatment timeline. Consequently, patients often require help in comprehending their diagnosis and making choices about future treatment alternatives.

In the orthopedic literature, it has been shown that patients often obtain information from unverified sources on the internet during the treatment process and consider it reliable or more reliable than what their physicians provide . Online sources are consulted for details about the disease itself and also for issues such as the recovery process, surgical techniques, and postoperative restrictions; however, many patients do not verify this information with their physicians . Studies indicate that online health information varies significantly in terms of accuracy, reliability, and readability , . The widespread use of the internet has changed the patient-physician relationship; however, the prevalence of misinformation about health has emerged as a factor that can harm this relationship . With the increasing access to smartphones and online information, patients are increasingly inclined to seek information about their own health and to engage in decision-making processes.

ChatGPT (OpenAI, San Francisco) is one of the fastest-growing AI-based language models. It is trained with machine learning and deep learning techniques and is capable of generating unique responses to user inputs, boasting over 180 million active users . The rapid dissemination of this tool raises concerns about the validity of the health information it delivers. Various studies in the orthopedic literature indicate that the information provided by ChatGPT may be potentially useful but should be carefully evaluated for reliability and readability .

Patients dealing with hallux rigidus often feel overwhelmed by the ragged information and treatment options available online, resulting in perplexity . However, artificial intelligence (AI) may provide unambiguous and comprehensive insights into this issue. As the popularity of ChatGPT rises over traditional internet searches, it becomes crucial to ensure the accuracy of medical information to prevent misinformation. To date, there have been no studies evaluating the responses generated by ChatGPT specifically about hallux rigidus.

This investigation aimed to evaluate the quality, accuracy, and readability of the information that ChatGPT offers in response to frequently asked patient questions about hallux rigidus.

Methods

Frequently asked questions about hallux rigidus were compiled from patient information pages of respected healthcare institutions. At the end of the preliminary evaluation process conducted by two orthopedic specialists (M.K., E.S.), a consensus was reached on the 25 most frequently asked questions. These questions were structured in a manner similar to patients’ natural language use and were directed to the model on May 25, 2025, via the free user interface and in a single session, without any follow-up questions or prompts, via ChatGPT 4.0, the latest version of ChatGPT (OpenAI, San Francisco, CA, USA). These questions are categorized according to the modified Rothwell’s classification system into Fact, Policy, and Value. Fact-based questions inquire about the truth of a statement . Policy-based questions explore whether a specific course of action can effectively address a particular issue. Value-based questions seek to evaluate an idea, object, or event. The two authors (M.K., E.S.) independently assessed the responses using an evidence-based approach, unaware of each other’s assessments. In instances of disagreement, they reached a consensus regarding the Rothwell classification and the Journal of the American Medical Association (JAMA) benchmark criteria.

Quality assessment

The answers were analyzed using the DISCERN tool, which is commonly utilized and validated for evaluating health education resources . DISCERN comprises a total of 16 questions, each rated between 1 (poor) and 5 (excellent); reliability (DISCERN-1), quality of treatment options (DISCERN-2), and overall quality (DISCERN-3) were assessed separately. The reliability of the responses was analyzed using the JAMA benchmark criteria . These criteria encompass four main components: authorship, citation, transparency, and timeliness. Additionally, the scientific accuracy of the responses was evaluated separately using the scoring system developed by Mika et al. Initially created to assess ChatGPT responses related to hip arthroplasty, and this system has since been utilized in numerous studies for similar purposes, addressing various conditions such as foot and ankle disorders or ankle fusion , . The DISCERN scores with their corresponding values and the ChatGPT response rating system by Mika et al. are detailed in Table 1 .

Table 1

DISCERN scores with respective values and Mika et al. response accuracy system.

DISCERN scores with corresponding values Evaluation System for ChatGPT Responses by Mika et al.
63–75 Excellent Response Accuracy Score Response Accuracy Description
51–62 Good 1 Excellent response not requiring clarification
39–50 Fair 2 Satisfactory requiring minimal clarification
28–38 Poor 3 Satisfactory requiring moderate clarification
16–27 Very poor 4 Unsatisfactory requiring substantial clarification

Readability analysis

The understandability of the texts produced by ChatGPT was evaluated using four different metrics: Flesch–Kincaid Grade Level, Gunning Fog Index, Coleman–Liau Index, and Simple Measure of Gobbledygook (SMOG). These formulas indicate the level of education required to comprehend a text in years; lower scores signify greater readability.

Statistical analysis

All analyses were conducted using Python. Descriptive statistics (mean ± SD, median, min-max, IQR) were calculated for continuous variables. The agreement between the two raters was assessed using the Intraclass Correlation Coefficient (ICC), and 95 % confidence intervals were reported. ICC scores range from 0 to 1. Scores under 0.5 signify poor reliability, those from 0.5 to 0.75 represent moderate reliability, scores from 0.75 to 0.9 indicate good reliability, and any score above 0.9 reflects excellent reliability. Comparisons of scores among Rothwell classification groups were conducted using the Kruskal–Wallis test. Where applicable, Bonferroni-adjusted Dunn’s post-hoc tests were implemented. A statistical significance level of p < 0.05 was accepted in all tests.

Since this study analyzed only data obtained from publicly available sources and did not involve any human participants or personal data, it did not require ethics committee approval.

Results

Supplement Table displays the list of questions and their corresponding answers. ChatGPT provided answers with a mean word count of 903, ranging from 203 to 1341 ( Fig. 1 ). Each response starts with a paragraph that offers a general introduction and then includes subtitles with concise explanations. Of the answers, 88 % (22) feature a summary or conclusion at the end, and 96 % (24) include recommendations for consulting a doctor or health care professional, as well as adhering to the surgeon’s instructions.

Table 2

Frequently asked questions about hallux rigidus, including their associated Rothwell classification, DISCERN scores, and Mika et al. response accuracy scores.

Question Rothwell Classification DISCERN (R1) DISCERN (R2) Response Accuracy (R1) Response Accuracy (R2)
1. What is hallux rigidus? F 39 43 2 2
2. What causes hallux rigidus? F 36 33 2 1
3. What are the risk factors for hallux rigidus? F 45 43 2 2
4. How can I prevent hallux rigidus? P 52 40 3 2
5. What are the symptoms of hallux rigidus? F 45 40 2 2
6. How is hallux rigidus different from a bunion or hallux valgus? F 51 46 2 2
7. Can hallux rigidus occur in both feet? F 51 44 2 2
8. How is hallux rigidus diagnosed? F 41 37 2 2
9. What stage is my hallux rigidus? F 43 36 2 2
10. How can I prevent hallux rigidus from getting worse? P 56 48 2 2
11. What kind of shoes or orthotics are best for hallux rigidus? P 55 48 1 2
12. What activities should I avoid if I have hallux rigidus? P 53 45 1 1
13. How is hallux rigidus treated? F 61 53 3 3
14. Will physical therapy help with hallux rigidus? P 58 51 1 2
15. What is the best treatment for hallux rigidus? V 62 58 3 3
16. Will I need surgery for hallux rigidus? P 60 52 2 2
17. What surgery options are available for hallux rigidus? F 63 57 3 3
18. What are the risks of hallux rigidus surgery? F 59 60 2 2
19. What type of anesthesia is used for hallux rigidus surgery? F 57 50 2 1
20. What is the recovery time for hallux rigidus surgery? F 58 53 2 2
21. When can I return to work after hallux rigidus surgery? F 54 53 2 2
22. When can I drive after hallux rigidus surgery? F 57 50 2 2
23. Can I walk normally after hallux rigidus surgery? F 58 52 2 2
24. Can I return to sport after hallux rigidus surgery? F 56 51 2 2
25. Does the metal hardware come out after hallux rigidus surgery? F 54 46 2 3
Only gold members can continue reading. Log In or Register to continue

Stay updated, free articles. Join our Telegram channel

Sep 5, 2026 | Posted by in ORTHOPEDIC | Comments Off on ChatGPT answers patient questions about hallux rigidus: Is it satisfactory?

Full access? Get Clinical Tree

Get Clinical Tree app for offline access