Skip to Content

Informatics

Informatics is an international, peer-reviewed, open access journal on information and communication technologies, human–computer interaction, and social informatics, and is published monthly online by MDPI.

JCR - Q2 (Computer Science, Interdisciplinary Applications)| Cite Score - Q1 (Communication | Computer Networks and Communications)

Get Alerted

Add your email address to receive forthcoming issues of this journal.

All Articles (873)

Diabetic foot ulcers (DFUs) impose a substantial burden on people with diabetes and healthcare systems. The post-healing “transition phase” remains clinically challenging with limited guideline support. While large language models (LLMs) are increasingly proposed as clinical decision-support tools, their reliability in evidence-scarce scenarios is largely untested. This exploratory study benchmarked leading LLMs against European clinician consensus for the evidence-scarce scenario of DFU transition-phase management. Six LLMs (ChatGPT-4o, ChatGPT-5.0, Gemini 2.5 Flash, Gemini 2.5 Pro, Claude Sonnet 4.0, and Perplexity) were evaluated for accuracy and hallucination using a two-stage framework. Benchmarks were derived from an online survey of European DFU experts reflecting European-level and national practices (Denmark, Netherlands, UK). Binary outcomes were summarized as proportions with Wilson 95% confidence intervals. Paired within-item comparisons across models were assessed using Cochran’s Q, followed by exact McNemar tests with Holm correction (α = 0.05); inferential results were considered supportive due to the limited number of paired items (n = 8). Across 192 accuracy assessments and 384 reference checks, LLM accuracy ranged from 50–75% and declined when simulating national practices. Hallucination rates exceeded 50% in several models. LLMs rely on generic recommendations, which contrast with clinicians’ contextual, patient-centered reasoning, suggesting current limitations in their suitability for clinical decision support in DFU transition phase clinical care.

Informatics

20 July 2026

Schematic of the study protocol.

This study investigates consumer perceptions of procedural and distributive fairness/unfairness in service interactions involving AI versus human providers. It also examines how these perceptions are influenced by technology-related discomfort and the perceived responsibility of the service organization. A scenario-based online experiment was conducted using participants recruited through the Prolific research platform. Participants were asked to review a small business loan application process at a fictional digital bank where the application is processed either by an AI algorithm or a human loan officer. The study found that consumers experienced a higher level of satisfaction when they encountered procedural unfairness from humans compared to AI. The negative effect of procedural unfairness on satisfaction was amplified as algorithm discomfort and perceived company responsibility increased. The findings suggest that AI unfairness requires special attention to address differences in consumer reaction.

Informatics

20 July 2026

Background/Objectives: The increasing use of artificial intelligence (AI) chatbots for obtaining health-related information has raised concerns regarding the quality and reliability of the information they provide. This study aimed to compare the quality and reliability of responses generated by ChatGPT free tier (GPT-4o, with GPT-4.1 mini as the fallback model after the usage limit) and Gemini 2.5 Flash (Google, free version) regarding newborn screening tests. Methods: A total of 31 questions were developed based on international and national newborn screening guidelines and were posed to both chatbots. Responses were independently evaluated by two researchers using the DISCERN instrument and the Global Quality Score (GQS), and inter-rater reliability was assessed. Descriptive statistics and non-parametric tests were used to compare chatbot performance. Results: Both evaluators assigned significantly higher DISCERN total scores to Gemini than to ChatGPT free tier. For Evaluator 1, the mean DISCERN scores were 48.8 ± 10.2 for Gemini and 42.4 ± 5.9 for ChatGPT free tier (p < 0.001); for Evaluator 2, the corresponding scores were 53.7 ± 8.9 and 42.4 ± 5.8, respectively (p < 0.001). For the GQS ratings, Evaluator 1 rated Gemini significantly higher than ChatGPT free tier (3.6 ± 0.6 vs. 3.0 ± 0.5, p < 0.001), whereas Evaluator 2 found no statistically significant difference between the two chatbots (3.3 ± 1.3 vs. 3.4 ± 0.8, p = 0.877). Inter-rater reliability for DISCERN scores was excellent for ChatGPT and moderate for Gemini, whereas agreement for GQS ratings was low for both chatbots. Conclusions: Gemini demonstrated consistently higher DISCERN scores than ChatGPT free tier; however, its superiority in overall GQS ratings was not consistently supported across the two evaluators. Neither chatbot consistently achieved the highest levels of information quality. AI chatbots should therefore be considered supplementary sources of health information rather than substitutes for healthcare professionals or official health information resources.

Informatics

17 July 2026

Causality extraction is an important task in natural language processing, yet it remains underexplored in informal Arabic social media text, particularly in dialectal contexts. This study investigates causal-reason extraction from Saudi Arabic tweets related to sick-leave requests. A gold-standard dataset was annotated for multiple causality-related tasks, including cause-presence detection, cause-span extraction, cause-category classification, causal-marker detection, and causal marker text identification. The study compares two modeling paradigms: fine-tuned BERT-based models, represented by SaudiBERT and AraBERT, and prompting-based large language models (LLMs), represented by GPT-4.1-mini and Gemini-2.5-flash. The descriptive analysis showed strong class imbalance, substantial implicit causality, and uneven cause-category distributions. Results showed that SaudiBERT generally outperformed AraBERT when macro-level and minority-class performance were considered. Among LLMs, Gemini-2.5-flash achieved the strongest overall performance, particularly under natural 10-shot single-tweet prompting, while balanced few-shot prompting improved macro-F1 for cause-category classification. However, step-wise prompting did not consistently improve performance and may have introduced error propagation. Overall, the findings show that causality extraction in informal Saudi Arabic remains challenging, especially for implicit causal expression. The study highlights the complementary strengths of dialect-specific transformers and LLM-based prompting for Arabic causality extraction.

Informatics

17 July 2026

Highly Accessed Articles

News & Conferences

Latest Issues

Open for Submission

Journal Sections

Reprints of Collections

Advances in Construction and Project Management
Reprint

Advances in Construction and Project Management

Volume III: Industrialisation, Sustainability, Resilience and Health & Safety
Editors: Srinath Perera, Albert P. C. Chan, Dilanthi Amaratunga, Makarand Hastak, Patrizia Lombardi, Sepani Senaratne, Xiaohua Jin, Anil Sawhney
Advances in Construction and Project Management
Reprint

Advances in Construction and Project Management

Volume II: Construction and Digitalisation
Editors: Srinath Perera, Albert P. C. Chan, Dilanthi Amaratunga, Makarand Hastak, Patrizia Lombardi, Sepani Senaratne, Xiaohua Jin, Anil Sawhney
XFacebookLinkedIn
Informatics - ISSN 2227-9709