The Dawn of Artificial Intelligence in Healthcare
Artificial intelligence (AI) is a branch of computer science that develops computational systems capable of performing tasks that typically require human intelligence, including pattern recognition, prediction, and decision support. Artificial intelligence has become increasingly integrated into healthcare, offering new opportunities to improve diagnostics, prognostic assessment, and clinical decision-making. Recent advances in machine learning (ML), deep learning (DL), and large language models (LLMs), together with the increasing availability of digital health data, have accelerated the development of AI applications across medicine (1).
In infectious diseases and clinical microbiology (IDCM), AI has attracted growing interest because it can improve laboratory diagnostics, antimicrobial stewardship, prognostic assessment, outbreak surveillance, and clinical decision support. Although numerous studies have reported promising results, important challenges related to external validation, data quality, model generalizability, and clinical implementation continue to limit the routine adoption of these technologies.
This review was prepared based on the thematic discussions, expert presentations, and literature evaluations conducted within the 2025 Basic Education Program of the Assistants and Young Specialists Commission of the Turkish Society of Clinical Microbiology and Infectious Diseases, entitled AI in Infectious Diseases and Clinical Microbiology. Building on these educational activities, this review critically synthesizes current evidence on AI applications in IDCM, highlighting both their clinical potential and the methodological, practical, and ethical challenges that must be addressed before widespread implementation.
The Role of AI in Medicine
AI in medicine has evolved from early rule-based expert systems to modern ML and DL models capable of analyzing complex clinical data (2,3). These approaches have been widely adopted in specialties such as radiology and pathology, where image-based analysis has demonstrated substantial clinical utility (4,5). In contrast, AI applications in IDCM remain at an earlier stage of implementation because they require integrating microbiological, clinical, and epidemiological information. Accordingly, AI systems in this field should be evaluated according to their specific clinical applications, evidence base, and limitations rather than as a single technology.
ML and DL in Infectious Diseases and Clinical Microbiology
Machine learning and DL applications in IDCM have expanded rapidly over the past decade, driven by advances in computational methods and the increasing availability of digital laboratory and clinical data. Current applications can be broadly categorized into four major areas: diagnostic automation, clinical decision support, clinical documentation and symptom assessment, and epidemiological surveillance (Table 1). Although numerous studies have demonstrated promising results, most AI systems remain task-specific and have not yet undergone sufficient prospective or multicenter validation for routine clinical implementation.
Diagnostic Automation and Molecular Analysis
Diagnostic microbiology is among the most suitable areas for AI implementation because many laboratory procedures involve repetitive image interpretation and pattern-recognition tasks. Deep learning models, particularly convolutional neural networks (CNNs), have demonstrated high diagnostic performance in automated Gram stain interpretation and digital culture plate analysis, with reported agreement rates exceeding 95% under controlled experimental conditions (6). Nevertheless, further validation in diverse laboratory settings is required before widespread implementation.
Artificial intelligence has also advanced molecular diagnostics through automated analysis of whole-genome sequencing data. ML-based platforms, such as DeepARG and PARGT, enable rapid identification of antimicrobial resistance genes and facilitate genomic surveillance (7,8). However, their performance remains dependent on sequencing quality, reference database completeness, and the continuous updating of resistance gene repositories.
Clinical Decision Support Systems
Clinical decision support systems (CDSS) integrate laboratory findings with clinical information to assist with diagnosis, risk prediction, and treatment decisions. AI-based sepsis prediction models, including Sepsis Watch and the Epic Sepsis Prediction Model (Epic Systems, Verona, WI, USA) and Epic’s Sepsis Prediction Model, have demonstrated the potential to identify patients at increased risk before clinical deterioration occurs (9). However, inconsistent performance across institutions, false-positive alerts, and alert fatigue remain important barriers to routine implementation.
Similarly, AI-assisted diagnostic support platforms combine clinical findings with large medical knowledge bases to facilitate differential diagnosis, whereas natural language processing (NLP)-based applications improve access to evidence-based recommendations and clinical guidelines (10,11). Although these tools may improve efficiency, they should be regarded as supportive technologies rather than autonomous diagnostic systems.
Symptom Analyzers, Triage, and Documentation
Recent advances in NLP and LLMs have expanded AI applications to symptom assessment and clinical documentation. Symptom analyzers may assist with preliminary patient triage, whereas AI-powered documentation systems can automatically generate clinical notes and reduce administrative workload (12–15). Despite these advantages, current LLM-based applications remain susceptible to hallucinations, incomplete contextual understanding, and variable performance across different clinical scenarios, emphasizing the need for careful clinician review.
Epidemiological Surveillance and Real-World Performance
Artificial intelligence has become an increasingly valuable tool for infectious disease surveillance by supporting genomic epidemiology, outbreak monitoring, and pathogen evolution analysis. Platforms such as Nextstrain have played an important role in monitoring viral evolution during the COVID-19 pandemic (16,17). At the same time, the pandemic highlighted the importance of representative datasets, continuous model updating, and external validation, demonstrating that algorithmic performance alone does not guarantee successful implementation in real-world public health settings (18–20).
From Microscope to AI: Current Evidence on AI Models in Medical Parasitology
Microscopic diagnosis remains the cornerstone of medical parasitology but is often labor-intensive and dependent on experienced personnel. Consequently, AI-based image analysis has emerged as a promising approach to improve diagnostic efficiency and standardization, particularly in settings with limited access to expert microscopists (21).
Current evidence is strongest for blood-borne parasites, particularly Plasmodium spp., where CNN-based models have consistently demonstrated high diagnostic accuracy under experimental conditions (22,23). Similar automated ML frameworks are being developed to identify intestinal parasites in stool samples. For instance, DL algorithms trained on large datasets of helminth eggs have shown significant potential to differentiate species with high morphological similarity, such as Ascaris lumbricoides and Trichuris trichiura (24).
More recently, multimodal LLMs, including GPT-4 and Gemini, have been explored for parasitological image interpretation. However, unlike CNN-based models specifically trained for image classification, these general-purpose systems have shown inconsistent performance in recognizing complex microscopic structures and remain susceptible to hallucinations, particularly when image quality or contextual information is limited (25).
Recent literature largely agrees that while advanced computational models offer promising support for education and second-opinion assessment, they cannot yet function as primary diagnostic tools in parasitology. Therefore, current evidence supports using AI as an adjunct to, rather than a replacement for, expert microscopic evaluation, particularly in diagnostically challenging cases (26).
DL in the Diagnosis of Exanthematous Infectious Diseases
Deep learning has been increasingly investigated for automated recognition of infectious skin diseases because visual lesion assessment is an important component of clinical diagnosis. Most studies have focused on CNN-based image classification models that can distinguish common exanthematous infections from non-infectious skin disorders (27,28).
Recent studies have explored CNN-based models for classifying infectious skin diseases, including mpox, Lyme disease, and measles (29–32). While these studies report promising diagnostic performance in controlled experimental settings, their generalizability to real-world clinical environments remains uncertain and requires further validation (33).
Deep learning-supported systems also hold potential in dermatology by improving diagnostic consistency in atypical presentations and supporting remote evaluation through teledermatology applications (34). However, these systems’ performance depends heavily on the quality, representativeness, and diversity of training datasets. A major limitation is algorithmic bias, particularly the underrepresentation of darker skin tones in publicly available datasets, which may reduce diagnostic accuracy across diverse populations (35,36).
In addition, many current DL models rely exclusively on visual inputs and do not incorporate critical clinical information such as associated symptoms (e.g., fever, pruritus) or disease course, which are essential for differentiating infectious exanthems from non-infectious conditions such as drug eruptions.
Several commercially available mobile applications provide AI-assisted assessment of skin lesions; however, they should be regarded as screening tools rather than diagnostic systems because performance varies by image quality, lesion characteristics, and patient population (35,36).
Future research should prioritize prospective multicenter validation and the integration of clinical, epidemiological, and imaging data to improve the robustness and clinical applicability of AI models for infectious skin diseases.
AI-Based Antimicrobial Management
Antimicrobial stewardship (AMS) is one of the most promising applications of AI in infectious diseases because it requires the integration of microbiological, clinical, and epidemiological data to support timely treatment decisions. Current AI applications focus primarily on optimizing empirical antibiotic selection, identifying patients at increased risk of resistant infections, and supporting antimicrobial prescribing through CDSS (37,38).
ML models, particularly those analyzing longitudinal electronic health record data, have demonstrated encouraging performance in predicting healthcare-associated infections and identifying patients at high risk of multidrug-resistant organism (MDRO) infections, including carbapenem-resistant Enterobacterales (CRE) (39,40). These models may enable earlier risk stratification than conventional clinical scoring systems and support more targeted empirical antimicrobial therapy. However, most available evidence originates from retrospective studies, and prospective validation across different healthcare settings remains limited.
Although DL models can identify complex relationships within large clinical datasets, their predictions are often difficult to interpret in routine clinical practice. Explainable AI (XAI) approaches may help clinicians understand which variables contribute most to model predictions and thereby improve confidence in AI-assisted recommendations. However, XAI does not provide complete transparency into the internal processes of complex models and should not be interpreted as establishing causal relationships (41–43).
Large language models have recently emerged as potential AMS tools by summarizing clinical guidelines, assisting with literature retrieval, and supporting educational activities. Nevertheless, current LLMs remain susceptible to hallucinations, outdated knowledge, and limited access to real-time patient-specific information, limiting their use for independent therapeutic decision-making (44,45).
Overall, current evidence suggests that AI has considerable potential to optimize AMS by improving risk prediction and supporting antimicrobial prescribing. However, despite encouraging predictive performance, robust evidence demonstrating improvements in prescribing quality, antimicrobial consumption, or patient outcomes remains limited. Future progress will depend on high-quality clinical data, rigorous external validation, and successful integration of AI tools into AMS programs and routine clinical workflows (46).
ML Applications in Predicting the Prognosis of Infectious Diseases
Predicting disease progression is one of the most promising AI applications in infectious diseases because timely risk stratification can improve patient management and resource allocation. Unlike diagnostic models, prognostic models aim to estimate future clinical outcomes by integrating clinical, laboratory, and increasingly multi-omics data. Recent advances in multi-omics technologies have enabled the integration of genomic, transcriptomic, proteomic, metabolomic, and conventional clinical data into ML-based prognostic models, supporting more individualized risk prediction (47,48).
Machine learning algorithms, including Random Forest, Support Vector Machines (SVM), and gradient boosting methods such as XGBoost, have been applied to predict outcomes in sepsis, COVID-19, and Crimean-Congo hemorrhagic fever (CCHF) by integrating clinical and molecular data (47–50). For instance, in targeted metabolomic analyses of patients with CCHF, ML models have been used to analyze amino acid profiles, enabling the prediction of intensive care requirements with high reported accuracy. In parallel, specialized computational frameworks have facilitated biomarker discovery using transcriptomic and metabolomic datasets, expanding opportunities for precision prognostic modeling (50).
Despite these encouraging findings, most prognostic models have been developed using retrospective datasets and lack prospective external validation. Challenges related to data imbalance, heterogeneous patient populations, and differences in clinical practice continue to limit model generalizability. Future studies should prioritize multicenter validation, standardized reporting, and assessment of clinical impact to determine whether improved predictive performance translates into better patient outcomes.
AI in Scientific Writing and Scholarly Publishing
Artificial intelligence is increasingly incorporated into the research and publication process, supporting activities ranging from literature retrieval and evidence synthesis to manuscript preparation and editorial screening. Unlike clinical decision-support applications, these tools primarily aim to improve research efficiency rather than generate new scientific evidence. Their growing adoption has also raised important questions regarding transparency, accountability, and responsible use. AI-powered literature discovery platforms use NLP and citation network analysis to facilitate literature screening, identify relevant publications, and support evidence synthesis (51). Although these applications can substantially reduce literature review time, researchers still need to appraise their output critically. The ethical framework and permissible use cases of AI in scientific publishing are summarized in Table 2.
Beyond literature discovery, LLM-based applications are increasingly utilized for document analysis, including summarizing scientific manuscripts, extracting key findings, and assisting in evidence synthesis. These tools may enhance researchers’ ability to process large volumes of clinical data; however, their outputs require careful verification because LLMs may generate inaccurate interpretations, unsupported statements, or fabricated references (52).
Furthermore, LLMs such as ChatGPT and Claude are increasingly used as writing assistants to improve grammar, language fluency, and manuscript organization. They may be particularly valuable for researchers writing in a second language by facilitating clearer scientific communication. Nevertheless, authors should always critically review and edit AI-generated text to ensure scientific accuracy and appropriate interpretation (53).
Publishers increasingly employ AI-assisted tools to detect plagiarism, image manipulation, and potential ethical concerns during manuscript screening (54). Leading editorial organizations, including the International Committee of Medical Journal Editors (ICMJE) and the World Association of Medical Editors (WAME), have issued clear guidance stating that LLMs cannot be credited as authors, as they lack accountability, responsibility, and the capacity to provide informed consent (55–58).
Transparency regarding AI use has become a fundamental principle of responsible scientific publishing. Current editorial policies require authors to disclose the use of AI-assisted technologies during manuscript preparation while maintaining full responsibility for the accuracy, interpretation, and integrity of the published work (53–58). In parallel, emerging reporting frameworks, such as the ChatGPT, generative artificial intelligence and natural large language models for accountable reporting and use (CANGARU) guidelines, aim to standardize the responsible use and reporting of generative AI in research, emphasizing the need for human oversight throughout the scientific publication process (58). As these technologies continue to evolve, standardized reporting and transparent disclosure will be essential to ensure scientific credibility and maintain trust in AI-assisted research.
Current Challenges and Limitations of AI in Clinical Practice
Despite the rapid expansion of AI applications in healthcare, several methodological, technical, and ethical challenges continue to limit their routine implementation in IDCM. Clinical decision-making is inherently multidimensional, requiring integration of patient-specific factors, comorbidities, microbiological findings, and clinical context that algorithmic models may not fully capture (59). Key characteristics of black-box and explainable AI approaches relevant to clinical practice are summarized in Table 3.
Technical and Methodological Limitations
The performance of AI models depends heavily on the quality, completeness, and representativeness of the data used for model development. Although many studies have reported excellent predictive performance, most published models have been developed and validated using retrospective datasets from single institutions, limiting their generalizability across different healthcare settings (60). Differences in patient populations, disease prevalence, laboratory practices, and healthcare infrastructure may substantially affect model performance when AI systems are applied outside their original development environment.

Table 4. Performance comparison of sepsis prediction models: AI-based SPM versus traditional clinical scores.
Another important challenge is dataset shift, whereby changes in patient characteristics, pathogen epidemiology, diagnostic practices, or clinical workflows reduce the accuracy of previously developed models over time. These limitations are particularly relevant in infectious diseases, where evolving pathogens, emerging antimicrobial resistance, and changing treatment strategies continuously alter the clinical landscape. Table 4 summarizes the practical implications of these challenges for AI-based sepsis prediction systems and their major limitations compared with traditional clinical scores.
In addition, many AI models are evaluated primarily using measures of predictive performance, such as discrimination or accuracy, without adequately assessing their impact on clinical decision-making or patient outcomes. Consequently, excellent performance in retrospective validation studies does not necessarily translate into meaningful clinical benefit in routine practice. Future research should therefore prioritize prospective multicenter validation, standardized reporting, and evaluation of clinical effectiveness alongside algorithmic performance (60).
Bias, Fairness, and Equity
Artificial intelligence systems may inadvertently reproduce or amplify biases present in the datasets used for model development, resulting in unequal performance across different patient populations (61). Such biases may arise from underrepresentation of specific demographic groups, differences in disease prevalence, incomplete clinical data, or the use of proxy variables that fail to accurately reflect patients’ healthcare needs.
A well-recognized example is the Optum algorithm, which underestimated the healthcare needs of Black patients because healthcare expenditure was used as a proxy for illness severity rather than direct measures of health status (61). Similar concerns have been reported in medical imaging, where AI models trained predominantly on lighter skin tones demonstrated reduced diagnostic accuracy in individuals with darker skin pigmentation (36). These examples illustrate that algorithmic performance may not be uniform across populations and highlight the importance of evaluating AI systems for fairness before clinical implementation (62).
Reducing algorithmic bias requires diverse, representative datasets, continuous post-implementation monitoring, and external validation across different populations and healthcare settings. Incorporating fairness assessments into model development and evaluation is therefore essential to promote equitable and reliable AI-assisted healthcare.
Explainability and Clinical Trust
Many AI systems, particularly DL models, remain difficult to interpret because their internal decision-making processes are not directly observable. XAI methods aim to explain which variables or features influence model predictions and may improve clinicians’ confidence in AI-assisted recommendations. However, XAI does not necessarily provide complete transparency into model behavior, establish causal relationships, or fully explain how complex models generate their predictions (63).
These challenges have also been observed in real-world implementations. For example, the Epic sepsis prediction model showed limitations, including delayed alerts and high false-positive rates, which contributed to alert fatigue and reduced clinician confidence in the system (9). These experiences highlight the importance of rigorous clinical evaluation, continuous monitoring, and careful integration of AI systems into existing healthcare workflows.
Ethical and Human Factors
Beyond technical performance, implementing AI in healthcare raises important ethical and professional considerations. The World Health Organization has identified key principles for the responsible use of AI in health, including respect for human autonomy, beneficence, non-maleficence, transparency, accountability, inclusiveness, and sustainability (64). Adherence to these principles is essential to ensure that AI technologies are implemented safely, ethically, and equitably.
Although AI can support clinical decision-making by rapidly analyzing large volumes of data, it cannot fully replicate human judgment, empathy, ethical reasoning, or effective communication with patients. These human attributes remain particularly important when managing complex clinical situations, communicating diagnostic uncertainty, or making value-sensitive treatment decisions. Consequently, AI should be regarded as a tool that supports, rather than replaces, clinical expertise.
Taken together, these limitations suggest that the major challenge facing AI is no longer algorithm development alone, but generating robust clinical evidence demonstrating safe, equitable, and effective implementation in routine clinical practice. Future research should therefore emphasize prospective multicenter validation, transparent reporting, fairness assessment, and evaluation of clinical impact rather than predictive accuracy alone.
Conclusion and Future Directions
Artificial intelligence is rapidly reshaping many aspects of IDCM, with applications spanning diagnostic microbiology, prognostic modeling, antimicrobial stewardship, epidemiological surveillance, and scientific research. Current evidence indicates that AI can improve data analysis efficiency, support clinical decision-making, and facilitate the interpretation of increasingly complex clinical and laboratory information. However, evidence strength varies considerably across applications, and many reported models have yet to demonstrate consistent clinical benefit in routine practice.
The successful integration of AI into infectious diseases will depend not only on continued technological advances but also on the availability of high-quality and representative datasets, rigorous external validation, transparent reporting, careful evaluation of clinical impact, and adherence to ethical principles. Importantly, future research should move beyond reporting predictive accuracy alone and focus on whether AI-assisted systems improve patient outcomes, antimicrobial prescribing, healthcare efficiency, and health equity across diverse populations.
Rather than replacing clinicians, AI should be viewed as a complementary technology that supports evidence-based decision-making while preserving the essential role of clinical expertise, professional judgment, and patient-centered care. Close collaboration among clinicians, microbiologists, data scientists, and policymakers will be essential to ensure that AI is implemented safely, responsibly, and in a manner that delivers meaningful benefits to patients and healthcare systems.


