I.INTRODUCTION

Artificial intelligence (AI) chatbots are being adopted in higher education, but their real contribution to inclusive teaching is still scattered across disciplines and study designs. We systematically reviewed studies published from 2021 to 2025, following PRISMA 2020 guidance. From an initial corpus of 338 records, 11 studies met the inclusion criteria for qualitative synthesis. Overall, the evidence suggests that AI chatbots can support inclusion mainly by personalizing explanations and feedback, providing accessible interaction channels, and assisting teachers with scaffolding and monitoring learning.

At the same time, the literature remains limited in scope, with uneven coverage of disability profiles, learning contexts, and evaluation outcomes. We highlight methodological gaps and propose priorities for more rigorous, equity-focused research and implementation. Methodologically, the studies included in this review comprise systematic reviews, scoping reviews, meta-analyses, adoption studies, implementation-focused research, and contextual qualitative investigations. Most studies adopted qualitative descriptive, analytical, or mixed-method approaches, while others emphasized systematic evidence synthesis.

The methodologies frequently relied on narrative thematic analysis, perception-based evaluations, contextual interpretation, and exploratory implementations of AI-supported educational interventions. Overall, findings consistently indicate that AI chatbots may strengthen inclusive educational practices by facilitating adaptive learning pathways, accessible communication channels, and scaffolded instructional support. However, researchers also report concerns related to privacy, bias, academic integrity, technological dependency, and unequal access to digital infrastructure.

Despite the rapid growth of research in this field, significant gaps remain regarding long-term effectiveness, participation-oriented outcomes, and the implementation of AI chatbots for underrepresented learner populations, particularly students with mobility-related needs. Therefore, this systematic review addresses the following research question: What contributions of AI chatbots can be evidenced in inclusive pedagogical practices? The objectives are to (1) identify and synthesize the main contributions of AI chatbots to inclusion, accessibility, and equity, (2) describe research trends through bibliometric mapping, and (3) analyze the theoretical, methodological, pedagogical, and ethical gaps reported in the literature.

The remainder of this article is organized as follows. Section II presents the Materials and Methods, including the PRISMA 2020 protocol, search strategy, eligibility criteria, and bibliometric procedures. Section III presents the Results, including descriptive findings, bibliometric analyses, and thematic synthesis of the 11 included studies. Section IV discusses the pedagogical implications, scaffolding contributions, accessibility challenges, ethical considerations, and research gaps associated with AI chatbot integration in inclusive education. Finally, Section V presents the conclusions, highlighting the potential of AI chatbots to strengthen inclusive pedagogical practices when implemented through equity-oriented, accessible, and ethically grounded educational frameworks.

II.LITERATURE REVIEW

Inclusive education is widely recognized as a priority for educational quality and equity, as reflected in Sustainable Development Goal 4 (SDG 4) and the Education 2030 Framework for Action, which advocate for inclusive and equitable quality education for all learners [1,2]. In this study, inclusive pedagogical practices are understood as educational approaches aimed at removing barriers to learning and participation, promoting accessibility, and ensuring meaningful engagement for diverse learners within mainstream educational settings [3]. These practices are closely associated with Universal Design for Learning (UDL), which recommends multiple means of engagement, representation, and action/expression in order to address learner variability from the outset [4]. In cases where barriers persist, reasonable accommodations are also necessary to guarantee equitable participation and educational access [5]. Contemporary perspectives on disability further emphasize participation, contextual adaptation, and rights-based approaches instead of deficit-oriented interpretations [6,7].

Parallel to these developments, AI has emerged as a transformative force in education. Among AI applications, chatbots have gained increasing attention due to their conversational capabilities, scalability, and potential to provide personalized learning support. Educational institutions and researchers have increasingly explored how AI chatbots can enhance teaching, accessibility, and learner engagement while supporting inclusive educational environments. UNESCO’s ICT Competency Framework for Teachers also highlights the importance of digital competencies and ethical technology integration in inclusive educational practices [8].

Recent literature demonstrates growing interest in the relationship between AI chatbots and inclusive education. For instance, Ali D et al. [9] conducted a systematic review showing that ChatGPT can support tutoring, feedback, and adaptive explanations for diverse learners when pedagogically mediated. Alyoussef I et al. [10] examined factors influencing inclusive AI adoption in higher education and found that institutional support, collaboration, and perceived usefulness strongly shape implementation processes. Coughlan T et al. [11] analyzed disability-related student narratives and suggested that AI systems may help identify and organize support strategies, although ethical concerns regarding oversimplification remain. Beyond empirical studies, policy frameworks also shape how these tools are adopted: Colombia’s Decree 1421 (2017), for instance, establishes regulatory expectations for inclusive education and disability support [12].

Similarly, Dizon J and Prudente M [13] demonstrated through a meta-analysis that AI-powered chatbots can improve achievement when integrated as scaffolding, feedback, and guided-practice tools. Complementing these classroom-focused findings, emerging studies propose teacher professional-development approaches specifically oriented toward inclusive AI integration [14]. Hamzah R et al. [15] proposed an autism-focused individualized education chatbot model designed to support communication and developmental profiles. In a related line of work, culturally tailored health-promotion chatbots have also been examined, illustrating how accessibility considerations can be embedded for diverse and underserved populations [16]. In low-resource educational contexts, Kouam A and Muchowe R [17] argued that AI technologies may help mitigate educational inequities by extending learning support where institutional resources are limited.

Other studies further reinforce the potential of AI for accessibility and inclusion. Jaime-Vargas J [18] conducted a systematic review indicating that generative AI can enhance accessibility in higher education through personalized explanations and alternative content formats. Lin T and Riccomini P [19] explored technology-enabled inclusive learning in rural mathematics education, highlighting improved participation opportunities for students with disabilities. In addition, competency frameworks addressing educators’ digital skills and ICT use offer a further lens for guiding responsible implementation [20]. Pagliara S et al. [21] synthesized evidence showing that AI integration can support adaptive learning, accessibility tools, and learner-centered assistance in inclusive environments. Likewise, Pergantis P et al. [23] found that AI chatbots may contribute to executive-function scaffolding and cognitive regulation processes, while Potapiuk L and Sasiuk A [24] emphasized the importance of ethical safeguards, teacher preparation, and accessibility adaptation when implementing AI in inclusive educational systems.

Collectively, these studies suggest that AI chatbots can contribute to inclusive pedagogical practices through three major dimensions: personalization and instructional adaptation, functional accessibility, and equity-oriented educational support. In particular, scaffolding emerges as a recurring pedagogical mechanism, since chatbots can provide guided explanations, adaptive feedback, executive-function support, and differentiated learning assistance for students with diverse educational needs. Nevertheless, the evidence base remains fragmented due to methodological heterogeneity, limited longitudinal evaluation, and uneven representation of disability profiles and educational contexts.

III.MATERIALS AND METHODS

A.REPORTING STANDARDS AND PROTOCOL

This systematic review was conducted and reported in accordance with PRISMA 2020 [22]. The PRISMA flow diagram is provided (Fig. 1), and the PRISMA 2020 [22] checklist is provided as Supplementary Material S3 (PRISMA 2020 Checklist). The annual distribution of publications in the initial corpus is shown in Fig. 2.

Fig 1. PRISMA 2020 flow diagram for the identification, screening, eligibility assessment, and inclusion of studies in the systematic review (n = 11).

Fig 2. Annual distribution of publications in the initial corpus (n = 338); records outside 2021–2025 were removed prior to screening (see Fig. 1).

Protocol registration: The review protocol was not prospectively registered.

B.STUDY DESIGN

We conducted a documentary systematic review guided by PRISMA 2020, complemented by bibliometric analyses in VOSviewer to characterize the research landscape.

C.DATA SOURCES AND SEARCH STRATEGY

A sensitive search was performed in November 2025 using Scopus and Dimensions. The initial retrieval yielded 338 records (Scopus, n = 59; Dimensions, n = 279). Eligibility screening then applied the review’s temporal criterion (2021–2025) and an open-access filter (as defined by each database), and records outside the time window were removed prior to screening (see Fig. 1). To preserve reproducibility, both databases were queried using the same conceptual core (“inclusive education” AND “artificial intelligence chatbots”), after which each platform’s native filters were used to delimit publication years, document types, and subject categories. In Dimensions, the search incorporated Education-related categories together with Artificial Intelligence, Machine Learning, Psychology, and adjacent fields, and restricted results to open-access articles, chapters, and edited-book records. In Scopus, the search was executed in TITLE-ABS-KEY and limited to 2021–2025 open-access records from relevant subject areas (e.g., social sciences, computer science, engineering, psychology, chemistry, and chemical engineering) and selected document types. This broader “landscape” search served two complementary purposes: first, to identify studies for PRISMA screening; and second, to assemble a corpus suitable for bibliometric mapping of research growth, keywords, and collaboration patterns. The exact search strings, filters, export decisions, and database-specific limits are reproduced in Supplementary Material S1 so that readers can distinguish the broad bibliometric corpus from the smaller set retained for qualitative synthesis.

D.ELIGIBILITY CRITERIA

Inclusion criteria covered (i) empirical or review studies addressing AI chatbots in education, (ii) explicit relevance to inclusive education, accessibility, or equity, (iii) publication types including articles, books, or book chapters (as returned by the databases), and (iv) open-access availability. Exclusion criteria removed non-educational AI development papers, opinion pieces without evidence, and inclusive education studies without reference to AI chatbots or closely related technologies.

E.QUALITY APPRAISAL AND RISK-OF-BIAS CONSIDERATIONS

To strengthen methodological transparency, the included studies (n = 11) were appraised using a structured methodological coherence checklist aligned with the review’s research question and inclusion focus. The checklist was derived from the screening questions already defined by the review team and covered the following domains: (1) clarity and alignment of the title/aim with the topic; (2) study context and author affiliation; (3) publication type and relevance; (4) adequacy and credibility of sources used; (5) explicitness of conceptual/theoretical framework; (6) sample characteristics and appropriateness; (7) data collection techniques and instruments; (8) internal methodological coherence among design, sampling, instruments, and analysis; (9) claims of generalizability and supporting evidence; and (10) relevance and sufficiency of conclusions in relation to the review objective and inclusive education implications.

Each domain was coded qualitatively (Yes/Partially/No) to identify potential threats to validity and interpretive limitations rather than to exclude studies a priori. Given the heterogeneity of designs and outcomes across the evidence base, no single quantitative risk-of-bias score was calculated. Instead, appraisal outcomes informed the narrative synthesis by indicating where conclusions should be treated with caution (e.g., small samples, short-term implementations, and limited outcome measures).

F.SCREENING AND SELECTION PROCESS (PRISMA FLOW)

From 338 records, 13 duplicates were removed in Zotero, leaving 325 unique records. Records were then restricted to the 2021–2025 window, excluding 10 records (n = 315). Title and abstract screening focused on whether each record explicitly connected AI chatbots or closely related conversational AI systems with education, inclusion, accessibility, disability support, or equity. Full-text eligibility checks then verified the educational context, the relevance of the chatbot component, and the presence of evidence or analytical substance beyond opinion-based commentary. After this staged process, 11 studies met the inclusion criteria for rigorous analysis. This distinction is important: the initial corpus (n = 338) was used to map the broader research landscape, whereas the final PRISMA-included set (n = 11) was used to synthesize evidence on inclusive pedagogical practices.

G.DATA EXTRACTION AND SYNTHESIS

For each included study, we extracted educational level/context, target populations, inclusion/accessibility focus, chatbot function, study design, sample characteristics, instruments, outcomes, and reported limitations. Extraction was organized around a common coding sheet to reduce interpretive dispersion across heterogeneous study designs (Table I). The synthesis followed a narrative thematic approach centered on inclusion-related contributions (personalization, accessibility, and equity), alongside reported gaps and ethical challenges. In addition, we coded whether the chatbot was described as a general-purpose conversational tool or as a more specialized assistive or advisory system, because this distinction shaped both the type of pedagogical contribution claimed and the nature of the limitations reported by authors. We also noted whether the study foregrounded learner-facing uses, teacher-facing uses, or institutional/administrative support functions, which helped clarify the level at which “inclusion” was being operationalized.

Table I. Technical coding scheme used to characterize AI chatbot approaches and implementation attributes

CategoryCoding and operationalization
(1) Study metadata & contextYear; venue/source; country/region; educational level/discipline; study type (design/implementation, empirical evaluation, review).
(2) Inclusion & accessibility scopeTarget learners/needs (e.g., disability/NEE); barriers addressed; accessibility features (TTS/STT, captions, simplified language, multimodal supports, assistive compatibility).
(3) Use case & interactionPrimary purpose (tutoring, practice/feedback, Q&A, guidance/support); modality (text/voice/multimodal); platform (web/mobile/LMS/messaging); personalization/adaptation (if reported).
(4) AI approach & architectureModel family (rule/retrieval/classical ML/neural/transformer/LLM/hybrid); grounding (KB/ontology, RAG); tool use/agents; prompt or context strategies (if reported).
(5) Data & resourcesData sources (public/proprietary; curriculum/KB); dataset size/labeling (if reported); privacy handling (anonymization/consent) when applicable.
(6) Evaluation design & metricsDesign (experimental/quasi/observational/qualitative/mixed); baseline/comparator; sample (n), duration, setting (lab/field); metrics (learning outcomes, engagement logs, usability/satisfaction, accessibility-related measures, qualitative themes).
(7) Dependability, safety & governanceReported/assessed risks (hallucinations, bias, privacy leakage, academic integrity, robustness/jailbreak, transparency); mitigations (human oversight, guardrails/filtering, data minimization, secure storage, auditing/logging, risk-based evaluation); key findings, limitations, and future directions mapped to gaps.

H.BIBLIOMETRIC ANALYSIS

All RIS files were imported into Zotero and then processed in VOSviewer to generate (i) keyword frequency and co-occurrence networks, (ii) co-authorship by country, (iii) author citation networks, and (iv) co-citation maps. Before visualization, terms were standardized where necessary to reduce duplication caused by singular/plural variants, translation differences, and near-synonymous labels. The VOSviewer maps were interpreted descriptively: node size indicates relative occurrence or citation strength, link thickness indicates the strength of connections, and clusters suggest topical or intellectual proximity. These bibliometric outputs were not treated as evidence of effectiveness; rather, they contextualized the growth and structure of the field. Accordingly, the maps are presented as a landscape complement to the PRISMA synthesis, helping distinguish between a broad publication surge and the much smaller set of studies that explicitly address inclusion, accessibility, or equity in a way that met the review criteria.

The full search strings and database filters for Scopus and Dimensions are provided in Supplementary Material S1.

IV.RESULTS

A.DESCRIPTIVE RESULTS AND RESEARCH GROWTH

Across the initial corpus (n = 338), publication activity shows a steep recent increase, with peaks in 2024 and 2025 (as reported in the dataset trend analysis) (Fig. 2). This pattern suggests that AI chatbots and related conversational AI tools have become a rapidly expanding topic in educational research, although the final set of 11 included studies shows that only a small subset of this growth explicitly examines accessibility, equity, or inclusive pedagogical practice. Therefore, Fig. 2 should be read as evidence of field expansion, not as evidence that robust inclusion-focused interventions are already consolidated.

The distribution also helps explain the small number of studies retained for qualitative synthesis: much of the recent literature discusses AI in education broadly, while fewer publications meet the narrower criteria of chatbot relevance plus explicit inclusion, accessibility, or equity focus.

B.BIBLIOMETRIC HIGHLIGHTS (KEYWORDS AND NETWORKS)

Within the PRISMA-included studies (n = 11), the most frequent keywords include AI, chatbots, and inclusive education, suggesting a shared conceptual core around AI-mediated inclusion. Additional terms such as disability, accessibility, students’ perceptions, higher education, ChatGPT, and adaptive technologies show how the corpus connects technical tools with participation barriers, learner diversity, and pedagogical adaptation. Keyword frequency is summarized in Fig. 3.

Fig 3. Keyword frequency across the 11 included studies (n = 11). Author keywords were extracted from each included paper and standardized (translation where needed, plus singular/plural harmonization) for consistency.

This visualization clarifies that the review is not centered on chatbots as generic automation tools but on their relationship with inclusive education, disability support, accessibility, and adaptive learning. The presence of “students’ perceptions” also indicates that part of the evidence base remains perception-oriented, which reinforces the need for stronger outcome-based studies.

Network-based bibliometric visualizations were generated from the initial corpus (n = 338): author citation (Fig. 4), author co-citation (Fig. 5), country co-authorship (Supplementary Fig. S1), and keyword co-occurrence (Supplementary Fig. S2). In these maps, larger nodes indicate authors or terms with stronger presence in the corpus, while links indicate citation, co-citation, collaboration, or keyword association depending on the map. The networks are therefore useful for locating visible authors, intellectual proximity, and emerging conceptual clusters, but they should not be interpreted as direct evidence that a given chatbot intervention improves inclusion.

Fig 4. Author citation network (VOSviewer) derived from the initial corpus (n = 338).

Fig 5. Author co-citation network (VOSviewer) derived from the initial corpus (n = 338).

Taken together, Figs. 4 and 5 show that the broader field is still organizing around dispersed author groups rather than a single consolidated research tradition. This supports the need for the qualitative synthesis in Table II, which narrows the interpretation to the studies that directly address inclusive pedagogical practices.

Table II. Summary of key findings from the 11 studies included in the qualitative synthesis

StudyType/contextMain inclusion-related findingCaution or gap
Ali et al. [9]Systematic review of ChatGPT in teaching and learningChatGPT can support explanations, feedback, tutoring, and flexible learning interactions that may assist diverse learners when pedagogically mediated.Evidence is heterogeneous; inclusive value depends on instructional design, teacher guidance, and safeguards.
Alyoussef et al. [10]Higher education adoption studyInclusive learning adoption is shaped by collaboration, perceived usefulness, institutional support, and technology acceptance factors.Adoption readiness does not by itself demonstrate improved participation or accessibility outcomes.
Coughlan et al. [11]Analysis of disability descriptions and student suggestionsLearner-generated descriptions help identify barriers and inform support strategies that AI systems could help organize or route.AI-based interpretation of disability narratives may oversimplify needs or reproduce deficit-based assumptions if not governed carefully.
Dizon and Prudente [13]Meta-analysis of AI-powered chatbots in science educationChatbots can improve achievement when they are pedagogically integrated as scaffolding, practice, or feedback tools.Most effects are not specific to disability or inclusion; transfer to inclusive education requires targeted evaluation.
Hamzah et al. [15]Autism-focused individualized education chatbot modelSpecialized chatbot designs can be tailored to communication and developmental profiles, supporting individualized educational interaction.Prototype and modeling emphasis; classroom validity, usability, and long-term impact require further testing.
Kouam and Muchowe [17]Equity-focused contextual study in ZimbabweAI technologies may help reduce access gaps by extending learning support where institutional resources are limited.Infrastructure, connectivity, cost, language, and digital competence can reproduce inequity if not addressed.
Jaime-Vargas [18]Systematic review on generative AI, accessibility, and inclusion in higher educationGenerative AI can support accessible content, alternative explanations, and personalized learning assistance in higher education.Risks include privacy, bias, academic integrity, over-reliance, and uneven access to premium tools.
Lin and Riccomini [19]Technology-enabled inclusive learning in rural math educationTechnology can expand access and support improved math learning outcomes for students with disabilities in rural contexts.Findings depend on local infrastructure, teacher mediation, and contextual scalability.
Pagliara et al. [21]Scoping review of AI integration in inclusive educationAI can contribute to inclusive education through adaptive supports, accessibility tools, and learner-centered assistance.The evidence remains fragmented across tools, populations, and outcome definitions.
Pergantis et al. [23]Systematic review on AI chatbots and cognitive controlChatbot interaction may support executive-function scaffolding, cognitive regulation, and guided learning processes.Cognitive benefits require rigorous measurement and careful attention to dependency and autonomy.
Potapiuk and Sasiuk [24]Implementation-focused study in Ukrainian inclusive educationAI is positioned as an advanced learning tool for inclusive educational environments and differentiated support.Local readiness, teacher preparation, ethical safeguards, and accessibility adaptation remain decisive conditions.

C.CHARACTERISTICS OF INCLUDED STUDIES

The 11 studies comprise a balanced mix of secondary evidence syntheses (systematic/scoping reviews and a meta-analysis) and primary empirical work, plus one implementation-focused contextual study. This composition is important because it reveals a field that is conceptually active but still empirically consolidating: several papers synthesize emerging evidence or propose interpretive frameworks, while fewer studies evaluate sustained intervention effects under clearly defined inclusive conditions. Geographically, the evidence spans 10countries (two studies from the USA and one each from the Philippines, Saudi Arabia, Italy, Malaysia, Zimbabwe, Greece, Mexico, the UK, and Ukraine), suggesting broad international interest but also a fragmented knowledge base shaped by very different educational systems and resource conditions. Most studies are situated in higher education or inclusive schooling contexts and examine chatbot-mediated personalization, communication support, or equity-oriented assistance in resource-constrained settings. Across the sample, chatbots are described either as general-purpose LLM-based tools (e.g., ChatGPT) used for learning support and feedback, or as task-specific systems tailored to particular needs (e.g., autism support and executive-function coaching). Methodologically, the sample also reflects uneven maturity: some studies are broad reviews synthesizing patterns across settings [9,18,21,23], some are adoption or implementation studies focused on attitudes and enabling conditions [10,17,24], and others are practice-oriented or prototype-oriented contributions exploring how chatbot functions might be adapted for specific learners or instructional problems [11,15,19]. Table II summarizes the key inclusion-related findings and cautions from each included study, while full study-level characteristics are reported in Supplementary Material S2.

The table shows that the most consistent evidence concerns personalization, accessibility affordances, and equity-oriented support, while the most persistent weaknesses concern short-term designs, limited accessibility metrics, and uneven coverage of disability profiles. These patterns inform the thematic synthesis that follows.

D.THEMATIC SYNTHESIS: THREE INCLUSION-CENTERED CONTRIBUTIONS

1).PERSONALIZATION AND PEDAGOGICAL ADAPTATION

AI chatbots are consistently described as supporting individualized instruction and adaptive learning pathways, including intelligent tutoring roles and cognitive scaffolding for students facing learning/participation barriers. Across the included studies, personalization operates in at least three ways. First, chatbots can reformulate explanations, generate examples, and provide immediate feedback at different levels of complexity, which may help learners who require repeated clarification or alternative representations of content [9,13,18]. Second, they can support teachers in differentiating instruction by suggesting materials, structuring practice opportunities, or drafting accessible prompts aligned with learner needs [19]. Third, in more specialized designs, conversational agents can be configured to respond to particular communication or developmental profiles, as illustrated by autism-focused individualized education proposals [15]. Taken together, the literature suggests that the main pedagogical value of chatbots is not merely speed or automation, but their potential to make instructional adaptation more continuous, responsive, and feasible within ordinary teaching workflows.

2).FUNCTIONAL ACCESSIBILITY (OPERATIONAL AND COMMUNICATIONAL)

Evidence highlights accessibility gains through assistive affordances such as text-to-speech (TTS) and speech-to-text (STT) for learners with sensory, speech, or communication limitations, contributing to participation and autonomy. Several studies frame accessibility not only as the presence of technical features but also as the availability of alternative modes of interaction that reduce friction in academic tasks and communication [18,21,24]. In that sense, chatbots can serve as mediators between learners and instructional content by offering conversational prompts, language simplification, reminders, or supportive dialog that makes participation less dependent on a single format. Teacher- and institution-facing studies also suggest that chatbot systems may reduce administrative burden and guide students toward more appropriate forms of support, although the validity and fairness of such guidance still require careful evaluation [11]. The strongest pattern here is that conversational AI may widen access when it expands communicative options, but only if those options are designed and evaluated as accessibility supports rather than as generic convenience features.

3).EQUITY-ORIENTED SUPPORT IN LOW-RESOURCE CONTEXTS

Chatbots are positioned as compensatory agents where institutional resources are constrained, offering scalable support that may improve engagement and access to learning guidance. This contribution appears most clearly in studies that discuss disadvantaged or under-resourced settings, where access to specialist support, timely feedback, or individualized tutoring is limited [13,17,19]. In such contexts, conversational AI is described as a potentially lower-cost layer of assistance that can partially extend teacher capacity and provide learners with additional opportunities to ask questions, rehearse ideas, or obtain clarification. However, the equity promise is explicitly conditional. The same literature warns that premium access models, infrastructure gaps, language barriers, and uneven digital competence can reproduce exclusion rather than alleviate it [17]. Thus, equity-oriented value does not derive from the chatbot itself, but from whether institutions can integrate it in ways that are affordable, accessible, and pedagogically meaningful for populations facing structural disadvantage.

E.REPORTED LIMITATIONS AND GAPS

Across the 11 included studies, the evidence base remains methodologically constrained. Most contributions rely on small or convenience samples, brief implementation windows, and self-reported outcomes, with limited use of validated accessibility measures or participation-oriented indicators. These constraints reduce robustness, limit causal inference, and make it difficult to determine whether reported benefits translate into sustained learning gains and durable participation. Coverage of inclusion needs is also uneven: studies most frequently address sensory and communication-related barriers, while mobility-related needs and intersectional barriers (e.g., language and socioeconomic constraints) are rarely examined. Heterogeneity is visible at several levels: the reviewed papers differ in educational stage, chatbot architecture, target population, outcome definitions, and even in what counts as “inclusive” use. Some focus on adoption perceptions, others on learning outcomes, and others on conceptual framing or prototype development, which complicates cross-study comparison and weakens the basis for strong generalizations. Future work would benefit from aligning outcome definitions and reporting with participation-oriented disability frameworks and rights-based inclusion constructs, rather than treating disability primarily as an individual deficit [6,7]. It would also benefit from clearer reporting on implementation conditions—such as teacher mediation, access infrastructure, language settings, and prior digital competence—because these contextual variables likely influence whether a chatbot functions as an inclusive support or as an additional layer of exclusion.

For this reason, the synthesis avoids treating the presence of a chatbot as equivalent to inclusive impact. Instead, the interpretation emphasizes the pedagogical mechanism, the accessibility barrier addressed, the learner population considered, and the implementation conditions reported by each study.

V.DISCUSSION

This review indicates that AI chatbots can contribute to inclusive pedagogical practices primarily by enabling (i) personalization, (ii) accessibility through assistive modalities, and (iii) equity-oriented support under resource constraints. Importantly, these contributions should not be interpreted as “automatic inclusion.” The same technologies can create new barriers if implementation ignores UDL principles, contextual constraints, or the lived realities of disability and exclusion.

A.PEDAGOGICAL IMPLICATIONS (INCLUSION-FIRST INTEGRATION)

From an inclusion-first perspective, chatbots should be framed as supports for participation and learning rather than as substitutes for instruction. Aligning chatbot use with UDL principles (multiple means of engagement, representation, and expression) helps ensure that design choices accommodate learner variability and reduce the need for reactive, individualized fixes [4]. However, the review also reinforces that technology does not deliver inclusion automatically: inclusion depends on pedagogical intent, barrier removal, and ongoing adaptation in context—consistent with inclusive pedagogy frameworks that prioritize participation and avoid separating “most” from “some” learners [3,6]. Accordingly, implementation should specify inclusion objectives (e.g., participation, autonomy, and accessibility), define who benefits and under which conditions, and combine chatbot use with teacher mediation and accessible learning resources. A practical implication is that educators and institutions should begin with barrier analysis rather than tool enthusiasm: before adopting a chatbot, they should clarify which participation barrier is being addressed, what alternative modality or scaffold is required, how human oversight will be maintained, and how benefits will be evaluated for different learner groups. Under this logic, chatbots become one component within a broader inclusive design strategy, not the centerpiece of it.

B.ETHICAL AND PRACTICAL RISKS

The reviewed literature flags risks that are directly relevant to the trustworthy and responsible deployment of AI-driven chatbots: academic integrity concerns (including plagiarism), privacy and data protection challenges, potential over-reliance that may undermine learner autonomy, and bias or inaccuracies in chatbot outputs. These issues become especially sensitive in inclusive education settings, where learners may be more exposed to harm from misinformation, intrusive data practices, or stigmatizing recommendations. Related reviews also highlight cognitive and executive-function considerations when chatbots are used as learning supports [23], and context-specific work in inclusive environments underscores the need to align AI tools with local accessibility realities [24]. Framing these risks within established AI ethics and governance frameworks can strengthen both interpretation and recommendations, linking concerns about bias, privacy, transparency, and integrity to actionable principles such as human oversight, accountability, data minimization, and risk-based evaluation [2527]. An additional concern emerging from the included literature is representational harm: when systems infer needs from disability descriptions, user prompts, or interaction histories, they may oversimplify complex support requirements or encode narrow assumptions about what counts as appropriate participation [11,21]. For inclusive education, this means that technical performance alone is insufficient; governance must also address fairness of interpretation, contestability of recommendations, and the possibility of unintended stigmatization.

C.RESEARCH AGENDA

Future work should move beyond describing “use” of chatbots and instead test how, when, and for whom chatbots contribute to inclusion. This requires treating inclusion as barrier removal and participation (not merely usability) and explicitly linking chatbot functions to inclusive pedagogical mechanisms (e.g., UDL-aligned multiple means of engagement/representation/action). The evidence reviewed here suggests that methodological progress will depend not only on larger samples but also on sharper conceptualization: studies need to identify what kind of chatbot is under examination, which barrier it is intended to address, which population is expected to benefit, and what inclusion-sensitive outcomes will indicate success. To strengthen the evidence base, future studies should prioritize:

  • •Longer-term and larger-scale evaluations (multi-site when possible) with clearly defined inclusion outcomes and accessibility metrics, including participation, autonomy, and equity indicators rather than satisfaction-only measures. In particular, studies should distinguish short-term usability gains from sustained pedagogical effects and should report implementation conditions in enough detail to support replication.
  • •Targeted research for underexplored populations, especially learners facing mobility-related barriers, and broader disability profiles beyond sensory/communication needs, with attention to intersectional constraints (e.g., disability plus socioeconomic or language barriers). Comparative work across rural/urban, low-resource/high-resource, and different linguistic settings would be especially valuable for testing the equity claims often made about conversational AI.
  • •More explicit theoretical integration (UDL plus participation-oriented disability and inclusive pedagogy frameworks) to avoid defaulting to deficit-based or medicalized interpretations of disability, and to operationalize inclusion as barrier removal and participation [37]. Studies should state the adopted framework, specify hypothesized mechanisms (e.g., scaffolding, reduced cognitive load, alternative modality access), and align measures accordingly. This would also help prevent the common slippage between “accessible,” “personalized,” and “inclusive” terms that are related but not interchangeable.
  • •Governance models for safe deployment (transparency, human oversight, data minimization, and integrity safeguards) grounded in established AI ethics and risk management frameworks [2527]. At minimum, future studies should report data handling, bias mitigation steps, and human-in-the-loop oversight practices to enable reproducibility and comparability. Research should also consider institutional readiness, teacher professional development, and procurement/access models, because implementation quality may be as consequential as technical capability for inclusive outcomes.

D.LIMITATIONS OF THIS REVIEW

This review is limited by database scope (Scopus and Dimensions), open-access filtering, and the small number of eligible studies (n = 11), which restricts the ability to generalize across educational systems and disability groups. The open-access restriction may have introduced selection bias by underrepresenting relevant evidence not indexed as OA. In addition, heterogeneity in study designs, populations, and chatbot architectures limits direct comparability across studies and constrains the feasibility of quantitative synthesis. Finally, bibliometric maps reflect the initial corpus (n = 338) and should be interpreted as a landscape view rather than as evidence of effectiveness, which is based on the PRISMA-included set (n = 11). A further limitation is that some included studies examined broader AI-enabled inclusion ecosystems rather than chatbot interventions in isolation; although these studies remained relevant to the review question, their broader scope may blur boundaries between chatbot-specific effects and more general digital-support dynamics. This was handled through narrative interpretation, but it should be considered when reading the synthesis.

V.CONCLUSION

AI chatbots show promise for inclusive education by supporting personalized instruction, functional accessibility (e.g., TTS/STT), and equity-oriented assistance in low-resource contexts. However, the current evidence base is still small and methodologically heterogeneous, and benefits should not be assumed to generalize across learners, contexts, or time. A key implication of this review is that “inclusion” should be treated as a theory-driven construct: chatbot adoption is most defensible when embedded within inclusive pedagogical designs—anchored in UDL and reasonable accommodation—and when outcomes are evaluated in terms of participation, autonomy, and accessibility rather than usability alone [47]. Finally, responsible scaling requires governance that addresses bias, privacy, transparency, and academic integrity through risk-based approaches and explicit human oversight [2527]. In practical terms, the literature reviewed here suggests that the question is no longer whether conversational AI can enter inclusive education, but under what pedagogical, ethical, and institutional conditions it can do so without reproducing the very barriers it claims to reduce.