I.INTRODUCTION
Since Turing’s 1950 question, “Can machines think?”, AI has evolved from infancy to rapid growth, with generative AI enabling tasks that once took hours or days to be completed in seconds [1,2]. Tasks that used to take hours or days now it takes few seconds with AI tools [3]. It is used in various domains starting from medicine [4,5], to engineering[6,7], and basic sciences [8–13]. Students and educators have been using these AI tools in studying [14,15], preparing exams or quizzes [16,17], and generating or evaluating courses’ syllabi [18,19]. The shift to online learning during events like COVID-19 and regional conflicts exposed vulnerabilities in STEM assessments, as students increasingly used AI tools like ChatGPT, raising academic integrity concerns [20]. Research shows a nearly 200% spike in use of illicit academic help services (e.g., Chegg) for chemistry during remote learning (April–August 2020), highlighting students’ reliance on third-party support even in exams [21]. Generative chatbots like ChatGPT (2022) and DeepSeek (2023) pose new challenges, as students may not see their use as cheating. Authorities advise assessment reform and AI ethics education rather than relying on pen-and-paper exams or unreliable detection [22]. For example, researchers successfully submitted AI-generated exam answers that passed human markers nearly undetected, prompting calls for redesigning assessment formats [23]. DeepSeek’s chain of thought transparency and multimodal capabilities reportedly boost its performance especially in technical and science domains, sometimes surpassing ChatGPT in elaborate reasoning and chemical problem-solving [24]. Other recent studies highlight that DeepSeek outperforms ChatGPT in generating bibliographic references and retrieving accurate academic data, though no model yet guarantees full accuracy [25]. Its growth is mirrored by major Chinese universities integrating DeepSeek into curricula both to teach AI and to address security and ethical concerns [26].
Bloom’s Taxonomy, developed by Benjamin Bloom in 1956 and revised in 2001, classifies educational goals by cognitive complexity [27,28]. Bloom’s Taxonomy, a hierarchy from basic recall to complex creation, guides learning objectives, instruction, and assessment. It promotes higher-order thinking: analysis, evaluation, creation, ensuring alignment between curriculum, teaching, and assessment for deeper, transferable learning [29]. Despite the growing use of LLMs in education and chemistry assessment, few studies have systematically evaluated ChatGPT and DeepSeek on chemistry exam questions across multiple disciplines, and existing work has not fully examined performance according to Bloom’s taxonomy. The aim of this study is to address these gaps by evaluating the effectiveness of ChatGPT vs. DeepSeek at answering chemistry exam questions (general, organic, inorganic, and physical chemistry exams), classification of the exam questions according to Bloom’s verbs, rephrasing the exam questions using both AI chatbots in order to alter the level of the questions and hence lowering cheating levels for any upcoming online assessments.
The remainder of the paper is structured as follows: Section II presents a review of prior studies on the application of large language models in education and question-solving, highlighting gaps addressed by this study. Section III details the methodology, including the chemistry exams and procedures for evaluating ChatGPT and DeepSeek. Section IV presents the results and discussion of the models’ performance in solving chemistry exam questions. Section V examines the ability of both chatbots to classify and rephrase exam questions according to Bloom’s Taxonomy. Finally, Section VI concludes the paper and outlines directions for future work.
II.RELATED WORK
Recent machine learning approaches have automated classifying exam questions by Bloom’s levels. Khalifa and Mohamed [30] used term frequency–inverse document frequency and a support vector machine to achieve ∼80% accuracy on 600 questions. Reviews highlight the importance of verbs and parts of speech for better classification [31]. Transformer-based classifiers like DistilBERT have demonstrated validation accuracies near 91%, notably improving performance on higher-order levels (Apply through Create) [32]. Hwang et al. studied AI methods for generating and evaluating Bloom’s Taxonomy-aligned multiple-choice questions (MCQs) in chemistry and biology using GPT-3.5, RoBERTa, and expert review, showing the models can create higher-order thinking questions and align with human complexity assessments [33]. Fergus et al. found that ChatGPT gives well-written chemistry answers but struggles with complex application, interpretation, and non-text information, suggesting assessments emphasize problem-solving and data interpretation [34]. Studies showed that ChatGPT helped chemistry students improve writing and lab reports, reduced grammar errors [35], support lesson planning [36], and assisted with general and analytical chemistry calculations [37]. A recent study found that ChatGPT correctly answered 80% of high school general chemistry entrance exam questions and aided in generating new questions through prompt engineering [38]. Clark [39] discussed ChatGPT’s ability to solve general chemistry exam questions. It excelled at recognizing concepts in closed-response questions but had a problem-solving success rate of 44%, below the class average of 69%. Its open-response answers showed strong language skills but often included flawed yet convincing explanations. Sorenson and Hanson [40] used Rasch analysis on general chemistry exams and accurately identified AI-influenced answers with few false alarms, showing AI-generated patterns are detectable. In an earlier study, they found no evidence of increased cheating in unproctored online exams compared to prior in-person exams [41]. Kassem et al. [13] compared ChatGPT and Gemini (formerly Bard) on basic organic topics, including nomenclature, classification, molecule drawing, and identifying resonance contributors. Other studies explored general chemistry applications of ChatGPT [42], but, to our knowledge, none evaluated its performance on organic, physical, and inorganic chemistry exams. Moreover, few studies investigated DeepSeek’s use in chemistry exams. To our knowledge, no peer-reviewed study compared ChatGPT and DeepSeek on chemistry exams under controlled Bloom’s taxonomy conditions. Bouchra et al. [24] provided a broad comparison in science education but lacked peer review, DOI, and analysis by question difficulty.
III.METHODOLOGY
All data used in this study were anonymized and reported as class averages. Permission to use these data for research and publication purposes was obtained from the Lebanese International University (LIU). The study did not involve any identifiable student information, and all procedures were conducted in accordance with institutional guidelines for research ethics. The exams selected for this study were administered during the Fall 2023–2024 semester (hereafter referred to as Fall 2023). One examination from each course was analyzed (one general, one physical, one organic, and one inorganic), and the number of questions in each exam is reported in Table I.
Table I. Number and type of questions analyzed per exam
| Exam | Number of MCQs/true or false/matching | Number of subjective questions |
|---|---|---|
| 22 | 5 | |
| 21 | 5 | |
| 16 | 7 | |
| 25 | 6 |
These chemistry courses at LIU enrolled students from prepharmacy, biomedical, nutrition, chemistry, biochemistry, and biology majors. Fall 2023 was the first fully on-campus semester after COVID-19, with exams adjusted for the prior online period. A similar online/hybrid approach occurred during Lebanon’s October 2024 war (Fall 2024), when students, more familiar with AI, likely used chatbots for cheating. Fall 2024 exam results were compared with Fall 2023 to assess potential cheating. The data were collected for both chatbots in December 2024 and January 2025. Every exam question was fed to the said AI chatbot, and the following prompts were asked in order:
- 1.Answer the question
- 2.Classify the question according to Bloom’s taxonomy
- 3.Use the above example to write the question in another way using the other verbs of Bloom’s taxonomy.
Sometimes, the chatbot did the paraphrasing of step 3 above for one verb only from Bloom’s taxonomy so we used a clearer prompt and asked to use the other Bloom’s taxonomy verbs by specifying the category as remember, understand, apply, analyze, evaluate, and create. The results of this study (grading and question paraphrasing) were tabulated so an easy comparison could be made between ChatGPT and DeepSeek. The results of all outcomes are summarized in the supporting information in tables as Table Si. For this study, DeepSeek operated in its deep-thinking mode, while ChatGPT utilized its internal reasoning mode to generate responses.
Statistical analyses were performed using IBM SPSS Statistics. To evaluate the performance of ChatGPT and DeepSeek on MCQs across the examined chemistry exams, we conducted a paired analysis at the individual question level. Each MCQ was coded as a binary outcome (1 = correct, 0 = incorrect) for both models. Because the same questions were answered by both models, differences in performance were analyzed using McNemar’s exact test, which is appropriate for paired binary data and provides an exact p-value for small sample sizes. Discordant pairs, where one model answered correctly while the other did not, were used to calculate the odds ratio, serving as a measure of effect size, with 95% confidence intervals computed using exact methods. Paired outcome tables and odds ratios were reported in the Supporting Information, while summary statistics, effect sizes, and p-values were presented in the main text to provide both descriptive and inferential insights into the relative performance of the two models. Statistical analysis was restricted to MCQs because they provide objective, binary outcomes, whereas subjective questions involve grading variability that limits reliable statistical comparison in a single exam setting.
IV.RESULTS AND DISCUSSION
This study evaluates ChatGPT-4.o and DeepSeek on four undergraduate chemistry exams (general, physical, inorganic, organic) at LIU, including their ability to classify and paraphrase questions by Bloom’s taxonomy. DeepSeek takes longer to respond (20 seconds to minutes) and sometimes has server issues, while ChatGPT responds instantly with no delays. DeepSeek can upload unlimited attachments, whereas ChatGPT requires a paid plan. Exam grades are summarized in Fig. 1. Both chatbots passed the general and physical chemistry exams with overall grades above 80%, but scored below 60% on inorganic and organic chemistry. On the physical chemistry exam, ChatGPT scored 98% and DeepSeek 100%. DeepSeek outperforms ChatGPT in inorganic (62% vs. 56%) and general chemistry (90% vs. 82%), while ChatGPT scores higher in organic chemistry (38% vs. 23%). Comparing with class averages in Fall 2023 and 2024, physical and general chemistry averages rise by ∼10%, with exams of comparable difficulty and students coming from online or hybrid backgrounds. Although the increase coincides with chatbot availability, their impact cannot be separated from instructional or cohort factors. Organic chemistry averages drop from 80% in Fall 2023 to 57% in Fall 2024, but no conclusions about chatbot effects can be drawn due to possible differences in instruction, assessments, or cohorts. Inorganic chemistry averages remain comparable, likely reflecting on-campus exams.
Fig. 1. The grades of ChatGPT-4.o and DeepSeek obtained for four chemistry exams of Fall 2023. A comparison with class averages of Fall 2023 and Fall 2024 is shown as well.
A.SOLVING ORGANIC CHEMISTRY EXAM
The main topics discussed in this organic chemistry exam are revision topics from general chemistry, resonance, acids and bases, stereochemistry, and reactions of alkanes. Both ChatGPT and DeekSeek perform poorly in most of the organic chemistry exam topics. The results are summarized in Table S1 (in supporting info) with the grades, and the comments on every question are shown. A summary of the main topics is shown in Table II. It has been found that both chatbots cannot handle drawing structures or analyzing molecules in space, which is reflected in lower accuracy in such problem types. Both chatbots also show difficulties with stereochemistry problems such as assigning configuration, showing relationship between structures, assigning chirality/optical activity of molecules, and in converting a perspective formula to Fischer projection. DeepSeek misanalyzes groups around the chiral center and can replace atoms by other atoms, and the result in most cases is either wrong or correct by chance. On the other hand, ChatGPT can identify clearly the groups around the chiral center, number them correctly and can assign configuration when the least priority group is at the back in perspective formulas (dash-wedge).
Table II. Main topics covered in the organic chemistry exam with comments on the outcomes of the answers
| Question Type | ChatGPT | DeepSeek |
|---|---|---|
| Weak | Weak | |
| Can describe movement of electrons but cannot draw | Analyze wrong structure | |
| Can discuss the factors but cannot manage to rank correctly in most of times | Can discuss the factors but cannot manage to rank correctly in most of times | |
| 1. Can predict product of easy molecules | 1. Can predict product of easy molecules |
However, when assigning configurations in Fischer projections, ChatGPT swaps the horizontal lines with the vertical lines. In other words, according to ChatGPT the vertical and horizontal lines in Fischer projection are the towards and backward positions, respectively. The latter ends up with structures having an opposite configuration. Both chatbots cannot convert perspective formulas to Fischer projections or a Newmans projection to line angle formula. This finding is particularly concerning given research by Hsin-Kai [43] who emphasized the critical importance of visual-spatial reasoning in chemistry education. Drawing resonance structures in organic chemistry is one of the most tedious tasks a student finds and that’s why too many papers have been published to aid students and instructors [44–46]. Both chatbots still have problems in drawing resonance structures in spite of the fact that both chatbots state the rules of drawing resonance but fails to draw it. DeepSeek always analyzes the wrong structure, and it investigates the resonance structures of ions already found in its data base as nitrate or carbonate ions. On the other hand, ChatGPT can analyze the structure (in most cases), explain how electrons must move, but cannot provide a drawing (Table S1). Another important topic in organic chemistry is the ranking of acidity of molecules without having the pKa values. Drawing the conjugate bases and examining their stability is the key point to compare acidity. Both chatbots do not draw conjugate bases in their answers and sometimes misinterpreted the structures being analyzed. However, the factor affecting the acidity is correctly stated in the examples for both ChatGPT and DeepSeek but the outcomes of both are partially correct: either a problem in analyzing the molecule (for both) or choosing the wrong choice in MCQs as is the case of ChatGPT. Likewise, analyzing basicity of molecules faces similar problems. Predicting the product of organic reactions is an important task that a student should learn. The free radical substitution reaction of alkanes is one of the first organic reactions that we teach to students. Students have to draw the products and show the mechanism. It has been shown that both chatbots can correctly predict the product of such reaction by providing the name of the product but without drawing it. The free radical mechanism of the reaction, showing the initiation, propagation, and termination steps, is correctly provided by both chatbots but without the use of fish-hook arrows. The radical is placed on the wrong atoms in the structure in some cases. The latter example is provided for 2-methylbutane. Both chatbots cannot manage to analyze more complicated structures such as 1,1-dimethylcyclobutane or 2,2,3-trimethylbutane where a quaternary carbon is found in the structure. The products resulting from the substitution of the quaternary carbon is provided in both examples. Hence, once asked about the number or drawing the monohalogenated products, or the most stable radical obtained after removal of hydrogen, the wrong answer is provided. Table III summarizes the performance of ChatGPT and DeepSeek on the MCQs (n = 16) of the chemistry exam, along with paired statistical analysis. In this study, ChatGPT achieves a higher numerical accuracy on the MCQs than DeepSeek (37.5% vs. 25.0%), but the difference is not statistically significant (McNemar’s exact test, p = 0.63; OR = 3.0, 95% CI: 0.30–30.6). The wide confidence interval reflects the limited number of discordant items, indicating substantial uncertainty. These findings suggest a potential advantage of ChatGPT for MCQs on this exam, but they should be interpreted cautiously and cannot be generalized to other exams or question types.
Table III. MCQ performance and paired statistical comparison between ChatGPT and DeepSeek in the organic chemistry exam. Odds ratio and confidence interval were calculated from discordant pairs (ChatGPT correct/DeepSeek incorrect = 3; ChatGPT incorrect/DeepSeek correct = 1)
| Model | Correct (n) | Accuracy (%) | Statistical test | Effect size (95% CI) | p-value |
|---|---|---|---|---|---|
| ChatGPT | 6/16 | 37.5 | McNemar’s exact test | Odds ratio = 3.0 (95% CI: 0.30–30.6) | 0.63 |
| DeepSeek | 4/16 | 25.0 | — | — | — |
B.SOLVING INORGANIC CHEMISTRY EXAM
The inorganic chemistry exam covered 25 MCQs and several subjective problems, covering major topics including hybridization, bond order, molecular orbital (MO) theory, symmetry and point group assignment, magnetic properties, and periodic trends. The results of the questions are summarized in Table S2 (in supporting info). Overall, both ChatGPT and DeepSeek handle simple recall-based questions successfully but show consistent weaknesses when higher-level reasoning, orbital visualization, or symmetry analysis is required. The results are summarized in Table IV. Hybridization and simple bonding questions are among the best performed. For example, both chatbots correctly assign sp3 hybridization to phosphorus in PCl3 and sp hybridization to beryllium in BeCl2. However, ChatGPT frequently fails to connect hybridization with the correct geometry (e.g., describing BeCl2 as bent instead of linear). DeepSeek generally offers clearer justifications, though sometimes confusing electron domain geometry with molecular geometry. These issues parallel the organic chemistry results, where the models could recall terms but misapply them in structural reasoning. MO theory reveals much greater challenges. While both models correctly describe O2 as paramagnetic with a bond order of two, they struggle with less common diatomics such as Li2, He2, and Be2. ChatGPT tends to misapply the bond order formula, inflating the stability of unstable molecules. DeepSeek occasionally draws or describes MO diagrams correctly but gave contradictory bond orders. These inconsistencies highlight a gap in stepwise reasoning, similar to the organic part where resonance rules are recalled but not applied to actual structures. Symmetry and group theory are particularly problematic. In predicting IR and Raman activity of PCl3 vibrations, ChatGPT offers short answers without using character tables, while DeepSeek attempts symmetry-based reasoning but often misidentifies key symmetry elements. Both correctly recognize the inversion center in SF6, but analysis of lower-symmetry molecules such as trans-N2F2 is inconsistent, with ChatGPT missing a mirror plane while DeepSeek identifies most elements correctly. This mirrors the organic stereochemistry questions, where superficial recognition led to errors in spatial interpretation.
Table IV. Main topics covered in the inorganic chemistry exam with comments on the outcomes of the answers
| Question type | ChatGPT | DeepSeek |
|---|---|---|
| Hybridization (PCl3, BeCl2) | Correct hybridization, weak geometry link | Correct, better geometry link |
| σ and π bonds (C2H4) | Correct: 5σ + 1π | Correct |
| MO bond order (O2, Li2, He2, Be2) | O2 correct, others misapplied | O2 correct, others inconsistent |
| Symmetry/point groups (OF2, SF6, trans-N2F2) | OF2 and SF6 correct, trans-N2F2 incomplete | Slightly better in trans-N2F2 |
| IR/Raman activity (PCl3) | Short answer, no reasoning | Partial reasoning |
| Magnetic properties (O2, N2, others) | O2 and N2 correct, others inconsistent | O2 and N2 correct, others slightly better |
| Subjective (definitions, explanations) | Good on definitions, weak in applications | Good on definitions, better but still weak in applications |
Magnetic property questions show mixed results. Both chatbots correctly classify O2 as paramagnetic and N2 as diamagnetic, but ChatGPT sometimes ignores antibonding effects in borderline cases, while DeepSeek hesitates on spin states. Both recall definitions well (e.g., coordination number) but struggle with applied reasoning, such as explaining crystal field trends or coordination geometries, mirroring their organic chemistry performance.
On the 25 MCQs, ChatGPT answered 17 correctly (68.0%) and DeepSeek answered 21 correctly (84.0%). Paired comparison using McNemar’s exact test shows no statistically significant difference between the models (p = 0.72), with an odds ratio of 0.6 (95% CI: 0.12–3.0), reflecting a small, non-significant tendency favoring DeepSeek (Table V). The wide confidence interval highlights the uncertainty due to the limited number of discordant items, suggesting that while DeepSeek achieved higher numerical accuracy on this exam, the difference should be interpreted cautiously and cannot be generalized beyond this specific assessment.
Table V. MCQ performance and paired statistical comparison between ChatGPT and DeepSeek in the inorganic chemistry exam. Discordant pairs: ChatGPT correct/DeepSeek incorrect = 3; ChatGPT incorrect/DeepSeek correct = 5
| Model | Correct (n) | Accuracy (%) | Statistical test | Effect size (95% CI) | p-value |
|---|---|---|---|---|---|
| ChatGPT | 17/25 | 68 | McNemar’s exact test | Odds ratio = 0.6 (95% CI: 0.12–3.0) | 0.72 |
| DeepSeek | 21/25 | 84 | — | — | — |
C.SOLVING GENERAL CHEMISTRY EXAM
The results of the answers to questions are summarized in Table S3 in supporting information and in Table VI (general summary of topics). Both AI systems demonstrate strong performance in kinetics problems, for instance finding the rate law of a reaction using the method of the initial rate (Questions 1a–1e and 2a–2c), with perfect or near-perfect scores on most sub-questions. However, a critical difference emerges in Question 2b related to the calculation of the rate constant of the reaction from the half-life, where ChatGPT received 0 points despite providing correct analysis methodology, while DeepSeek achieved full marks with both correct analysis and accurate final answer. This discrepancy highlights the importance of computational precision in chemistry problem-solving, where conceptual understanding alone is insufficient without accurate mathematical execution. A significant weakness identified in ChatGPT is its difficulty handling large numbers and complex calculations, particularly evident in calculations using the Arrhenius equation (MCQ-4). The evaluators noted that ChatGPT has problems dealing with large numbers, while DeepSeek has no problem dealing with calculations and large numbers, and the steps of calculations are clear. The discrepancy may reflect differences in training methodologies or computational architectures between the models. The practical implications are significant, as chemistry frequently involves calculations with scientific notation, logarithms, and exponential functions that require high precision. The most striking performance differences occurred in atomic structure problems (Questions 3a–3d), where both AI systems received zero points across all sub-questions. This consistent failure suggests fundamental limitations in visual-spatial reasoning and diagram interpretation capabilities.
Table VI. Main topics covered in the general chemistry exam with comments on the outcomes of the answers
| Question type | ChatGPT | DeepSeek |
|---|---|---|
| Kinetics and mechanism | 1. Strong in initial rates method | 1. Strong in initial rates method |
| Atomic structure | Failure in analyzing orbital energy diagram | Failure in analyzing orbital energy diagram |
| Bohr Model | Good abilities in calculations involving calculation of energies and wavelength of transitions | Good abilities in calculations involving calculation of energies and wavelength of transitions |
| Electron configuration | Good capabilities in writing expanded and shorthand electron configuration | Good capabilities in writing expanded and shorthand electron configuration |
The evaluators note that both systems misinterpret the orbital energy diagram, with ChatGPT providing “1s22s22p4” instead of the correct electron configuration, while DeepSeek does not calculate the number of electrons in the p orbitals. This systematic error propagates through all related questions, highlighting the interconnected nature of chemistry concepts. The inability of both AI systems to accurately interpret orbital diagrams suggests a fundamental limitation that can significantly impact their utility in chemistry instruction and assessment. Both AI systems demonstrate strong performance on multiple-choice questions, with ChatGPT scoring 39/45 and DeepSeek scoring 42/45. This pattern aligns with previous research by Goorts et al. [47] who found that AI systems generally perform well on structured, multiple-choice assessments that rely on pattern recognition and knowledge recall. The consistent performance across MCQ items suggests that both systems possess robust foundational chemistry knowledge, despite the computational and visual-spatial challenges identified in other question types. This finding has important implications for the types of assessment tasks where AI systems may be most effectively employed. In calculations requiring finding the wavelength or energy from transitions in Bohr model dealing with transitions in hydrogen or hydrogen-like ions, both chatbots can successfully handle such problems. Good performance of both chatbots is observed as well in writing the electron configuration of elements or ions.
On the 15 MCQs, ChatGPT answered 13 correctly (86.7%) while DeepSeek answered all 15 correctly (100%). Paired comparison using McNemar’s exact test shows no statistically significant difference (p = 0.50), with an odds ratio of 0 (95% CI: 0–3.8), reflecting a small number of discordant items (Table VII). These results suggest a numerical advantage for DeepSeek on this exam, but the difference should be interpreted cautiously and cannot be generalized beyond this assessment.
Table VII. MCQ performance and paired statistical comparison between ChatGPT and DeepSeek in the general chemistry exam. Discordant pairs: ChatGPT correct/DeepSeek incorrect = 0; ChatGPT incorrect/DeepSeek correct = 2
| Model | Correct (n) | Accuracy (%) | Statistical test | Effect size (95% CI) | p-value |
|---|---|---|---|---|---|
| ChatGPT | 13/15 | 86.7 | McNemar’s exact test | Odds ratio = 0 (95% CI: 0–3.8) | 0.50 |
| DeepSeek | 15/15 | 100 | — | — | — |
D.SOLVING PHYSICAL CHEMISTRY EXAM
The tested physical chemistry exam was mainly discussing simple topics and fundamentals in physical chemistry including the first law of thermodynamics. The detailed study is summarized in Table S4. The main topics (MCQs and subjective questions) discussed in the exam are questions asking to calculate for work, heat, internal energy, and enthalpy (Table VIII). Both chatbots can successfully tackle such questions with the correct reasoning. Adding, both chatbots could manage to calculate enthalpy using Hess’s law and bond energies (Subjective questions). It is found that both AI platforms almost got a full mark (100% DeepSeek and 98 % for ChatGPT).
Table VIII. Main topics covered in the physical chemistry exam with comments on the outcomes of the answers
| Question Type | ChatGPT | DeepSeek |
|---|---|---|
| Very good skills | Very good skills | |
| Very good skills | Very good skills |
On the 21 MCQs, ChatGPT answered 20 correctly (95.2%) while DeepSeek answered all 21 correctly (100%). Paired comparison using McNemar’s exact test shows no statistically significant difference (p = 1.00), with an odds ratio of 0 (95% CI: 0–5.3), reflecting only a single discordant question (Table IX). These results suggest a numerical advantage for DeepSeek on this exam, but the difference is minimal and should be interpreted cautiously.
Table IX. MCQ performance and paired statistical comparison between ChatGPT and DeepSeek in the physical chemistry exam. Discordant pairs: ChatGPT correct/DeepSeek incorrect = 0; ChatGPT incorrect/DeepSeek correct = 1
| Model | Correct (n) | Accuracy (%) | Statistical test | Effect size (95% CI) | p-value |
|---|---|---|---|---|---|
| ChatGPT | 20/21 | 95.2 | McNemar’s exact test | Odds ratio = 0 (95% CI: 0–5.3) | 1.00 |
| DeepSeek | 21/21 | 100 | — | — | — |
V.CLASSIFICATION OF THE EXAM QUESTIONS ACCORDING TO BLOOM’S TAXONOMY
To reduce cheating in online chemistry exams and help instructors quickly generate alternative questions of varying difficulty, ChatGPT and DeepSeek classified exam questions using Bloom’s taxonomy and then rephrased them across categories to build a question bank. The study covered general and organic chemistry exams, with detailed results in the supporting information (Tables S5–S6).
A.CLASSIFICATION OF THE ORGANIC CHEMISTRY EXAM QUESTIONS ACCORDING TO BLOOM’S TAXONOMY
No questions were classified as “remember.” GPT classified one as “understand,” both chatbots classified six as “apply,” and DeepSeek classified six as “evaluate.” Most questions fell under “analyze” (20 by GPT, 15 by DeepSeek). Results indicate incomplete coverage of Bloom’s taxonomy, suggesting improvements for future exams. The chatbots agree on ∼60% of classifications; most differences are minor and likely due to image-analysis limits or classification errors. Some application-level questions (e.g., optical activity, Newman projections, and chiral configurations) are occasionally misclassified as analysis. GPT shows better ability to adjust question levels across Bloom’s taxonomy, while DeepSeek often keeps levels unchanged except at higher categories. For resonance structure questions, GPT-4 appropriately rephrases items at the remember and understand levels, with little change between apply and analyze, and higher difficulty at evaluate and create levels. DeepSeek instead simplifies the structure to the nitrite ion (), likely misinterpreting the original question. Although difficulty increased across categories, it was based on an incorrect structure. For acidity comparisons within a molecule, GPT-4 effectively adjusts question difficulty across Bloom’s categories. In contrast, DeepSeek shows little variation from remember to apply, repeatedly asking for acidity ranking and conjugate-base stability. Similar trends appear in MCQs, where GPT-4 better matches question difficulty to category, while DeepSeek largely maintains the same level.
B.CLASSIFICATION OF THE GENERAL CHEMISTRY EXAM QUESTIONS ACCORDING TO BLOOM’S TAXONOMY
In the general chemistry exam, ChatGPT-4.o outperforms DeepSeek in classifying and generating questions across Bloom’s taxonomy, especially at higher cognitive levels. Both perform similarly well at the remember and understand levels, aligning with prior findings on AI strength in recall and basic comprehension. [32]. ChatGPT shows better contextual explanations that support understanding, while DeepSeek emphasizes procedural accuracy. The largest differences appear at higher cognitive levels, where ChatGPT generates more authentic, complex tasks. It creates realistic application scenarios and stronger critical thinking questions at the analyze and evaluate levels, while DeepSeek tends toward more formulaic tasks, an important finding for AI support of higher-order thinking in education [48]. At the create level, the gap is greatest, with ChatGPT generating tasks requiring experimental design, modeling, and concept synthesis.
VI.CONCLUSION
Both ChatGPT and DeepSeek performed well on undergraduate chemistry assessments, with variation across subdisciplines. DeepSeek slightly excelled in general chemistry, while both models succeeded in physical chemistry but struggled in organic and moderately in inorganic chemistry. Weaknesses appeared in tasks requiring numerical reasoning, spatial analysis, and diagram interpretation. ChatGPT better adjusted question difficulty across Bloom’s levels. These findings highlight both the potential and current limitations of LLMs in chemistry education. Future work will evaluate newer models, course-aligned problems, and comparisons with other chatbots to support higher-order learning outcomes.