The Assessment of ChatGPT and DeepSeek to Solve Chemistry Exams and the Classification of the Exam Questions According to Bloom’s Taxonomy in order to Alter the Level of Exams. A Comparative Study
The Assessment of ChatGPT and DeepSeek to Solve Chemistry Exams and the Classification of the Exam Questions According to Bloom’s Taxonomy in order to Alter the Level of Exams. A Comparative Study
Authors
Ali Ghadban
Chemical Sciences Laboratory (CSL@LIU), Department of Biological and Chemical Sciences, School of Arts and Sciences, Lebanese International University, Beirut, P.O. Box 146404, Lebanon; Department of Chemistry and Biochemistry, Faculty of Sciences, Lebanese University, Zahle, Lebanon † Authors contributed equally to this work.
Fatima Al Khatib
Chemical Sciences Laboratory (CSL@LIU), Department of Biological and Chemical Sciences, School of Arts and Sciences, Lebanese International University, Beirut, P.O. Box 146404, Lebanon
https://orcid.org/0009-0000-5091-4183
Hanan Rahal
Chemical Sciences Laboratory (CSL@LIU), Department of Biological and Chemical Sciences, School of Arts and Sciences, Lebanese International University, Beirut, P.O. Box 146404, Lebanon
https://orcid.org/0000-0002-1516-5049
Sanaa Khaled
Chemical Sciences Laboratory (CSL@LIU), Department of Biological and Chemical Sciences, School of Arts and Sciences, Lebanese International University, Beirut, P.O. Box 146404, Lebanon
https://orcid.org/0000-0002-6403-1194
This study examines the performance of ChatGPT and DeepSeek in solving undergraduate chemistry exams (general, organic, inorganic, and physical) and their ability to classify questions by Bloom’s Taxonomy to adjust cognitive levels. Both models perform well on general chemistry, with DeepSeek scoring about eight points higher, and both succeed on direct calculation tasks (e.g., kinetics, atomic structure, and Bohr model). ChatGPT struggles with complex numerical problems (e.g., Arrhenius equation) and does not handle diagram based tasks (e.g., orbital diagrams). In physical chemistry, both achieve near-perfect scores on thermodynamics questions, while performance on organic chemistry is low (38% for ChatGPT vs. 23% for DeepSeek), reflecting difficulty with stereochemistry, acid-base chemistry, resonance, and spatial structures. In inorganic chemistry, both perform moderately (62% vs. 56%), solving simple but not symmetry or group theory-related problems. Statistical analyses across multiple-choice questions show that differences between models are small and not statistically significant for any exam: organic chemistry (37.5% vs. 25.0%, p = 0.63), inorganic (84% vs. 68%, p = 0.72), general (100% vs. 86.7%, p = 0.50), and physical chemistry (100% vs. 95.2%, p = 1.00), indicating only potential advantages that cannot be generalized. Both classify most questions correctly by Bloom’s levels, though errors occur at the apply/analyze levels. ChatGPT is more effective in adjusting question difficulty by moving through Bloom’s verbs.
Author Biographies
Fatima Al Khatib, Chemical Sciences Laboratory (CSL@LIU), Department of Biological and Chemical Sciences, School of Arts and Sciences, Lebanese International University, Beirut, P.O. Box 146404, Lebanon
Assistant Professor
Chemical Sciences Laboratory (CSL@LIU), Department of Biological and Chemical Sciences, School of Arts and Sciences, Lebanese International University, Beirut, P.O Box 146404, Lebanon
Hanan Rahal, Chemical Sciences Laboratory (CSL@LIU), Department of Biological and Chemical Sciences, School of Arts and Sciences, Lebanese International University, Beirut, P.O. Box 146404, Lebanon
Assistant Professor
Chemical Sciences Laboratory (CSL@LIU), Department of Biological and Chemical Sciences, School of Arts and Sciences, Lebanese International University, Beirut, P.O Box 146404, Lebanon
Sanaa Khaled, Chemical Sciences Laboratory (CSL@LIU), Department of Biological and Chemical Sciences, School of Arts and Sciences, Lebanese International University, Beirut, P.O. Box 146404, Lebanon
Assistant Professor
Chemical Sciences Laboratory (CSL@LIU), Department of Biological and Chemical Sciences, School of Arts and Sciences, Lebanese International University, Beirut, P.O Box 146404, Lebanon
Ghadban, A., Al Khatib, F., Rahal, H., & Khaled, S. (2026). The Assessment of ChatGPT and DeepSeek to Solve Chemistry Exams and the Classification of the Exam Questions According to Bloom’s Taxonomy in order to Alter the Level of Exams. A Comparative Study. Journal of Artificial Intelligence and Technology, 6, 646–653. https://doi.org/10.37965/jait.2026.0942
We are thrilled to announce that the CiteScore 2024 for the Journal of Artificial Intelligence and Technology is 11.0, which ranks 62 out of 450 journals in the Artificial Intelligence category.
We are delighted to announce that the CiteScore 2023 for the Journal of Artificial Intelligence and Technology is 8.7, which ranks it 77 out of 350 journals in the Artificial Intelligence category.
The Journal of Artificial Intelligence and Technology has been accepted for inclusion in Scopus, the world’s largest abstract and citation database of peer reviewed literature. The acceptance for inclusion by Scopus underlines the journal’s consistent high quality.
We are pleased to announce that the Journal of Artificial Intelligence and Technology is now indexed in IET Inspec, the definitive engineering, physics, and computer science research database.
We are delighted to announce that ORCID has been integrated into submission system. It's encouraged that authors can authorize our journal's submission system to obtain your ORCID ID. We would send e-mail to request ORCID authorization from contributor before publication.