IIHS Undergraduate Researchers Publish Timely Study on ChatGPT’s Role in Higher Education

Student-led research reveals that more advanced AI models may be increasingly accurate, but incomplete information and poorly framed questions remain significant educational challenges
Welisara, Sri Lanka – 5 August 2026: The International Institute of Health Sciences (IIHS) is pleased to announce the publication of an important student-led research article examining the accuracy, completeness and educational reliability of ChatGPT when responding to microbiology-related questions.
The article, titled “Assessment of the efficacy of ChatGPT responses to bacterial species-specific questions in microbiology,” was published in Access Microbiology, a peer-reviewed journal of the Microbiology Society, as part of its Pedagogy collection. It was published online on 4 August 2026.
The research was undertaken by a team of undergraduate Biomedical Science students from IIHS:
- Withanage Dona Manushi Dinasha
- Nissanka Mudiyanselage Tanuri Ayanga Nissanka
- Chamudhi Prabashi Wickramasinghe
- Warnakulasuriya Palakuttige Pasindu Damsara Fernando
The students conducted the study under the academic guidance of Mr. Gayan Gunatilake, Head of the Department of Biomedical Science at IIHS, and with the specialist mentorship of Dr. Vindya Perera, External Consultant and academic from the Department of Microbiology, Faculty of Medicine, Sabaragamuwa University of Sri Lanka.
Addressing a critical issue in the age of generative AI
Large language models such as ChatGPT are becoming increasingly capable and are now widely used by students to obtain explanations, summarise information and support academic work. However, the IIHS study demonstrates that increasing model accuracy does not automatically guarantee that every response will be sufficiently complete, appropriately contextualised or suitable for academic and health-science education.
A particularly important challenge identified by the research was the prevalence of mixed or incomplete information. An AI-generated response may contain several correct statements while still omitting essential concepts, terminology, qualifications or context required for a complete academic answer.
This can be more difficult for students to detect than an obviously incorrect response. Because the information appears credible and is presented confidently, a learner may accept it as complete without recognising what has been excluded.
The study therefore highlights an important distinction for higher education: information can be factually correct in part without being educationally sufficient.
Newer models improved accuracy, but incompleteness remained substantial
The research compared responses generated by ChatGPT 3.5 and ChatGPT 4.1 mini. The questions covered 15 bacterial species, with 18 questions developed for each organism. Questions were structured across low, moderate and high language-proficiency levels to simulate differences in how learners may communicate with an AI system.
Responses were evaluated against an established microbiology reference guide and categorised as accurate, mixed or incomplete, or inaccurate.
The results showed that ChatGPT 4.1 mini performed better than ChatGPT 3.5:
- 56.3% of ChatGPT 4.1 mini responses were classified as accurate, compared with 40.4% for ChatGPT 3.5.
- Mixed or incomplete responses declined from 58.1% with ChatGPT 3.5 to 43.2% with ChatGPT 4.1 mini.
- Inaccurate responses were comparatively uncommon, accounting for 1.5% of ChatGPT 3.5 responses and 0.5% of ChatGPT 4.1 mini responses.
These findings demonstrate clear progress between the models. However, they also reveal why accuracy percentages alone should not determine whether an AI system is educationally reliable. Even with the better-performing model, more than four in every ten responses were assessed as mixed or incomplete.
The major risk identified by the research was therefore not simply the generation of entirely false information. It was the production of answers that appeared plausible and contained correct information but failed to provide all the details necessary for a complete understanding.
The quality of the question influences the quality of the answer
The research also found that the specificity, clarity and language proficiency of the question affected the quality of ChatGPT’s response.
More precise questions were more likely to generate accurate answers, while imprecise questions frequently produced only partially correct or incomplete responses. Questions formulated at a higher proficiency level also demonstrated a greater proportion of accurate outputs.
This finding is highly significant for higher education. Large language models generate answers based on the context and instructions supplied by the user. When a student provides an overly broad, ambiguous or poorly structured request, the system may not identify the exact depth, academic level, disciplinary focus or scope required.
For example, asking an AI tool to “explain a bacterium” may result in a general answer. A more specific request asking for its morphology, transmission, virulence factors, clinical manifestations, diagnostic methods, treatment and prevention is more likely to produce a structured and comprehensive response.
However, even a well-written prompt does not guarantee complete accuracy. AI-generated information must still be evaluated against textbooks, peer-reviewed literature, clinical guidelines and other authoritative sources.
The research therefore suggests that responsible AI use requires a combination of:
- Relevant subject knowledge;
- The ability to formulate clear and specific questions;
- Critical evaluation of the response;
- Recognition of missing information;
- Verification using authoritative academic sources; and
- Appropriate guidance from educators.
Moving beyond simple “right or wrong” evaluations
The findings challenge the assumption that AI outputs should be evaluated only as correct or incorrect.
In an educational setting, an answer can be broadly accurate but still be inadequate because it omits a mechanism, exception, key technical term, diagnostic consideration or important contextual factor. Such omissions may affect examination performance, conceptual understanding and, in health-related disciplines, future professional decision-making.
As LLMs continue to improve, completely inaccurate responses may become less frequent. Nevertheless, partially correct and incomplete responses may remain a major concern because they are less visible and more likely to be accepted without further scrutiny.
This makes AI literacy increasingly important. Students must be taught not only how to use AI tools but also how to question their outputs by asking:
- What information is missing?
- Does the response address the full scope of the question?
- Are important scientific terms or qualifications absent?
- Is the answer appropriate for the required academic level?
- Can each major claim be verified using a credible source?
The study consequently supports an educational approach in which AI is treated as a tool for inquiry and discussion rather than as an unquestioned source of knowledge.
Demonstrating the value of undergraduate research at IIHS
The publication is a significant achievement because the research was carried out by undergraduate students with structured academic guidance and specialist mentorship.
Through the project, the students gained practical experience in identifying an emerging educational problem, reviewing scholarly literature, developing research questions, designing a methodology, collecting and evaluating data, interpreting findings, preparing an academic manuscript and participating in the peer-review and publication process.
Rather than interacting with ChatGPT only as users, the students were given the opportunity to examine the technology scientifically and evaluate its strengths and limitations through a structured research process.
This reflects IIHS’s commitment to hands-on, research-informed learning and to providing students with meaningful exposure to research, innovation and academic publication from an early stage of their education.
The project also demonstrates that undergraduate students, when provided with suitable supervision and opportunities, can make valuable contributions to current international discussions concerning artificial intelligence, health-science education and the future of higher education.
Preparing students for an AI-supported future
The study does not suggest that ChatGPT or other LLMs should replace lecturers, textbooks, peer-reviewed evidence or professional judgement. Instead, it highlights the need for educational institutions to prepare students to use these tools responsibly.
The future of higher education will require more than access to increasingly advanced technology. It will require learners who can provide appropriate context, ask precise questions, evaluate the completeness of an answer, identify omissions and independently verify information.
Through research of this nature, IIHS continues to develop students who are not merely consumers of technology but critical investigators capable of evaluating and contributing to its responsible use.
Reference and full article:https://www.microbiologyresearch.org/content/journal/acmi/10.1099/acmi.0.001137.v5


